mini-beatrix-3

The 32-block rung of the mini-beatrix ladder, with its stage arms. 376M parameters, byte-level (vocab 256), with a multi-constellation CausalSplatHUB (signed-address linear attention over learned codebook blackboards) and an anchored expert bank in every one of its 32 blocks. No softmax-over-positions attention anywhere. Trained on 64.4B bytes (245,674 steps, about 295 hours on two RTX 5090s) through a staged curriculum, completed 2026-10-05.

The package ships both states in one repository: the bare model (arms off) and the model with its stage arms mounted (arms on). An arm is a small detachable adapter trained on this exact frozen core. The arms come as groups: arms that were trained switched on together and are mounted together. There are two, the four arms of stages 1 to 4 and the eight arms of stages 1 to 8 (those four, held fixed, with four more fitted over them), and an experimental third: the eight, held fixed, with a ninth arm for image captions fitted over them, a prototype that is listed and mountable like the others but is not the default. Each stage also has a solo arm, trained alone, for use one at a time. Arms are guests on the core, never a change to it: the core's weights are the same files either way, and detaching restores the bare model bit for bit.

import torch
from transformers import AutoModelForCausalLM

m = AutoModelForCausalLM.from_pretrained(
        "AbstractPhil/mini-beatrix-3", trust_remote_code=True).eval()

m.say("Who are you?")            # arms off: the bare core, its own chat frame

m.mount_arm("stages-1-8")         # arms on: the stage arms, all on together
m.detach_arm()                   # the bare core again, verified bit-exact

Arms on from the first call:

m = AutoModelForCausalLM.from_pretrained(
        "AbstractPhil/mini-beatrix-3", trust_remote_code=True,
        default_arm="stages-1-8").eval()

The model reads raw UTF-8 bytes: input_ids are byte values 0–255. There is no tokenizer to download.

Architecture

  • d_model 1024 Β· 32 layers Β· 16 heads Β· ctx 4096 Β· byte-trigram embedding (raw UTF-8 bytes; input ids are byte values 0–255)
  • Hubs (all 32 blocks): 4 constellations Γ— 64 anchors @ D=128 per block, read by an exact chunked scan (chunk 256) and budget-composed (numerators and agreement masses sum before one divide: reconstructive, never comparative, with no argmax and no top-k). Constant-size prefix state: each layer encodes the sequence onto a fixed-width addressed blackboard rather than caching it.
  • Anchored banks: 3 full-width experts per block (ff 1024), signed dispatch, expert outputs born at zero.
  • Dual head: linear readout + a signed aleph read (256 anchors @ 256), fitted to the first batch at step 0 so that it works from birth (8.26 β†’ 5.45 bpb on that batch).
  • Held-out cost of switching a mechanism off (the toggle ledger, in bpb), at the last boundary before the end and at the end: hubs +5.86 and +6.41 Β· head +1.94 and +1.90 Β· banks +3.40 at the former. The banks' reading at the end (+6.13) is nearly twice every earlier one and is being repeated before it is relied on.

The stage arms

The curriculum taught the model nine kinds of text, one stage after another, and the finished model moved on from each stage's text as the later stages came. The stage arms bring that text back without touching the core. They were fitted on the finished, frozen core as always-on groups: the arms of stages 1 and 2 were trained first and then held fixed, the arms of stages 3 and 4 were attached over them one after the other and trained in their presence, and then, with all four held fixed and on, the arms of stages 5, 6, 7 and 8 were attached over them the same way, one after the other, each for 4,000 steps and each kept quiet on every other arm's stage text. The eight-arm group's first four arms are the four-arm group unchanged; the four-arm group is also packaged on its own. m.arms() returns the tables below as data, with each row's recipe and full measurements.

The group stages-1-4 (4 arms, 54.8M parameters; two seeds):

arm its stage's text stage text, arms off β†’ the group on (bpb) this arm alone (bpb) stage items in the stage's own form, off β†’ group on stage items in a new form, off β†’ group on
stages-1-4/s1_perspective one small event retold from each side: seen from a person, said to them, told about them 0.554 β†’ 0.048 0.052 55% β†’ 96% 46% β†’ 48%
stages-1-4/s2_concept kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ 0.721 β†’ 0.112 0.350 30% β†’ 98% 68% β†’ 78%
stages-1-4/s3_rules if-then rules over made-up words, followed step by step to what follows 0.930 β†’ 0.224 0.824 16% β†’ 90% 7% β†’ 30%
stages-1-4/s4_arith small arithmetic worked out in text 1.234 β†’ 0.338 1.127 34% β†’ 61% 24% β†’ 24%

With the group on, held-out web text moves by +0.0017 bpb (the limit set for it was +0.012), and the nine-suite probe mean reads 0.411 with arms off and 0.419 with the group on.

The numbers are those of the packaged weights. A second run of the same recipe from other seeds read the stage texts with its group on at 0.048, 0.121, 0.222, 0.338 bpb and moved web text by +0.0011.

Each member mounted alone, and the whole group, on every stage's held-out text (bpb):

mounted stage 1 text stage 2 text stage 3 text stage 4 text web text
nothing (arms off) 0.554 0.721 0.930 1.234 0.951
only stages-1-4/s1_perspective 0.053 0.239 0.509 0.723 0.952
only stages-1-4/s2_concept 0.527 0.350 0.619 0.880 0.952
only stages-1-4/s3_rules 0.554 0.724 0.824 1.234 0.951
only stages-1-4/s4_arith 0.553 0.721 0.928 1.127 0.951
the whole group stages-1-4 0.048 0.112 0.224 0.338 0.952

The group stages-1-8 (8 arms, 109.5M parameters; two seeds):

arm its stage's text stage text, arms off β†’ the group on (bpb) this arm alone (bpb) stage items in the stage's own form, off β†’ group on stage items in a new form, off β†’ group on
stages-1-8/s1_perspective one small event retold from each side: seen from a person, said to them, told about them 0.554 β†’ 0.049 0.052 55% β†’ 89% 46% β†’ 44%
stages-1-8/s2_concept kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ 0.721 β†’ 0.111 0.350 30% β†’ 98% 68% β†’ 72%
stages-1-8/s3_rules if-then rules over made-up words, followed step by step to what follows 0.930 β†’ 0.224 0.824 16% β†’ 86% 7% β†’ 31%
stages-1-8/s4_arith small arithmetic worked out in text 1.234 β†’ 0.328 1.127 34% β†’ 64% 24% β†’ 24%
stages-1-8/s5_causal short cause-and-effect text: what happened and why 0.660 β†’ 0.043 0.091 not measured not measured
stages-1-8/s6_tryfail an attempt, its failure and the revised attempt 0.673 β†’ 0.055 0.164 not measured not measured
stages-1-8/s7_mixed the mixed stage's diet: the earlier stages' text beside ordinary prose 0.783 β†’ 0.720 0.724 not measured not measured
stages-1-8/s8_register the same content in different registers of speech 0.491 β†’ 0.021 0.068 not measured not measured

With the group on, held-out web text moves by +0.0019 bpb (the limit set for it was +0.012), and the nine-suite probe mean reads 0.411 with arms off and 0.422 with the group on.

The numbers are those of the packaged weights. A second run of the same recipe from other seeds read the stage texts with its group on at 0.048, 0.121, 0.222, 0.327, 0.043, 0.054, 0.720, 0.021 bpb and moved web text by +0.0010.

Each member mounted alone, and the whole group, on every stage's held-out text (bpb):

mounted stage 1 text stage 2 text stage 3 text stage 4 text stage 5 text stage 6 text stage 7 text stage 8 text web text
nothing (arms off) 0.554 0.721 0.930 1.234 0.660 0.673 0.783 0.491 0.951
only stages-1-8/s1_perspective 0.053 0.239 0.509 0.723 0.463 0.495 0.773 0.379 0.952
only stages-1-8/s2_concept 0.527 0.350 0.619 0.880 0.659 0.673 0.783 0.469 0.952
only stages-1-8/s3_rules 0.554 0.724 0.824 1.234 0.668 0.674 0.782 0.491 0.951
only stages-1-8/s4_arith 0.553 0.721 0.928 1.127 0.657 0.659 0.783 0.489 0.951
only stages-1-8/s5_causal 0.536 0.710 0.917 1.225 0.091 0.666 0.782 0.484 0.951
only stages-1-8/s6_tryfail 0.552 0.722 0.930 1.237 0.651 0.164 0.783 0.490 0.951
only stages-1-8/s7_mixed 0.497 0.715 0.868 1.195 0.558 0.593 0.724 0.483 0.950
only stages-1-8/s8_register 0.537 0.713 0.925 1.227 0.642 0.651 0.783 0.068 0.951
the whole group stages-1-8 0.049 0.110 0.224 0.328 0.043 0.055 0.720 0.021 0.952

The experimental group stages-1-8-caption (9 arms, 123.2M parameters; two seeds; experimental):

Experimental: a prototype for the diffusion line, not a curriculum stage. The ninth arm was fitted for 4,000 steps over the eight (held fixed and on) on image captions, text the core never trained on: five sources, each image one document in a labelled frame (tags, caption, short, attributes, scene, photo, structure, medium, long, rating). With all nine on, held-out caption rows read 0.783 bits per byte against 1.265 with arms off (0.780 on the second seed); the arm alone reads 0.785, so it carries the whole gain itself, and the eight read caption rows like the bare model (within 0.001). With it on, the eight's stage texts lose at most 0.002, web text moves by +0.003 (+0.002 on the second seed; the ninth arm's own share under 0.001) and its largest write on a partner's stage text is +0.002. Its training curve was still falling at 4,000 steps (0.708 to 0.693 over the last 800, where every stage arm had flattened) and its gate opens about twice as wide as any stage arm's (0.45 against 0.11 to 0.28), so a longer arm is an open item. What it does to generated captions, and to diffusion conditioning, is not measured here.

arm its stage's text stage text, arms off β†’ the group on (bpb) this arm alone (bpb) stage items in the stage's own form, off β†’ group on stage items in a new form, off β†’ group on
stages-1-8-caption/s1_perspective one small event retold from each side: seen from a person, said to them, told about them 0.554 β†’ 0.050 0.052 55% β†’ 86% 46% β†’ 44%
stages-1-8-caption/s2_concept kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ 0.721 β†’ 0.111 0.350 30% β†’ 96% 68% β†’ 70%
stages-1-8-caption/s3_rules if-then rules over made-up words, followed step by step to what follows 0.930 β†’ 0.225 0.824 16% β†’ 84% 7% β†’ 29%
stages-1-8-caption/s4_arith small arithmetic worked out in text 1.234 β†’ 0.329 1.127 34% β†’ 65% 24% β†’ 25%
stages-1-8-caption/s5_causal short cause-and-effect text: what happened and why 0.660 β†’ 0.043 0.091 not measured not measured
stages-1-8-caption/s6_tryfail an attempt, its failure and the revised attempt 0.673 β†’ 0.055 0.164 not measured not measured
stages-1-8-caption/s7_mixed the mixed stage's diet: the earlier stages' text beside ordinary prose 0.783 β†’ 0.722 0.724 not measured not measured
stages-1-8-caption/s8_register the same content in different registers of speech 0.491 β†’ 0.022 0.068 not measured not measured
stages-1-8-caption/s9_caption image captions in one labelled frame (tags, the source's caption, short, attributes, scene, photo, structure, medium, long, rating: one label per line): photographs with structured captions, CC12M alt text and generated descriptions, COCO captions, danbooru tags 1.265 β†’ 0.783 0.785 not measured not measured

With the group on, held-out web text moves by +0.0027 bpb (the limit set for it was +0.012), and the nine-suite probe mean reads 0.411 with arms off and 0.422 with the group on.

The numbers are those of the packaged weights. A second run of the same recipe from other seeds read the stage texts with its group on at 0.049, 0.122, 0.223, 0.327, 0.043, 0.055, 0.722, 0.021, 0.779 bpb and moved web text by +0.0017.

Each member mounted alone, and the whole group, on every stage's held-out text (bpb):

mounted stage 1 text stage 2 text stage 3 text stage 4 text stage 5 text stage 6 text stage 7 text stage 8 text captions text web text
nothing (arms off) 0.554 0.721 0.930 1.234 0.660 0.673 0.783 0.491 1.265 0.951
only stages-1-8-caption/s1_perspective 0.053 0.239 0.509 0.723 0.463 0.495 0.773 0.379 1.266 0.952
only stages-1-8-caption/s2_concept 0.527 0.350 0.619 0.880 0.659 0.673 0.783 0.469 1.266 0.952
only stages-1-8-caption/s3_rules 0.554 0.724 0.824 1.234 0.668 0.674 0.782 0.491 1.265 0.951
only stages-1-8-caption/s4_arith 0.553 0.721 0.928 1.127 0.657 0.659 0.783 0.489 1.264 0.951
only stages-1-8-caption/s5_causal 0.536 0.710 0.917 1.225 0.091 0.666 0.782 0.484 1.264 0.951
only stages-1-8-caption/s6_tryfail 0.552 0.722 0.930 1.237 0.651 0.164 0.783 0.490 1.265 0.951
only stages-1-8-caption/s7_mixed 0.497 0.715 0.868 1.195 0.558 0.593 0.724 0.483 1.264 0.950
only stages-1-8-caption/s8_register 0.537 0.713 0.925 1.227 0.642 0.651 0.783 0.068 1.264 0.951
only stages-1-8-caption/s9_caption 0.552 0.723 0.925 1.230 0.660 0.665 0.786 0.519 0.785 0.952
the whole group stages-1-8-caption 0.050 0.111 0.224 0.329 0.043 0.055 0.722 0.022 0.783 0.953

How to read the tables. Stage text is the loss, in bits per byte, on held-out text of the arm's own stage, with arms off and with the whole group on. This arm alone is the same loss with only that one arm mounted. Stage items are short test questions from that stage, scored as the share answered correctly, first in the form the stage text uses and then in a form it does not. The web text figure is the change on held-out ordinary web text (fineweb-edu) with the group on: the arms were trained to leave it alone.

A group is the unit. The eight-arm group is the four-arm group with four more arms fitted over it, and it keeps the four as they were: with all eight on, the text of stages 1 to 4 reads within 0.011 bits per byte of the four-arm group's own reading (stage 4 slightly better, the others the same). The new arms take their stages: with the whole group on, stage 5 text reads 0.043 bits per byte against 0.660 bare, stage 6 0.055 against 0.673, stage 8 0.021 against 0.491, on both seeds. Each of those three carries about two thirds to three quarters of its stage's gain itself and the fixed earlier arms supply the rest (the first arm, a generalist, reads stage 5 and 6 text 0.2 and 0.18 below bare on its own), so again the group works as a whole and is measured that way. The mixed stage (7) is the exception: its text is a blend of the other stages' kinds, the bare model reads it at 0.783, and no arm moves it much (the group reads it at 0.720, its own arm alone at 0.724). The group costs 0.001 to 0.002 bits per byte on ordinary web text, no arm costs more than 0.0006 of that, and no arm writes on another arm's stage text above 0.0012. Nothing an earlier arm learned was lost when a later arm trained over it (the largest loss of an earlier stage's gain at any later close was 2 percent). The four new arms trained for 4,000 steps each, about where a solo arm's curve flattens; the first four kept their shorter training (3,200, 2,400, 800 and 800 steps). For one stage by itself, use its solo arm below.

What the arms do, plainly. A group recovers each of its stages' own kind of text. The gain is in each stage's own forms. Whether it carries over to new forms is shown in each table's last column, where it was measured.

Arms off and arms on

m.arm                        # None: arms off
h = m.mount_arm("stages-1-8")       # arms on: the whole eight-arm group
with h.all_off():            # every member masked: the bare core's logits
    m.say("Hello there.")
m.detach_arm(verify=True)    # raises if the restored core is not bit-exact
m.mount_arm("stages-1-4/s1_perspective")       # one member alone (the "alone" readings)

A group mounts its members in the order they were trained, always on, one after another at every block, with no mixer between them. default_arm in from_pretrained (or in config.json) mounts a group as the model loads. save_pretrained refuses while an arm is mounted, so an arm can never be written into the core's weight file.

Solo arms

Each stage also has an arm that was trained alone on the same frozen core, for when one stage's text is all that matters. A solo arm holds its whole stage by itself.

arm its stage's text stage text, arm off β†’ on (bpb) web text (bpb) stage items in the stage's own form stage items in a new form seeds
solo/s1_perspective one small event retold from each side: seen from a person, said to them, told about them 0.554 β†’ 0.047 (a second seed: 0.047) +0.0003 55% β†’ 89% 46% β†’ 42% two
solo/s2_concept kinds, properties and differences: what a thing is a kind of, what its kind can do, how two things differ 0.721 β†’ 0.075 (a second seed: 0.075) +0.0001 30% β†’ 100% 68% β†’ 70% two
solo/s3_rules if-then rules over made-up words, followed step by step to what follows 0.930 β†’ 0.180 (a second seed: 0.179) +0.0003 16% β†’ 82% 7% β†’ 22% two
solo/s4_arith small arithmetic worked out in text 1.234 β†’ 0.267 (a second seed: 0.267) +0.0002 34% β†’ 83% 24% β†’ 21% two
solo/s5_causal short cause-and-effect text: what happened and why 0.660 β†’ 0.043 (a second seed: 0.043) +0.0006 not measured not measured two
solo/s6_tryfail an attempt, its failure and the revised attempt 0.673 β†’ 0.057 (a second seed: 0.057) +0.0003 not measured not measured two
solo/s7_mixed the mixed stage's diet: the earlier stages' text beside ordinary prose 0.783 β†’ 0.715 (a second seed: 0.717) +0.0000 not measured not measured two
solo/s8_register the same content in different registers of speech 0.491 β†’ 0.021 (a second seed: 0.021) +0.0003 not measured not measured two
m.mount_arm("solo/s1_perspective")  # one solo arm; mounting another swaps it

Solo arms are for use one at a time. Each was trained to leave ordinary web text alone, but nothing taught it to stay out of the other stages' text. When the four solo arms of stages 1 to 4 were switched on together, three of the four stage texts read worse than with no arm at all. To have several stages on at once, mount the group: its arms were trained together for exactly that.

What an arm is

Each arm places one module after every decoder block. At a site it projects the block's output into 16 slots of dimension 8, reads them through a 16-atom aleph address (a closed-form signed read, never a selector), and adds a gated patch back into the residual stream:

slots   = proj(x).view(B, n, 16, 8)
m_hat   = sum_k sinh(u_k) A_k / sum_k cosh(u_k),  u = (x_hat . A) / tau
x       = x + sigmoid(gate) * consume(m_hat)

13.7M parameters per arm over the 32 blocks. Every arm was trained with a quiet term: beside each chunk of its own stage's text it saw a chunk of text that is not its own and was penalised for changing the model's predictions there. For a group arm that text was ordinary web text and a partner arm's stage text, which is why the group can be left on and why no member harms a partner's stage. For a solo arm it was web text only.

Detach is exact

Mounting swaps the model's block list for wrapped blocks and keeps the originals. detach_arm() puts the originals back and re-runs a fixed probe: the restored logits must equal the pre-mount fingerprint bit for bit, or it raises. Before release, every packaged row was checked to produce logits identical to the adapter library the arms were trained with (amoe-lora).

The special-token control plane

Thirteen ids that valid UTF-8 can never produce carry structure and are trained:

id token meaning
0xFF DOC document boundary (taught from step 0)
0xFE / 0xFD USER / MODEL turn openers (taught in the final chat phase)
0xFC END universal block close
0xFB SYS system block opener
0xF7+b MODE register tag (+1 ASCII byte)
0xF5+b ESC 254 extended slots
0xFA 0xF9 0xF8 0xF6 0xC0 0xC1 THINK DATA SEP CUE RES reserved/instrument

Chat format: [SYS] text [END] [USER] text [END] [MODEL] text [END]. The frame is unforgeable (encoded text cannot contain a special). m.say() renders it for you; the chat frame was taught in the final 4.0B bytes, so treat it as a young capability.

Plain generation

prompt = "The history of astronomy begins"
ids = torch.tensor([list(prompt.encode("utf-8"))])
out = m.generate(ids, max_new_tokens=200, do_sample=True, top_p=0.95)
print(bytes(int(i) for i in out[0]).decode("utf-8", errors="replace"))

Weight files

file precision what it is
model.safetensors bf16 the final weights (step 245,674); the default, and the core every arm was trained on
model.fp32.safetensors fp32 the trainer's master weights at the same step; load with variant="fp32"
arms/<group>/<arm>.safetensors fp32 one file per arm of a group, as trained
arms/solo/<arm>.safetensors fp32 one file per solo arm, as trained
arms/index.json the arm table as data

Training ran under bf16 autocast over fp32 master weights. The bf16 file is those masters rounded to bf16, tensor for tensor. Each file's header names its precision and step. The arms were fitted on the bf16 file, so that is the core their measurements belong to.

Training

64.4B bytes on two RTX 5090s (data-parallel, 262,144 bytes per step): wikitext warmup (0.3B) β†’ fineweb-edu (20.9B) β†’ a nine-stage early-life curriculum, s0–s8 (35.2B: narrative, perspective, concepts, rule-chains, arithmetic, causal, try-fail, mixed, register) β†’ a two-phase anneal (4.0B of distribution shift without the chat frame, then 4.0B with it). Muon + pure Adam split, flat LR, bf16 autocast, zero loss spikes and the gradient clip never reached across the entire run. Every boundary shipped a report (held-out bpb, the toggle ledger, the chat frame, a probe suite, the rank profile): the reports, every checkpoint, the resume states, the TensorBoard runs and the pinned training code live in alephllm-mini-beatrix-training under mini-beatrix-3/.

Final validation: 0.9861 bpb on the fineweb-edu holdout. The web pretraining phase closed at 0.9838; the curriculum's stage diets moved the web holdout as high as 1.15, and the anneal brought it back.

The arms. During curriculum stages 1–4 an arm was trained beside the trunk for each stage; those arms stayed nearly empty, because the trunk took each stage's text in before its arm could, and they were detached at step 148,000 (their files are in the training repo under mini-beatrix-3/arms/, bound to those earlier trunk states). The arms in this repository were fitted afterwards on the finished, frozen core: as always-on groups and as solo arms. The four-arm group in two steps (first the arms of stages 1 and 2, in a run that added the four stages' text one stage at a time with every attached arm training under one pure Adam, 800 steps per stage; then, with those two held fixed and on, a new arm for stage 3 and after it a new arm for stage 4, 800 steps each). The eight-arm group by continuing that run: with the four held fixed and on, a new arm for stage 5, then 6, then 7, then 8, each attached over every arm before it and trained for 4,000 steps in their presence. Every group arm was kept quiet on ordinary web text and on the stage text of every other arm in its group. Each solo arm was trained alone for 4,000 steps (two seeds), kept quiet on web text; and, as an experimental third group, with the eight held fixed and on, a ninth arm for image captions trained over them for 4,000 steps on the caption pack, kept quiet on web text and on the eight's stage text.

Lineage

Code: AbstractEyes/alephllm (this repo vendors the model files; alephllm 0.10.5) and AbstractEyes/amoe-lora (the arm library; arms.py here is a self-contained re-expression of its runtime). Siblings: mini-beatrix-2s (237M, 20 blocks, the previous rung), mini-beatrix-2.5s (the 2s core with its arms) and mini-beatrix-1 (112M, the 3-hub hybrid).

Downloads last month
467
Safetensors
Model size
0.4B params
Tensor type
BF16
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using AbstractPhil/mini-beatrix-3 1

Collection including AbstractPhil/mini-beatrix-3

Article mentioning AbstractPhil/mini-beatrix-3