Baseline µop tier: compile everything fast, optimize what is hot

Status: on toyvm, phase 0 measured and phase 1 steps 1-2 built (§6: the E1 arena with links, and the µop-only arm that runs every instruction through it). Nothing is in the main emulator; it ports there (07d/07e) only if the toyvm gates pass.

1. The idea

Today both VMs run x86 in two tiers that do not share anything:

x86 --decode--> threaded code (L1 / $next)   every block, cheap, runs until hot
                   |  hot (sampled in toyvm; K block entries in the main emu)
                   v
x86 --decode AGAIN + 12 passes--> µop program (E1 / $uop_run)   loops and traces only

The two costs of this shape:

The proposal is to make the µop engine the only execution model:

x86 --decode--> BASELINE µops (one linear pass, block-local)     every block
                   |  hot (per-program counter)
                   v
       REGION µops (today's optimizer: promote, flagfwd, flaglive, guards, licm, clock...)
threaded code stays only as the fallback for instructions the µop set cannot express

Tier-up rewrites programs that are already µops; there is no second x86 decode. The existing optimizer does not change. It just gets a new input source.

2. Why the engine can afford this

These are toyvm measurements (memory note project_toyvm_uop_tier, uop-tier-design §1):

engine native instructions per transition (Ion, arm64 / x86_64)
E1 loop engine (one function, br_table) 11 / 12
threaded tail-call dispatch 33 / 38

The ~25-instruction difference is function-entry bookkeeping: frame setup, signature check, stack-limit check, the interrupt check, and the tail-call epilogue. Branch prediction does not account for it; that has been measured three ways. So one E1 transition costs about a third of one threaded transition.

The catch is the number of µops per x86 instruction.

Breakeven against threaded code is roughly 33/11 ≈ 3 µops per x86 instruction. The naive lowering is above that line, so it would lose; the baseline has to land clearly below it. That is the central design constraint, and it is the first thing phase 0 measures.

3. The baseline lowering: one linear pass, block-local

Everything below is chosen to be one forward pass over one decoded block, with no CFG, no fixpoint and no whole-region analysis. Each rule is the local version of a pass the optimizer already has.

concern naive lowering today baseline
guest registers GETR/PUTR through the register file on every access toyvm: promote within the block (load on first use, write back at block exit). Main emu: nothing, because vregs 0–7 already are the register file
addressing separate add/shift temps the memory op's own [base + (idx << s) + disp] form, filled in at decode (local addrfold)
flags a REC per producer, a GETCC per consumer producer→consumer forwarding inside the block (cmp+jcc → bcc). At block exit, one rec of the last producer, only if a consumer outside the block could read it. With no liveness information, assume it can
memory full-semantics checked access (non-native in toyvm E1, so a BAIL) a per-access native check: window test, then the slow path through the translator. No hoisted guards, because there is no loop to hoist to
clock STEP 1 per instruction, a budget test per transfer one charge per block and one budget test at the block's exit, the same accounting the threaded path uses
unsupported instruction declines the region the block ends there and hands that one instruction to threaded code

Target: ≤ 2 µops per x86 instruction on the corpus's hot blocks. Phase 0 measures what these rules actually reach.

4. Block chaining keeps execution inside the engine

A baseline program is one basic block. Every exit is a transfer to another guest address. If each exit went back to the host dispatch loop, the design would recreate the boundary cost it exists to remove. So:

Invalidation stays exactly as it is for threaded code today: self-modifying code retires a block, and retiring a block also unlinks every target word that points at it. That needs a back-edge list per program, or an epoch check on linked exits (§7, question 3).

5. Tier-up

Where tier-up goes next: once calls stay inside the engine (§4), the region former can follow call/ret pairs through linked baseline programs. That gives the inlining the main emulator's §11.3 census asks for, without a separate call mechanism.

6. Plan on toyvm, with gates

toyvm is the right place to test this. It already has the pieces: the naive lowering (uop-ir.js), the optimizer with per-pass switches (uop-opt.js, ablationConfigs), the E1 engine (uop-wasm.js), the live installer (uop-live.js), a timing harness (uop-speed.js) and a 199-program corpus whose frames and dispatch counts serve as an exact oracle.

Phase 0: measure before building (no new tier)

  1. µops per x86 instruction for the baseline rules.
    • Build a baseline config in uop-opt.js from the block-local versions of promote, addrfold and flagfwd, with no liveness, guards, LICM or clock.
    • Report the µops per x86 instruction on every corpus head.
    • Gate: median ≤ 2.5.
  2. Speed per x86 instruction: baseline E1 against L1, on the same snapshots, using uop-speed.js.
    • Use snapshots at non-loop blocks too, not only hot heads; blk_mix-style working sets.
    • Gate: geomean ≤ 0.8× L1's ns per instruction, on both V8 and SpiderMonkey.
  3. Compile cost. Time the baseline lowering and encoding per x86 instruction against L1's compile.js per instruction on the same blocks.
    • Gate: ≤ 3× L1. Code that runs once pays this, so it bounds the boot and load regression.
  4. Size. Compare the encoded bytes per x86 instruction for baseline µops and L1 threaded words. This decides whether the arena holds a whole working set (§7, question 2).

If gate 1 or 2 fails, stop and write down why. The baseline lowering is then the thing to fix, not the plan.

Phase 1: baseline tier live (--uop-baseline)

Phase 2: tier-up (--uop-baseline --uop)

Phase 3: main emulator

Port the design only if phases 1 and 2 pass.

Phase 0 results (2026-09-29, quiet box, not yet committed)

Density (item 1): tools/toyvm/uop-baseline-census.js, 193 programs with samples. Every sampled L1 block is lowered as a straight line (discover {straight}) and its path µops counted, weighted by samples (guest steps):

naive baseline (OPT.BASELINE) all passes baseline, resident regs all, resident
median µops / x86 insn 7.62 4.59 3.59 3.66 2.62

L1 takes a median 0.97 dispatches per instruction. Gate 1 (≤ 2.5) fails for every variant. "Resident" drops the per-block getr/putr reloads, the model where guest registers live in E1's value file permanently; it is the closest (2.62). On a short block (median 5–6 instructions) the promote model's entry/exit register traffic is most of the cost.

Coverage is not the obstacle it was assumed to be: of the sampled steps, 49% sit in L1 spin folds (nearly free in time already), 37% are whole straight lines in the µop set, 12% run a prefix and stop at an unsupported instruction (top stops: movs, leave, out, a self-patching cs: store, call far), and only 2% cannot start.

Speed (item 2): uop-shell-bench.js --configs=all,baseline, top live head of each of the 67 installing programs, node and SpiderMonkey. A loop region compiled with only the baseline passes, i.e. block-local code quality with perfect in-region chaining and a per-header clock, so an upper bound on a real baseline tier:

So far, on toyvm: density fails the proxy gate, but time on the clean set passes, because an E1 µop is much cheaper than an L1 dispatch.

E1 bails removed (2026-09-29). There were two native gaps, both now filled in uop-wasm.js:

In test-toyvm-uop-live the bails went from FLAGS 4512, MULDIV 8934, SEGLOADS 178 and DSHIFT 73 to 0. Two new cases were added to test-toyvm-uop (every config, via the reference interpreter) and test-toyvm-uop-live (E1 against pure L1: registers, RAM and dispatch count):

Rerun on the box (bt5, quiet, --configs=all,all,baseline,allRP; the second all is the run's own null band), 34 clean programs, zero bail rows anywhere:

node SM
null band (all#2/all, geomean) x1.000 x0.998
all over L1 x1.274 x1.164
baseline over L1 x1.153 x1.052
baseline time / all x1.105 x1.107
allRP time / all x0.999 x0.988

NM2 baseline went x0.03 → x2.10 (node) and RUNME2ND x0.02 → x2.17. The baseline passes keep most of the loop tier's win but not 70–80%: they are 10% slower than all in time. Gate 2 (≤ 0.8× L1's ns, i.e. ≥ x1.25) fails on this set, which is still dominated by heads that barely enter E1 (21 of 34 within ±5% of L1).

Capture ladder (bt6). In bt5, 27 of the 67 corpus heads "never reached" the loop head at a flat 20M-step capture. uop-speed now tries a ladder of budgets around it (B, B/2, B/4, B/8, 2B, 4B), each from a fresh run, and stands at the head from the first budget it recurs after. Only 2 heads still miss (B-STEEL, BAZIRRE), and the set grows to 48 programs (node) / 47 (SM), still with zero bail rows. The bt5 conclusions hold: all x1.302 / x1.193 over L1, baseline x1.138 / x1.027, allRP x1.003 / x0.990 of all, null band x0.996 / x0.997.

The ladder also exposed a toyvm cache bug, not a tier one. A new CodeCache restarts its arena but kept the previous session's jump table and return stack. Capture opens a second session on a warm VM, which then dispatched into overwritten arena words: "table index is out of bounds" on DINO, ANSWER, ALCHMSB, CONTAGIO, DENTROCF and ACME-VIC. The fix is in dos-loop.js: the constructor now clears both. Programs whose window ends on SMC or an int after a few thousand steps (ANSWER, DOPE, POLLY) are too short to time and should be read as no data.

What resident is worth to a chained tier (bt7). Phase 1 would chain per-block programs, which is exactly where the promoted model pays a reload at every block entry and a flush at every exit. The baselineBF proxy puts those loads and stores on every fast body block of an otherwise-baselineRF region: the same µops plus the block-boundary traffic.

node (48) SM (47)
null band (all#2/all, geo; p90) x1.003; 2.1% x0.998; 2.4%
baselineRF over L1 x1.138 x1.024
baselineBF over L1 x1.062 x0.956
BF time / RF (geo) x1.071 x1.071
BF time / RF, p90 x1.324 x1.263
BF > 5% slower 17 of 48 15 of 47

Per-block reload and flush cost 7% of E1 time on both engines. That is three times the null band's p90, and on SM it pushes the baseline below L1. It is also a lower bound: mergeStraight has already fused straight-line runs, so the proxy has fewer boundaries than a per-L1-block tier would. Resident is the register model for phase 1.

Resident guest registers (2026-09-29, measured)

Idea: make guest registers the first vregs, so a block never reloads or writes them back. L1's register block was reordered (isa.js: selectors first, then the 8 registers and 6 segment bases as one run of 14 dwords). A program built with resident (OPT.ablationConfigs allR/baselineR) puts its vreg file at REGFILE_BASE (uop-wasm.js VFILE), so vregs 0..13 are the registers, and finalize drops every RELOAD and FLUSH. A register that mergesink keeps narrow must keep its real upper bits now, so every write to it is tagged dw. The reference interpreter merges it, and E1 lowers it to a store-narrow op variant (add.h, mov.hi = AH). A partial merge into the register itself becomes one mov.h/mov.b/mov.hi.

Density: it works. Census, 193 programs, median E1 ops per x86 instruction (the census now counts lowered engine ops, not only optimizer µops, and those differ: flag materialization in exit stubs expands):

baseline baselineR all allR
E1 ops / insn 6.00 4.79 4.09 2.93

Time: narrow stores lose, full-width merges break even. uop-shell-bench, same 67 heads, clean rows (no bails), time relative to all (<1 = faster):

node (n=35) SM (n=34)
allR (narrow i32.store16/8) x1.084 x1.077
allRF (32-bit load-merge-store) x0.998 x1.005

The narrow store is the problem. The register's next read is a 32-bit load of the same slot, and a wider load cannot forward from a narrower in-flight store: it is the partial-register stall moved into memory. DIZZY_FI and anarchy go from x2.9 to x1.5 over L1 (node), and from x1.46 to x0.84 (SM). With the merge done as a full-width RMW (resident: 'full', allRF) those recover completely, and the geomean is flat.

Flat, not faster, because these heads are loops: one entry and one exit per thousands of iterations, so the reload and flush it removes were never on the hot path. Inside the loop, the promoted model keeps a narrow register in a vreg with garbage upper bits and pays nothing. The resident model pays a load-merge-store per write (RUNME2ND x5.14 → x3.99). Where resident should pay is the case the census prices: short blocks entered and left constantly, i.e. a baseline tier with block chaining. Loops can't show that, so it waits for phase 1. First verdict (superseded below): keep resident: 'full' as the baseline tier's register model, never use narrow stores into a slot that is read full-width, and keep the loop tier on the promoted model.

One register model: resident, plus per-region promotion of narrow writes (resident: 'promote', allRP). Two register models aren't needed. The only thing promotion still buys a loop is full-width writes to a register mergesink keeps narrow. So finalize keeps the register file as every register's home and renames only those registers (narrow and written in the region) to a temp: t = mov r at FASTENTER, merge16/merge8l r, t (one mov.hf/mov.bf) at each exit and deopt. Everything else stays resident. The baseline tier needs none of this: a block compiles each operand to its slot.

Rerun on a quiet box (load 0.0; the allR/allRF table above ran while another agent held the box, so treat it as indicative). Same 67 heads, all clean, time relative to all:

node (n=35) SM (n=34)
allRF x0.999 x1.007
allRP x0.992 x1.009

This run has no all-vs-all repeat, so there is no measured null band. Per-program swings of ±10% show up in both directions for all three configs (ASYLUM, NM2, ANARCHY), so read them as noise. The narrow-write heads return to all under allRP: on SM RUNME2ND is x2.603 / x2.266 / x2.603 (all / allRF / allRP) and anarchy x1.35 / x1.46 / x1.353; on node RUNME2ND is x3.80 / x4.01 / x3.86. Verdict: allRP matches the promoted model on loops, and it is a single register model that a baseline tier can share. Tests: test-toyvm-uop (2904 differential runs) and test-toyvm-uop-live under allRP.

Phase 1, step 1: the arena and links (2026-09-29, 4738fc82)

The mechanism §4 needs now exists (uop-wasm.js E1Arena, uop-live.js chain, CLI --uop-chain):

It links nothing, and that is the finding. Six demos were run for 60M dispatches with --uop-chain and an exit census (exitsAt):

program early exits / all exits where the early ones go
RUNME2ND 87 / 2707 a 1-instruction line
DREAM 100 / 1683 1-instruction lines
NM2 64 / 2578 a 6-instruction line
AUTUMN 1 / 1563 a 1-instruction line

96-99.9% of loop-program exits are the end of the slice (budget spent). They are not a transfer a chain could continue. Chaining between loop programs has nothing to gain, for a structural reason: discovery explores everything reachable from a head. So a loop program's static exits land only where exploration was cut, at an unsupported instruction or the node cap, and never at another program's head.

The time the loop tier misses (DREAM 44%, AUTUMN 45%) is in heads it declines: functions ending in ret, straight lines, in al,dx. The reason is that a non-loop head costs a JS round trip per entry. That is §4's argument, now measured from the other side: the links pay off only once there are per-block programs to link to. The next steps are:

  1. Lazy link fill: a miss returns to the arena loop, not the session, which builds the target and patches the word.
  2. An L1 fallback that hands back after the unsupported instruction. Phase 0's 12% "cut" share (movs, leave, out, far call) goes through it on every execution.

Phase 1, step 2: the µop-only arm (2026-09-29, 4738fc82, fe26bbfb, 00306e68)

Both next steps now exist in one place, tools/toyvm/uop-only.js (run-dos uopOnly, CLI --uop-only, arm-bench arm only). It is the far end of this design: every instruction the guest runs goes through a µop program on E1. It is a fourth arm beside l1, uop (L1 + the loop tier) and jit, not a replacement for any of them. L1 is still the oracle: same guest, same dispatch clock, same frame.

Results (100M dispatches, 12 demos, load ~11 so the times are rough; checksums exact):

arm CPU vs l1, geomean p10 p50 p90
uop x1.40 x1.12 x1.39 x1.62
jit x1.40 x1.09 x1.27 x1.99
only x5.63 x3.63 x6.37 x8.34

11 of 12 programs are exact in every arm. The µop share is 86-100% per program. Most of the only gap is build time, not execution: BRW spent 148s building in a 110s run, ACCIDENT 13.7s of 17.4s. The engine is fast enough once built; building everything is what costs.

Bugs it found while becoming exact:

  1. A handback for an unbuilt site re-cut the slice. Each handback re-derives the Sound Blaster block-end cut, so an extra one with budget left moves every later IRQ date (BRW). Fix: build inline.
  2. The continue rule could run past an endSlice cut made by a port write. Fix: the drive loop checks machine.sliceCut.
  3. A program over an L1 spin loop left the loop after as many iterations as the port took to change. L1 charges the rest of the slice in whole iterations, so the handback landed on a different dispatch (BRW's SB poll, 3 dispatches early). Fix: where L1 compiles a spin op (lone self-branch, 3DAh poll, general port poll), the site is that one L1 block with every other head a handback. This is the spin answer for this arm: exact and at L1 speed. A µop spin terminator would buy nothing here, so none was built.
  4. A region that fails with too many vregs (> 2047) paid a ~1s failed build on every SMC rebuild. Fix: lineCap remembers the line length that worked (BRW 30M builds: 34.6s to 17.6s).
  5. resident: true computes wrong values in this arm. ACCIDENT exits to DOS at 1.2M dispatches, on E1 and on the reference alike, so the program is wrong, not its lowering. The arm now refuses anything but 'promote'. baselineRP (baseline passes, promote) is exact and is arm only-bl. (Fixed in step 7 below: narrow reads now see the slot's upper bits masked off; arms only-R and only-RF.)

BRW still disagrees, and it is a retiming, not a wrong computation. L1 marks code bits over everything compileProgram compiled, i.e. the whole reachable program including code never executed. The µop arm marks only what its sites cover. BRW's self-patching extender (110:18f..1b1) stores into bytes L1 has compiled and this arm has not, so L1 breaks the slice there and charges a dispatch the µop arm does not. The same computation lands one dispatch apart, and IRQ dates follow.

Where build time goes (OPT.build(reg, { timing }), BRW 15M, 788 builds):

pass share
constprop 34%
forwardFlags 16%
sinkDeoptDefs 14%
mergeStraight 10%
forwardMemory 6%

No pass is superlinear. A 12-instruction line replays in 18-30ms. The 150-500ms outliers were GC and box load, and a replay of the same build spiked to 193ms once. It is a fixed per-program price paid for every site, and most of BRW's builds are distinct sites, not rebuilds:

So the lever for this arm is building less, not optimizing faster: tier cold lines, or share programs across sites.

Fallback coverage (00306e68): RCL/RCR (register destination), CLC/STC/CMC and LEAVE are now µops, exact to L1's handlers. A new differential shape (rotcarry) covers every width and count form. The flag helpers go through callh, which can now read a flag: the carry is materialized ahead of it, and E1 treats a CF-reading helper as an effect. At 30M, CMA_SHRT's clc/stc (1.8M per 100M before), DTM2's leave and rcl, CYCLE's rcl/rcr (620K) and B-STEEL's rcr are gone, and all four stay exact.

What is left, and why it was not done yet:

Phase 1, step 3: a cheap cold tier and tier-up (2026-09-29, 21f4c54e + this)

A cheap build is a pass set, not a new lowering. The earlier worry was that OPT.build's fixed structure (fast and slow halves, cfg, mergeStraight, finalize) was the floor, since a no-pass build cannot be resident. It isn't the floor. With promote alone, the arena accepts every program. A replay of ACCIDENT's 2001 captured builds (same inputs, configs alternated) prices them:

config ms per build static µops per insn
allRP 2.3-2.9 15.5
baselineRP 3.0 19.3
promoteLiveRP 0.73 17.7
promoteRP 0.55-0.66 18.5

The cold tier is promoteLiveRP (promote + flaglive), not promoteRP. Without flaglive, a program materializes flags L1 leaves stale because nothing reads them. PMENTRY's last xor dx,cx leaves DX=0, so ZF=1 by the architecture, but L1 and allRP both skip the dead record. That is correct, and it still differs from the oracle.

Planar VGA in place (21f4c54e). A full-checked access that hit planar VGA used to hand its whole block to the JS reference interpreter. That was 1.1M hand-backs on DREAM, and runRef + enter made up 38% of the run. ldfv/stfv now call L1's own uop_vga_rd8/wr8, imported into E1. DREAM went from x14.4 to x4.4 of L1.

The fast half still deopts at every planar access. On DREAM, 65% of the hot program's blocks run in the slow half, so optimizing that loop buys nothing. That is the next VGA lever: predict planar from the segment base at build time, and make the op a barrier to forwarding.

Tier-up (tier: { passes, after }, arms only-tK):

Results, cpu against l1, exact everywhere:

dispatches only (allRP) only-min only-t1k only-t10k
10M, 12 programs x12.3 x6.6 x8.9
100M, DTM2 / B-STEEL / CMA_SHRT x5.0 x3.7 x3.7 x3.5

allRP code does run faster. In B-STEEL's profile, its E1 time is 0.09s against promoteLiveRP's 0.19s. But E1 execution is a sliver of a 10M run, and a hot site is a big region, so its allRP rebuild costs ~12ms. The garbage from those builds is also billed to cpuSecs through the concurrent GC threads, which buildSecs never sees. At 10M, tier-up cannot amortize. At 100M it breaks even (CMA_SHRT gains, x2.85 to x2.38).

So the gap to L1 at long runs is no longer builds. DTM2 on only-min spends 2.9s outside builds against L1's 0.74s total, with 608K one-instruction L1 fallbacks and every program entry a JS round trip. What's next:

Phase 1, step 4: far call and return as µops, linked across segments (2026-09-29)

In real and V86 mode, call ptr16:16, retf and retf imm16 now decode (uop-x86.js kinds callf/retf), with L1's exact semantics. call_far reads its immediates from the code again when it runs, as L1's handler does: FARCALL patches its own selector from the straight line that contains the call. The 32-bit and protected-mode forms stay in L1.

Every run agreed with l1 on dispatches and frame:

run only (allRP) only-min only-t1k only-t10k
10M, 12 programs x11.7 (was x12.3) x6.4 (x6.6) x8.1 (x8.9)
100M, DTM2 x6.9 x4.7 x4.8

DTM2's fallback entries at 100M fell from ~600K to 8,278. The top fallback is now spin, 40.9M of its 100M steps: its retrace wait runs in L1, and L1 folds it cheaply. The remaining fallbacks are pushf/popf, cli/sti, and movs/stos with rep.

Phase 1, step 5: planar VGA predicted at build time (2026-09-29)

Before this step, any fast-half access that hit planar VGA deoptimized. Its whole loop then ran in the slow half, whose ldf/stf bail to the reference interpreter. DREAM ran 65% of its hot blocks that way.

makeFast now predicts planar accesses:

Result: DREAM's slow half went from 65% of its hot blocks to 0. All 12 programs are still exact at 10M (only x11.9, only-min x6.4, only-t1k x8.3, at load 20-30).

A bug the new test found (VGAPLANAR, test-toyvm-uop-only.js). The bug was older than prediction. Both entry loops (enterOver and E1Arena.enter) read the VGA key once, when they were entered. Consider a program that enables planar mode with an out and then bails to the reference interpreter. When wasm resumed, it got the key from before the out. It then treated A000:xxxx as plain memory: a read-modify-write lost every VGA effect after its first iteration. The key is now read on every entry into wasm.

Phase 1, step 6: the tier-up K sweep (2026-09-29, 6fe267f7)

Arms only-tK for K = 64, 1k, 10k and 100k, against only (allRP) and only-min (promoteLiveRP). Every run was exact. The box was at load 6-14, so treat the timings as rough.

cpu x vs l1 (lower is better):

run only only-min t64 t1k t10k t100k
10M, 6 programs, geomean x14.7 x7.5 x11.3 x10.1 x8.6 x7.5
100M, B-STEEL x6.5 x4.3 x5.4 x4.5 x4.0
100M, CMA_SHRT x2.6 x2.1 x2.2 x3.0 (noisy) x2.8
100M, DTM2 x8.2 x5.2 x5.8 x5.4 x4.9

Build cost decides the ranking, not code quality. allRP code is much faster once it is built. Net of build time, on B-STEEL it runs 0.09s against promoteLiveRP's 0.96s; on DTM2, 0.84s against 1.67s. But a single allRP build costs as much as L1's whole 100M run (0.4-2s), which is not enough time to pay it back.

So, for this arm, the next lever after tiering is build cost and µops per instruction, not a better K.

Phase 1, step 7: resident: true fixed; arms only-R and only-RF (2026-09-30)

The bug. It was in the program, not the lowering. promote lets mergesink keep a register narrow: 8 bits if a program touches only cl, 16 if it touches only cx. The passes then assume the promoted model, in which the register's vreg holds a clean, zero-extended value.

How it was found. Three steps, with scratch probes:

  1. Bisect program start ips: the ip-0 programs.
  2. Keep masks for one consumer kind at a time: only add needed them.
  3. Bisect add sites: one site, head 0xEDC.

The fix. In finalize, for resident true and 'full' only, each narrow register read is of a zero-extended copy (andi). The copy is made once per block and made again after the register is written. Writes were already narrow stores (dw). On the wraps loop this is +4 µops per iteration (allR 43 against allRP's 39).

Tests. test-toyvm-uop-only.js NARROWHI is ch=1 set by an earlier program, then a cl-only loop of add dl,cl / adc bh,0. Without the masks, allR gives bx=2C00h; L1 gives 0. NARROWHI, DSHIFT, SHIFTS and FLAGS now also run on allR and allRF. uop-only accepts any resident model.

Results. arm-bench arms only-R (allR) and only-RF (allRF): all 12 programs are exact at 10M, ACCIDENT included. The three resident models are within noise of each other at load 12-19:

10M geomean only (allRP) only-R only-RF
programs 1-6 x14.5 x14.0 x13.8
programs 7-12 x10.3 x9.9 x10.1

The 10M runs are build-bound (step 6), so this says nothing about the three models' execution speed. That question stays with the resident-registers section above.

Phase 1, step 8: the naive lowering as the cold tier (2026-09-30)

Why a first build cost more than L1's whole compile. Measured with a phase probe and an optimizer replay over captured regions:

cost per x86 instruction
L1 compileProgram (decode, emit words) ~3 µs
µop optimizer, promoteRP (cheapest set) ~54 µs
µop optimizer, baselineRP ~405 µs
µop optimizer, allRP ~680 µs

That is before discover, IR.lower, lowerProgram and encode, and spinBlock's own L1 compile. The optimizer copies the region into a fast half, then walks every op 10-20 times: constprop alone runs up to five times, each up to 12 rounds. After two fixes (54555207, output-identical: mergeStraight was quadratic, and each op was cloned through JSON) no function stands out. GC is 12-14%, and the rest is spread thin.

The cold tier needs no passes. The naive lowering (IR.lower) is already valid in any machine state: every register is a GETR/PUTR of L1's register file, every flag a REC into L1's lazy-flag globals, every access the full accessor, and the budget is tested at every transfer. It names no guest vreg, so resident changes nothing in it but where its temporaries live. Pass set naiveR (uop-opt.js naiveResident) installs it as an arena program:

Arms. only-naive (naiveR everywhere) and only-nK (naiveR cold, rebuilt on allRP after K header counts). Tests: test-toyvm-uop runs naiveR as its 28th configuration (5040 differential runs agree), and test-toyvm-uop-only runs every program on naive and on naive-tier (tier after 4).

Results. All 12 programs are exact at 10M in every arm, load 13-20:

10M geomean vs l1 only (allRP) only-min only-naive only-n100k
programs 1-6 x15.6 x8.1 x6.9 x6.7
programs 7-12 x9.3 x5.1 x4.3 x4.8

A naive build is about 5-10x cheaper than the old one. At this run length the unoptimized program beats every optimized one outright: the optimizer never earns its build back. What is left, x4-7 of L1, is no longer build time. DTM2 builds for 0.07s of a 0.63s run, against L1's 0.08s total, so the next lever is the arm's execution and handback overhead.

Phase 1, step 9: no reference-interpreter blocks, and what Ion makes of E1

Bails. Profiling only-naive showed arena blocks the lowering refuses, all for one reason: full memory after an effect. A full-checked load or store can hand back, which reruns the block from its start, so it must not follow a side effect in the same block. Naive blocks are one instruction but still hit it (push stores after a register write; movsw does both). Those blocks ran on runRef, the JS reference interpreter: DTM2 93,144 in 10M, DEMO5 138,706, ACCIDENT 31,997. splitFullAfterEffect (uop-opt.js) cuts a body block at each such access, and br joins the two halves. Non-native blocks are now zero on every probed program. The remaining hand-backs are runtime ones: 0-135 per 10M.

10M geomean vs l1, all 12 exact, load 16-28 only-naive only-n100k only-min
after the split x4.61 x5.00 x5.97

That is down from about x5.4 for only-naive (step 8's two halves combined). Spin fallback is still large: DREAM hands L1 2.9M spin dispatches of 10M, RUNDEMO 6.2M, CYCLE 0.5M. Those run at L1 speed and cost this arm nothing. So the x4.6 is the arena's own per-µop cost.

What SpiderMonkey Ion makes of $run. Measured with tools/wasm-native.js on the E1 engine, with arms named through Ion's jump table. $run is one function: 785 br_table arms, 128 KB of arm64 for 68 KB of wasm. The dispatch head costs 10 instructions and 4 memory operations per µop:

ldr  w16,[x23,#56] ; cbnz        interrupt check, every µop
ldr  w0,[x20,#60]                $pc reloaded from a stack slot
ldr  w2,[x21,x0]                 opcode
cmp  w1,#0x311 ; b.cs            br_table bounds check
ldr  x3,=table ; ldr x16,[x3,x1,lsl#3] ; br x16

Every arm ends with ldr w6,[x20,#36] (the $shm param, reloaded), str w0,[x20,#60] ($pc spilled) and b head. The loop-carried locals live in memory, not registers: $pc makes a store-to-load round trip through the stack on every µop, on the critical path. Each operand costs two instructions (add x16,x0,#k; ldr w,[x21,x16]), because arm64 has no base+index+offset mode.

arm (naive-hot) insns loads/stores
mov 10 6
getr32 / putr32 / getr16 / putr16 13 8
add / sub / and / xor 14 8
step 10 5
link 34 12
exit 39 14
ldf16a 41 14
stf16a 52 15
all 785 arms: median / p90 / max 28 / 72 / 2304

So a register copy µop costs 20 native instructions with the head, and 10 memory operations, of which 2 are the work. The levers, in order:

  1. Fewer µops per x86 instruction. This is the naive tier's own lever: 84 µops per wraps iteration against allRP's 39. Fused naive shapes (getr+op+putr as one µop) remove heads without any pass.
  2. Keep $pc in a register. This is an Ion register-allocation outcome, not something the WAT can request directly. The candidate test is a smaller $run, with cold arms moved out to callees: if $pc stays in a register in a 50-arm engine, arm count is the cause.
  3. The interrupt check. Ion inserts it at the loop head. A loop that exits through $budget does not need it per µop, but wasm has no way to say so.

Phase 1, step 10: the $pc spill is Ion's, and it comes from the host calls

Cause. Rebuilding $run from a subset of arms (the first N in EOPS order) shows where the spill starts:

No single arm added to the first 90 triggers it. Arms 92-93 are ldfv8a and stfv8a, the first two arms that call a host import ($vga_rd8/$vga_wr8). One call site is fine. From two on, Ion stops splitting the loop-carried ranges around the calls and gives them stack slots everywhere. E1 has 175 call sites in 79 arms.

Fix: callSafe (uop-wasm.js). In the loop engine, every host call is wrapped so the loop-carried locals are stored to CALLSAFE before it and reloaded after it: $pc $steps $F plus the machine params and the slot registers. None of them is live across a call any more.

Ion, full 785-arm engine:

V8 is different. tools/wasm-native.js --engine=v8 reads TurboFan's code out of d8's --perf-prof jitdump. TurboFan keeps $pc in w0 in BOTH builds, and the hot naive arms are already tight without the change: getr16 is 8 instructions with 4 memory ops, step 6, no stack use. callSafe still reduces V8's stack use elsewhere: arms touching sp go from 440 to 81, and stack ops from 3,019 to 1,597, mostly in cold shift and move arms.

Timing. Nothing on this box resolves it.

callSafe stays on by default: it is correct everywhere (test-toyvm-uop, -uop-only and -uop-live all pass), and structurally it removes memory ops on both engines. TOYVM_CALLSAFE=0 and the @spill arms keep the old engine for the A/B on a quiet box.

Phase 1, step 11: fewer naive µops (fall-through layout, fused flags)

Census. TOYVM_E1HIST=1 makes the loop engine call a host.hist import before every dispatch. e1Hist() returns the op and op-pair counts. It is off by default and costs nothing then. Run over DTM2, DHADREN, B-STEEL and CMA_SHRT at 3M dispatches each with naiveR, the naive program executed 59.5M µops, about 5.7 per x86 instruction:

Fall-through layout (fallThrough in uop-wasm.js). Blocks are re-laid greedily along jmp chains from the entry. A trailing jmp to the next block is dropped. Arena and single programs use it; straightWat does not. The census fell to 53.2M (-10.6%).

Fused flag µops. A full six-flag rec from add/sub/add32/sub32/logic is one µop instead of seven:

Checks. test-toyvm-uop (all 28 configs agree), -uop-only and -uop-live pass. arm-bench (12 programs, 10M dispatches, load ~8, every program exact):

arm cpu vs l1 slice time vs l1
only-naive, this step x4.62 x3.78
only-naive, step 10 (load 20-60) x4.61-4.84 x5.1-5.9

The slice-time drop matches the µop cut, but the earlier runs were at a far higher load, so it is not a clean A/B. The cpu number includes the builds (up to 0.4s per program) and the L1 fallback, which this step does not touch.

Where the naive µops go now:

Phase 1, step 12: step fused into its predecessor

X_s is X followed by $steps -= i, for X in putr8/16/32, getcc*, the fused flag writers, wflags and stfv*. fuseSteps rewrites X, step into X_s in every lowered block. Not pout/pin, which read $steps. A bail inside X skips the step either way. TOYVM_STEPFUSE=0, or an arm-bench ARM@nofuse arm, turns it off.

BRW.EXE disagrees at 30M in every uop-only arm. It is older than steps 11-12 (the step-10 engine does it too), and uop and jit agree with l1.

Phase 1, step 13: BRW.EXE fixed, three clock leaks (2026-09-30)

There were three causes, and all three were handbacks or dispatches that depend on how the code is cut rather than on the guest.

  1. L1 billed its own volatile cut.
    • compile.js ends a straight line with end where it runs into bytes the host has learned are volatile, or where it runs out of words.
    • end is a dispatch, so it cost a step. So 110:18f..1b1 cost 6 once 110:1b1 was volatile, against 5 before, and 5 in uop-only.
    • end_cut is end that gives the step back. It is the same argument as jmp_syn. This is an L1 clock change, and all 28 toyvm tests still pass.
  2. A µop line exit handed back mid-block.
    • A naive program whose straight line stops at a fallback instruction (cli here) is in the middle of an L1 block. L1 tests no budget there.
    • Its exit is now why: 'line', lowered to exitl/linkl. linkl chains without the budget test and stops only on smc.
    • The driver carries on past a spent budget, as it already did for a fallback's end.
  3. An iret fallback always handed back.
    • L1's iret hands back only when it owes the host something: TF, or an IRQ held for IF. Otherwise it resolves the return through its block cache.
    • uop-only compiles nothing L1 can find, so the lookup always missed.
    • That extra handback in BRW's SB handler at 7.44M re-cut the slice. The SB transfer's first render landed 42 dispatches early (7,440,757 against 7,440,799). 17M dispatches later, in al,2 (the DMA position) read f1 where L1 read f4.
    • The drive loop now stays on an iret exit unless the window is owed.

How it was found:

Result: arm-bench, 12 programs at 30M, l1 against only-naive: 12 of 12 clean. Geomean x3.44 cpu and x2.97 slice, at load ~3.8.

Phase 1, step 14: fewer naive µops, cheaper builds, every arm measured (2026-09-30)

Dynamic µop census (TOYVM_E1HIST=1, BRW, CONTAGIO, DEMO5 and ACCIDENT at 10M each), and what each change took out:

change µops run where
step 13 140.2M
full-width getr/gets forwarded (17.8% of µops were these copies) 116.5M forwardFullGets, uop-opt.js
narrow masks fused into the op making their operand 107.4M fuseMasks, uop-wasm.js

Every arm, 12 programs at 30M, all clean. CPU geomean against l1, the two halves of the corpus run separately at load ~2.5-2.8:

arm cpu (1-6) cpu (7-12) slice (1-6) slice (7-12)
uop (l1 + µop loop tier) x2.07 x1.97 x1.00 x0.95
jit (l1 + region JIT) x2.66 x2.61 x1.06 x0.94
only (allRP) x10.0 x5.98 x9.73 x4.49
only-naive x3.82 x2.87 x3.07 x2.22
only-n1k (naive, tier-up to allRP) x5.83 x3.97 x4.92 x2.76

How to read it:

7. Open questions, each with how phase 0 or phase 1 answers it

  1. Local flags without liveness. Materializing a rec at every block exit may cost more than it saves in branch-dense code. Measure how many exit recs are actually read. The alternative is an eager flags word, as toyvm E1 already models it (flagof/getf/wflags).
  2. Arena capacity. A baseline tier stores all executed code, not just hot loops. The main emulator's churn with trace heads (Diablo: 2,491 kills) shows what an undersized arena does. Phase 0 item 4 gives the bytes per instruction. The arena must hold the working set or evict cold baseline programs only (LRU by entry count), and optimized programs must never be evicted for baseline ones.
  3. Unlinking on invalidation. Choose between a back-edge list per program and an epoch word checked by each linked exit. The epoch costs one load per chained exit but needs no bookkeeping. Measure both in phase 1.
  4. Blocks that only ever run once. If compile cost per instruction exceeds gate 3, compile on the second execution instead of the first, and let L1 run the first. That is a K=2 threshold for baseline, the cheapest possible tiering.
  5. Correctness surface. The baseline lowering is new code that runs everything, unlike the optimizer, which runs only hot loops. The corpus dispatch-count oracle in phase 1 is the check. Any divergence stops the phase.

8. What this is not