A toyvm-style micro-op tier for the main emulator
Status: design + phase-0 prototype, 2026-09-26. Off by default; nothing ships from this document until phase 0 says the engine is worth building.
1. Why the block executor did not pay, measured in machine code
The H454/H458 block executor (src/07c-block-exec.wat, --block-exec) was
timed on the quiet box for the first time in September 2026:
| app | executor off → on | native share |
|---|---|---|
| Heroes III gameplay | 13.52 → 13.10 ms/batch (−3%), boot +9%, total user CPU +7% | 97.6% |
| MechWarrior 3 gameplay | 553.8 → 519.1 ms/batch (inside a ~9% null band) | 99.87% |
| MCM race | 2.345 → 2.462 s (+5%) | 95.6% |
| Diablo II to Act I load | 21.99 → 23.01 s (+4.6%) | — |
Coverage was not the problem: 95–99.9% of ops ran natively. So the question is
what each native op costs, and tools/wasm-native.js --func='$th_block_exec'
(SpiderMonkey Ion, arm64) answers it directly.
- The per-op prefix is ~100–120 instructions and 3–4 indirect jumps before
the kind
br_tableis even reached. It covers:- the handler-hist test;
- the six descriptor fields stored to frame slots;
- three 15-arm register
br_tables (d,TU_B_SRC0, a); - the lane/width/no-flags/EA decode;
- the SIB range test.
- The arm then does real work plus overhead. Most arms call
$set_flags_*(mov sp/bl/ restore around each call), and a writebackbr_tablefollows. - The "registers in wasm locals" are not in registers. The function is 11 KB
and 2789 instructions. Ion spilled
r0..r14to stack slots (ldr w16,[x20,#192..220]), so every register read is a jump-table jump to a stack load.
A threaded handler is ~150 Ion instructions including its dispatch. The executor removed that dispatch and paid the same amount back in its own prefix. That is why it measured flat.
toyvm's E1 engine (tools/toyvm/uop-wasm.js) pays 11 instructions per
transition on the same compiler. It has the same one-function br_table
shape, but:
- An operand is the vreg's byte address, so
V(k) = load(load(pc + 4k)): two loads, no jump. - No op makes a call.
- No op writes lazy-flag globals; the optimizer forwarded them.
- The function has a handful of locals, so nothing spills.
toyvm measured the full effect at ns per x86 instruction (V8/SM):
| engine | ns / x86 insn |
|---|---|
| L1 interpreter | 3.71 / 2.82 |
| E1 loop engine | 0.64 / 0.92 |
| straight wasm (not an engine, upper bound) | 0.23 / 0.20 |
Engine shape and the optimizer passes are the lever, not op coverage.
(Memories: project_block_exec_native_cost, project_toyvm_uop_tier.)
2. Constraints this design keeps
- No runtime wasm codegen. The engine is one fixed function compiled with the build. A program is data in linear memory, as with E1.
- Exactness. Every exit leaves the machine exactly as the threaded path
would. That means:
- guest registers;
- the five lazy-flag globals, when live;
- EIP;
- memory, including SMC retirement.
- Threads. Every guest thread is its own wasm instance, and a program bakes
its instance's
$reg_base, temps and windows, so each instance has its own arena: the main thread uses$UOP_ARENA, workertiduses slottid-1of$UOP_THREAD_ARENAS(15 x 256 KB, paid for by shrinking each$THREAD_CACHE_BASEpartition from 1 MB to 768 KB).init_threadpoints the instance at its slot with$uop_set_arena, which empties the map. The compiler's scratch$UOP_CSCRATCHstays shared and is taken with an atomic try-lock; a thread that finds it held skips the compile (answer 1, "busy") and retries at the head's next hot bump rather than marking it dead.set_uopandset_branch_clockare inINHERITED_WASM_GLOBALS, so both--threadsworkers and cooperative threads inherit the switch. - Fail-soft. Anything the lowering does not model ends the program with an exit to the threaded interpreter at that instruction. That is never a crash and never an approximation.
3. The engine (E1 shape)
One function, $uop_run(pc) -> exit code: a loop over one br_table on
i32.load(pc).
Vregs live in memory.
- Vregs 0–7 are the guest registers themselves (
$reg_base+ 4r). Entry and exit need no GETR/PUTR, and falling back to threaded code needs no spill. This is the main emulator's advantage over toyvm, whose guest registers are separate. - Vregs 8+ are temporaries in a per-thread scratch page.
- An operand word is the vreg's absolute byte address, pre-scaled at lowering time, so reading or writing an operand is one load of the operand word plus one load or store of the vreg.
Ops are a small RISC set, all 32-bit, each doing only its own work:
| family | ops |
|---|---|
| move | movi, mov |
| ALU | add sub and or xor + immediate forms, shli shri sari, mul, sx8 sx16 zx8 zx16, merge8l merge8h merge16, ext8h |
| branch | compare-and-branch bcc cc,a,b,w,target (flag-forwarded cmp/test + jcc) and bccf (test one materialized flag bit) |
| memory | ld{8,16,32}{s,u} and st{8,16,32}, addressed as window-relative base + (index << s) + disp |
| guard | guard w, base, lo, hi, rw, exit |
| flags | rec op,a,b,res,shift: writes the five lazy globals; emitted only on exit paths where flags are live |
| control | jmp, clock n, exit (one per loop header), exit eip |
Memory: the hoisted range guard (the user's rule). Guest virtual mappings do not change between iterations of a loop that makes no calls out of the program. So translation is proved once per window, not once per access:
guardat the preheader calls$g2w_affine_span(lo, hi - lo). That function already proves a whole span shares one guest→wasm delta (direct window, DIB window, or a contiguous sparse run).- For a written window,
guardalso tests the code-page bitmap ($code_write_is_code) for each page in the span. - On success it stores
deltaand[lo, hi)in window slotw. On failure it takesexit, before any instruction has had an effect.
- For a written window,
- An access computes the guest address
gaand then doesif (ga - lo) >u (span - size) → slow; else load/store at ga + delta. That is one subtract and one compare: no call, no bitmap, no page-cross test. - Choosing
[lo, hi). When the lowering can bound the address range from induction variables and a trip count, the guard covers it exactly. When it cannot, the guard covers the page the stream starts in. The user's heuristic is that a fast start means the stream is almost always fast throughout. - The slow path re-guards on the page of
ga. It calls$g2w_affine_spanon that page and refreshes slotw; if that fails too, it exits at the instruction. So a stream that walks off its first page costs one re-guard per page crossed, not per access.
What the window proves, and for how long:
- A window is valid only within one
$uop_runcall. Every entry re-establishes it, so a mapping change between batches is always seen. - Nothing inside a program can change a mapping, because calls to APIs are exits.
- A decode of new code into a written window by another guest thread mid-slice is the one gap. x86 requires software serialization for cross-modifying code (SDM 8.1.3), so a program relying on it without a lock is already undefined there. This is documented rather than guarded.
Flags. The lowering starts naive, like toyvm: every flag producer is a
rec. Two passes then work on it:
- flag forwarding turns
cmp/test/sub/dec + jccintobcc; - flag liveness over the loop deletes every
recwhose flags die before an exit.
Surviving recs sit on exit stubs, so the steady-state loop writes no flag
globals at all.
Clock. The threaded path charges $steps per block transfer. The program
keeps the same accounting exactly: clock n at each loop header subtracts the
iteration's block transfers, and exits with a precise EIP when the budget runs
out.
4. The lowering and the passes
Input is the region the H458 machinery already discovers, a loop nest of
decoded blocks. But the lowering reads the x86 instruction, not the handler
index, for the reason tools/toyvm/uop-x86.js gives: a handler index hides its
operands. The main emulator's decoder already decodes ModRM and SIB; the
lowering calls into that decode rather than copying it.
Passes are ported from tools/toyvm/uop-opt.js, first those that toyvm
measured as paying:
- promote: guest regs stay in their vregs; temps coalesce.
- constprop + addrfold:
[esi+4]becomes one address expression. - forwardFlags + flagLiveness:
bcc, deadrecremoval. - guard hoisting: the per-access window checks above.
- licm: loop-invariant loads (
[esp+0x40], a LUT base). - clock: one check per header.
- dce.
Later, if the census says so: inlined call/ret (toyvm's call contexts), and
load forwarding (toyvm measured it as nearly free: project_toyvm_uop_tier).
Where the optimizer runs: in WAT (src/07e-uop-compiler.wat, since
2026-09-27). It started as lib/uop-compiler.js behind a synchronous host
import, which was the quickest route to a first measurement, and was ported so
browser workers and every host get it without an RPC and without a host call
per hot head. $uop_try (07d) calls $uop_compile directly; the compiler
decodes the guest bytes, forms the loop, lowers it and writes the encoded
program into $UOP_ARENA. That is data, not code, so the no-codegen rule
holds. Its working memory is the $UOP_CSCRATCH region; nothing in it
outlives one compile.
The port was checked word for word against the JS reference before the JS was
deleted: every instruction address of every test-uop-compiler.js case under
both clocks, and every head MW3, MCM, D2 and Heroes III compiled on their
benchmark routes, produced identical programs (or the same decline reason).
Exports: uop_compile(eip) (what $uop_try calls; answers the program, 1
when another thread holds the scratch, or 0), uop_cstat(k) (compiled,
declined, instructions, uops, flushes, words), uop_decline_count(reason).
5. Phases and the numbers that gate them
Phase 0: price the engine, not the compiler. Add the $uop_run engine plus
a bench-loops.js arm that runs a hand-lowered program for existing shapes
(lut, blk_mix512, and a Heroes III RLE-run shape) next to the threaded and
block-executor arms, in one process with alternating reps. Hand-lowered means
written the way the optimizer is expected to emit it, so phase 0 is toyvm's
"what is the engine worth on this loop" before any compiler work.
- Gate: the uop arm is at least 2x faster than threaded per x86
instruction on the non-periodic shape (
blk_mix512). - Also:
wasm-native.jsshows ≤ 20 instructions per transition for a register op, with no spills.
Phase 1: JS lowering + passes on real regions (CLI). Install on the loops H458 would install on, and compare final state against threaded code.
- Correctness gate: equal registers, flags and memory hash on every
installed program across the
block-exec-sweep.jsapp list. - Speed gate: box A/B with user CPU plus
--slice-spliton Heroes III, MW3, MCM and D2, with interleaved reps and a null band.
Phase 2: move the lowering into WAT; browser and worker coverage.
Phase 3: calls. Inline short leaf calls, so regions stop ending at every
call. The H458 census shows shortChain and unsafeOp as the biggest
decline rows, and MW3's regions are almost all one block because of calls.
6. What phase 0 does NOT prove
- A periodic microbench loop predicts every branch.
blk_mix512exists to remove that bias, but the real apps' hot code (Heroes III RLE runs,COMPLEX/DIAMONDclasses) is branchier than any synthetic shape. - The lowering and passes have to reach the hand-lowered program's quality on real code. That is phase 1's question, and toyvm's coverage census says it is usually the harder one: 117/191 programs never installed.
7. Phase 0 results (2026-09-26)
Engine code (Ion, arm64, tools/wasm-native.js --func='$uop_run').
$uop_run is 2992 bytes of Ion code, against 11 KB for $th_block_exec.
- Dispatch costs about 11 instructions: interrupt check, op load, bounds check,
table load,
br. - A register ADD arm is 13 more instructions, including a
pcspill to the frame. - So a register op costs ~24 instructions and one indirect jump, against ~110 for 07c's prefix alone before its arm even starts.
Timing (node tools/uop-engine-bench.js --bytes=4m --reps=7, laptop at
load ~5.5, V8). Every rep of every arm had identical registers, CF/ZF/SF/OF,
EIP and destination hash; no guard failed.
| shape | threaded ns/insn | block-exec | uop ns/insn | uop vs threaded |
|---|---|---|---|---|
lut (byte LUT blit) |
12.50 | x0.96 | 3.28 | x3.81 |
ckey (colour-key diamond) |
12.71 | x0.93 | 3.60 | x3.53 |
h3shadow (Heroes III 0x471da6, 16-bit) |
11.97 | x0.96 | 3.19 | x3.75 |
How to read it:
- Block-exec is x0.93–0.96 on the same code. That matches the whole-app timings in section 1 and the native-code explanation of them.
- The uop engine is ~3.7x. That is in line with toyvm's E1 against its L1, as the shape argument predicts.
- The
foldsarm matched threaded (not in the table): LUT_RUN did not fire on this encoding, so it is not a fold-vs-engine comparison. - Re-guards fired once per 4 KB page per stream, because the windows were guarded at their start byte only (the user's "fast start → fast whole" rule), and cost nothing measurable.
Not yet shown:
- A non-periodic block working set (
blk_mix512-style). - Deopt stubs that restore exact flags (here they only exit and never fire).
- Any real lowering. The programs are hand-written at the quality section 4's
passes must reach: forwarded
cmp/test/dec+jcc, flags materialized only on the exit path, one CLOCK per iteration.
Phase 1 is the question now.
8. Games, and the call-free engine (2026-09-27)
Games (node tools/uop-game-ab.js, each game's own gameplay route, both
arms --branch-clock, laptop at load ~2-4). Frames first: every game's final
frame matches off vs uop, or differs by no more than off vs off.
| game | frames | whole-run user CPU | gameplay phase |
|---|---|---|---|
| Heroes III | identical | -7.7% | -41% |
| Diablo (shareware) | identical | -6.2% | -34% |
| StarCraft | 1.48%, null band 1.43-1.51% | ~-5.5% | -- |
| Warcraft III (menu, software GL) | identical | -3.4% | flat (GL-bound) |
Warcraft III gameplay (wc3g, headless GL) |
identical | -8.8% | -20% (4.4 -> 3.5s); map load -7% |
| Warcraft III gameplay, software GL | identical | -5.7 to -7.5% | 0% and +13% in two concurrent pairs (noise; GL-bound) |
| Heroes II, Diablo demo menu | identical | flat | flat |
| Diablo II | nondeterministic off vs off too | -3.5% | -- |
The call-free split. With any call inside the dispatch loop, Ion kept
$pc and $budget in stack slots: 110 [x20,#28] references in the old
$uop_run, a store after every op and a reload on the pc chain every op
depends on. The loop called $uop_reguard from 16 memory ops plus
$uop_window_set, $get_cf (x2) and $eval_cc. Now $uop_fast makes no
call at all and hands any op that needs one back to $uop_run (a window
miss, GUARD, SAVECF, GETCF, BCC). A missed op is re-run from scratch after
the re-guard; nothing in it changed before its window check. Result: pc lives
in w0 with zero stack references in the loop, and a register MOV is 7
instructions plus an 8-instruction dispatch.
What it bought is small, and the reason is the finding:
- Microbench: h3shadow -6%, lut -3%, ckey +1%.
- Heroes III, Diablo and StarCraft gameplay: flat against the pre-split
engine (
--ref-wasm=), with identical frames. - A
--cpu-profof the Heroes III uop arm shows why.$uop_fastis 6.8% of self time. The whole enter/re-guard path ($th_uop_enter, with$uop_run/$uop_reguard/$uop_window_setinlined into it) is 0.6%.$g2w_affine_spanis 0.0%. The x87 handlers ($th_fpu_mem_ro,$fpu_exec_mem,$th_fpu_reg,$fpu_exec_reg) are ~18%.
So windows that survive across entries (an epoch bumped on every mapping
change and code-page mark, shared across Worker instances) would buy under 1%
of Heroes III. That is not worth a stale-window SMC hole. The remaining lever
is coverage: the declines are no-backedge and head-unsupported, and
Heroes III's hot threaded time is FPU code the tier does not lower.
(Section 9 built those windows anyway. The epoch closes the stale-window hole within a thread. Building them also found a real SMC hole in store windows.)
Where that FPU time is, measured (2026-09-28). It is not on the thread the
tier works on. --handler-hist per guest thread over gameplay batches
4100-4300 (--handler-hist-thread=1,3,2; the main thread over 4100-5101):
| thread | x87 dispatches | share of that thread's dispatches |
|---|---|---|
| T0 (game) | 585 of 1.19G | 0.00% |
T1 (start 0x8414a0) |
66.6M of 141M | 47% |
| T2, T3 | 0 | 0% |
So the ~18% x87 CPU in the profile is all T1, and "x87 in uops" would be
aimed at the wrong thread. The existing semantic x87 folds (--x87-fusion:
pipeline4, island, affine) already reach it. On T1 they cut unfused x87
dispatches from 66.6M to 14.1M (plus 4.2M fused: H449 pipeline4 3.08M and
H451 1.13M). That is 79% of x87 dispatches absorbed, and T1's total
dispatches fall from 141M to 93M. With --uop as well, T1 reads 15.3M raw
and 4.6M fused, so the fold and the tier compose: the tier is on T0, the fold
on T1. The remainder is mostly $th_fpu_mem_ro (8.4M).
The fold is off for Heroes III: x87Fusion: true is set only for
ut2003_demo in lib/apps.js. Fewer dispatches is not a time win by itself
(on MCM the fold was flat until the x87 file moved to memory; see
project_x87_fusion_mcm), so the time A/B is the fold / uopfold arms of
tools/uop-game-ab.js.
Bench box, 2026-09-28 (x86_64 V8, node 20, Ryzen 9 9950X, 4 vCPU, idle,
serial arms, HEAD 84e79bb4). gameplay is the --slice-split main-thread
guest slice, so guest-thread (T1) work shows only in user CPU.
| game | frames | uop vs off: gameplay | uop vs off: user CPU | null band (user / gameplay) |
|---|---|---|---|---|
| Heroes III | identical | -52.9% | -5.2% | 1.8% / 3.3% |
| Diablo shareware | identical | -43.2% | -4.1% | 0.2% / 0% |
| Warcraft III, software GL (wc3g)* | identical | -13.0% | -6.0% | 0.3% / 0% |
| StarCraft | nondeterministic (off~off2 differ too) | -5.9% (0.1 s resolution) | -2.3% | 0.9% / 0% |
*wc3g needs the then-uncommitted Game.dll ordinal-import linking from the
working tree; on bare HEAD, Game.dll's DllMain stops at KERNEL32.#00001.
Heroes III x87 fold, same box, frames identical in every pair:
uopfold vs uop user CPU -4.3% (86.58/86.53 s vs 90.46/90.49 s, null
band 0.03%), gameplay slice 0.0%. fold vs off is -5.3% (null 0.9%). All
of the saving is guest-thread time, as the dispatch counts predicted.
Moorhuhn, same box. These ran on the working tree as of 2026-09-28 morning;
routes are mh1/mh2/mhw/mh3 in tools/uop-game-ab.js, and every final
frame was looked at and shows a live round:
| game | frames | uop vs off: gameplay | uop vs off: user CPU | null band (user) |
|---|---|---|---|---|
| Moorhuhn | nondeterministic (off~off2 1.4%) | -14.3% | -30.0% | 1.2% |
| Moorhuhn 2 | identical | -4.5% (0.1 s resolution) | -14.0% | 0.5% |
| Moorhuhn Winter | nondeterministic (off~off2 17.6%) | -13.2% | -13.9% | 1.1% |
| Moorhuhn 3 | nondeterministic (off~off2 67%) | -25.3% | -32.4% | 0.2-2.5% |
On Moorhuhn 3 the x87 fold (uopfold vs uop) is inside its 3.5% null band,
so it has no measurable effect, even though an FPU MP3 filter is its hottest
gameplay block.
9. Entry cost: chaining measured, windows kept (2026-09-28)
Chaining would let a program exit that lands on another installed
program's head jump straight into that program. It was measured before
anything was built, with uop-game-ab.js --arms=uop on the gameplay routes
under --branch-clock:
| game | enters | exit lands on a live head | same program as the last entry |
|---|---|---|---|
| Heroes III | 7.81M | 2,188 (0.03%) | 83% |
| Diablo | 115.4M | 3,779 (0.003%) | 70% |
| StarCraft | 3.52M | 12,295 (0.36%) | 68% |
Programs exit into threaded code that is not another loop head, so chaining would remove at most 0.36% of entries. It was not built.
Where an entry's cost went. Every entry poisoned all of its windows,
2.1 to 3.7 per entry on average. The first access through each window then
missed and paid $uop_run → $uop_reguard → $uop_window_set →
$g2w_affine_span, plus the code-page walk for a written window. These
first touches after poisoning, rather than streams walking off a page, were:
| game | first-touch re-guards | share of all re-guards | per entry |
|---|---|---|---|
| Heroes III | 20.5M | 90% | 2.6 |
| Diablo | 168.8M of 256.6M | 66% | 1.5 |
| StarCraft | 7.9M | 86% | 2.3 |
Windows now outlive a run. Header +28 holds the $UOP_WIN_EPOCH under
which the windows were last poisoned. The enter op poisons only when the
shared epoch has moved since then. The epoch is bumped atomically, after the
change, in four places:
$guest_page_publish_rangeor$guest_page_clear_range, when a present PTE is replaced or removed;$code_page_mark, when a page first becomes code;$code_note_decode, when it widens the sparse generated-code span.
Each program now owns its window slots, placed after its code, instead of all programs sharing one set. Interleaved programs therefore keep their windows too.
The SMC hole this closed. It predates this change:
- The poison loop wrote
rw = 0into every slot, and nothing ever set it back to 1. - So a re-guard of a store window never asked
$code_write_is_code. - A program's store into a page holding decoded code therefore went straight to memory. It did not retire the blocks it overwrote, and it did not kill a program lowered from those bytes.
The fix:
- Slots carry
rwfrom compile time.$uc_encode_writemarks the windows of store ops, and poisoning leavesrwalone. - A failed
$uop_window_setnow leaves the slot poisoned instead of at span 0. For a 4-byte access, span 0 reads as "hit everything", which mattered only while a slot could outlive its run.
What the fix cost, and the retirement that pays for it. With rw
honoured, StarCraft's head exits went from 283 to about 770K. Four programs
(0x4b4417, 0x4c789f, 0x4b57a1, 0x4b43f6) store into pages that hold
decoded code on their first access. Each entry therefore exited at its own
head with zero blocks run: 720K entries that did no work at all. The
poor-retirement check only ran on non-head exits, so these programs were
never retired. It now also runs on a head exit that spent no block. Such an
exit made no progress, so the same enters >= 256 && blocks < 2*enters
rule applies to it unchanged. After the change, StarCraft has 2,891 head
exits and 35 programs retired as poor (15 before).
test-uop-compiler.js window-keep covers all three parts:
- A program proves a store window on a page.
- The page then gets decoded code. The program's next store there must exit,
so that threaded code invalidates the block. A build without the
$code_page_markbump fails here with the stale immediate. - Further stores into that page must get the program retired. It retires after 236 head exits, is not entered again, and the threaded store it leaves behind still lands.
Measured with uop-game-ab.js --arms=off,off2,uop,refuop, where refuop
is the base build. Load was 9-19, so timings are noisy.
| game | frames off~uop | re-guards base → new | windows kept | gameplay phase uop vs refuop | null band (off2/off) |
|---|---|---|---|---|---|
| Heroes III | IDENTICAL | 22.76M → 5.51M (-76%) | 7,807,899 of 7,808,452 | 15.2s vs 15.4s | 1.2% |
| Diablo | IDENTICAL | 256.6M → 147.3M (-43%) | 115,365,124 of 115,365,490 | 6.3s vs 6.4s | 0.0% |
| StarCraft | 0.94% (off~off2 1.41%) | 8.87M → 2.08M (-77%) | 3,609,198 of 3,689,426 | 1.7s vs 1.6s | 5.6% |
For Heroes III, enters, blocks and installs are identical between the two builds. Diablo differs by a few installs, and the base build alone varies by that much from run to run (179, then 177). StarCraft has 7.7% more enters. Those are programs whose stores now correctly exit partway through a trip on a code page.
The re-guard counters fall a great deal, but the time saved is inside the noise: -1.3% on Heroes III gameplay against a 1.2% band, and -0.9% on Diablo's whole-run user CPU. The per-entry saving is real but small next to what an entry already costs. It needs the quiet box to price.
What remains unguarded is the cross-thread case §3 already documents: another instance decoding or remapping during this instance's run. Between runs the epoch catches that case too, because the epoch is shared.
10. Coverage: what declined scans actually stop on (2026-09-28)
--uop-census now also records the instructions that end each declined scan:
kind-9 records, keyed by an opcode signature. tools/uop-census.js prints them
as "unsupported instructions in declined scans, by form", weighted by the
threaded block entries at the declined head in the histogram window. Before
this change, the top non-FPU forms were:
| game | top unsupported forms (weight = block entries at head) |
|---|---|
| StarCraft | mov eax,moffs 14.5M, sbb 8.6M, push imm8 7.5M, call 6.1M, setcc 5.3M, shr r,cl 4.3M |
| Diablo | mov eax,moffs 67M, mov al,moffs 32M, setcc 16M, mov moffs,eax 12M |
New lowerings in 07e, all without any call on the fast path:
A0-A3moffs. A kind-5 move with a disp32-only memory operand.rol/ror r/m, imm. Kind 11, ops 3 and 4.- Shift by CL (kind 18). The flag effect depends on the runtime count, and
a zero count leaves the flags untouched. The lowering therefore leaves the
record in the globals (state G):
BNZL(64, free) skips the flag write on a zero count, andSETSS(65, free) stores the sign shift. Exits need nothing extra. setcc(kind 19, cc at R+20).$uc_setvforwards cmp/sub, logic/test and inc/dec/add recipes into SLTU/SLT/EQ forms. Every other case materializes the record and runsGETCC(66), a service op that calls$eval_cc, so the result is bit-exact with the threaded path.- 32-bit
sbb(kind 20). Covers the ALU18-1Dforms and group80/81/83 /3, and writes flag_a/flag_b exactly as the threaded$set_flags_sub(a, b+CF, r)does.adcand the 8/16-bit forms are still declined.
Results. hu = head-unsupported, nb = no-backedge. Each census is one run of
the game's route. The hot-window share is the share of threaded entries the
tier did not take, by verdict.
| game | installs before → after | declines before → after | hot-window verdicts after |
|---|---|---|---|
| StarCraft | 578 → 711 | nb 1111 → 899, hu 248 → 205 | nb 15.4% → 3.9%, hu 6.2% → 1.2%, poor 4.3% → 0.2% |
| Heroes III | 290 → 432 | nb 1370 → 1205, hu 418 → 364 | unchanged (its hot code is x87) |
| Diablo | 249 → 254 | nb 835 → 804, hu 1663 → 1535 | hu 26.1% → 22.1% |
Every targeted form is gone from all three censuses. What still stops a scan
is almost entirely stack and control transfer: push r, call, push imm,
pop r, ret, ret imm, loop, jmp (as a head). After those come div,
rep movsd/rep cmpsb and lodsb. Push, pop and call cover the most
weight, well ahead of everything else. That is the next coverage step:
straight-line stack traffic, and possibly inlining a callee that returns.
The FPU is a separate step.
Speed (bench box, 2026-09-28). Candidate = 38144a80 + this section's
lowerings; reference = 38144a80. uop-game-ab.js --arms=uop,refuop,uop2,refuop2 --ref-wasm, serial, idle box (load 0.0), mean of two repeats per arm.
| game | frames | user CPU, candidate vs main | gameplay slice | null band (user) |
|---|---|---|---|---|
| StarCraft | nondeterministic (repeats of one build differ 1.2-1.5%) | 19.66 vs 31.46 s, -37.5% | 1.2 vs 1.75 s | 0.1% / 0.8% |
| Diablo shareware | identical | 195.5 vs 198.2 s, -1.35% | 5.15 vs 5.3 s | 0.4% / 0.5% |
| Heroes III | identical | 80.2 vs 85.0 s, -5.6% | 5.65 vs 5.7 s | 1.3% / 0.4% |
| Warcraft III, software GL | identical | 114.1 vs 137.7 s, -17.2% | 1.8 vs 2.0 s | 0.1% / 0.1% |
Most of the win is guest-thread time, which the main-thread gameplay slice
does not see: StarCraft's thread 1 runs 340M blocks in the tier against 84M,
and Warcraft III's Miles audio thread 432M against 90M (the moffs forms were
what kept its mixer loop out). Installs rise on every game (StarCraft 344 ->
425, Warcraft III 743 -> 791); StarCraft's head-unsupported declines fall
from 841 to 206.
11. Remaining bottlenecks (2026-09-28)
§8 and §9 asked whether the tier pays. This section asks the next question:
with the tier and the x87 fold both on by default (89c6890c), where does
gameplay CPU go now, and what should be built next? Every number is from the
gameplay window of a tools/uop-game-ab.js route (batches split..max-1), not
the whole run.
11.1 Method, and a tooling bug that invalidated earlier histograms
- Box. The quiet bench box (4 vCPU, load 0-1 throughout), worktree
~/profat 60a24244, which has the same tree as main 38144a80. Runs were serial, one at a time. - CPU.
--cpu-prof-window=SPLIT:ENDwith--cpu-windowuser CPU. It covers the main instance and every cooperative guest-thread instance. wasm-function indices were named withtools/wasm-func-name.js --dumpand bucketed by subsystem.--slice-split's "guest slice" is not the right denominator for a threaded game. On WC3g the main slice is 2.6s of a 26.4s window, and the rest is guest threads. - Coverage and loads.
--uop --uop-census --handler-hist-thread=N --hist-json --hot-block-dump --batch-stats --decode-stats. A local measurement-only--input=B:uop-statsaction snapshots theuop_statscounters of every instance at SPLIT and END. The tier's share of block entries is (uop blocks) / (uop blocks + threaded block hits) over the same window. None of these patches are committed. - The bug:
--handler-histturns the tier off on the thread it profiles.$handler_hist_enabledraises$dbg_any, which raises$dbg_chain_guard.$th_uop_enter(07d) and the$branch_end_atfast path both bail on$dbg_chain_guard. So every histogram window measured the tier-off state, including §8's per-thread histograms and anyuop-census.js --histcoverage figure. The box runs here patch the$th_uop_enterguard to(i32.and $dbg_chain_guard (i32.eqz $handler_hist_enabled)). This belongs in main. Until it lands, a histogram taken with--uopsilently describes a different program. Fixed since:$th_uop_enternow tests$dbg_tier_guard, which is$dbg_chain_guardwithout the histogram (13-exports.wat$dbg_recompute). The transfer fast paths still take the desk under--handler-hist, because$hot_block_hist_recordruns there; that costs time, not coverage.--break/--watch/--count/--trace-*still hold the tier off. test-uop-compilerhist-keeps-tierchecks both halves. - Uncovered: guest-thread verdicts.
--uop-censusrecords from guest threads are not in the log.uop-census.js --thread=Nfinds no[i32 TN]records, so the verdicts for WC3's audio thread below are inferred from its disassembly, not observed. Cause, since fixed: cooperative threads' records were in the log all along, tagged[i32 T<tid>]; theuop[thread 0x…]summary names the thread HANDLE, and--thread=wants the tid. The summary now printstid=N, anduop-census.js --thread=0xHANDLEmaps through it.--threadsworkers now forward their log under--uop-censustoo (they have no exit dump, so only live events appear).
11.2 Where the time goes, per game
Share of window user CPU, by self time, grouped by subsystem:
| group | H3 | Diablo | SC | MH3 | WC3g |
|---|---|---|---|---|---|
| window user CPU | 12.6s | 45.8s | 14.7s | 6.9s | 26.4s |
| threaded dispatch/handlers | 30.9% | 34.7% | 49.9% | 16.5% | 43.4% |
uop tier ($uop_fast…) |
19.3% | 6.1% | 20.4%¹ | 8.4% | 4.1% |
| x87 (fold + handlers) | 21.3% | ~0 | ~0 | 0.3% | 13.3% |
wasm other ($read_thread_word, $bx_hot_bump, paint scans…) |
10% | 26% | 15.1% | — | 16.6% |
memory translation ($g2w/$gl32/$gs32) |
8.4% | 8.1% | 7.3% | — | 12.4% |
decode/cache ($page_*, $code_page_test, decode) |
5.2% | 13.9% | 6.1% | — | 8.3% |
| DirectDraw handler | — | — | — | 60.1% | — |
JS (harness canvas + h.log) |
4.7% | 7.1% | — | — | — |
¹ SC's uop group is dominated by $uop_code_write (14.5%). $uop_fast
itself is 5.0%.
Named hot spots (self time unless marked incl):
- H3.
$uop_fast18.3%.$th_x87_island14.0% incl, pipeline4 3.4% incl,$th_fpu_mem_ro3.7% incl.$branch_end_at7.1% incl.$read_thread_word4.3%,$g2w4.0%.$decode_block3.4% incl,$invalidate_code_write2.0% incl,$bx_hot_bump1.3%.
- Diablo.
$win32_dispatch27.3% incl, of which$handle_PeekMessageA16.1% incl. The paint and non-client scans underPeekMessageA_fetchcost:$paint_flag_mine6.4% +$nc_flags_scan3.1% +$paint_select_next_dirty2.2% +$paint_drain_native_control_paints2.1% = 13.8% self.- Block transfer:
$branch_end_at6.3% (19.9% incl),$page_resolve3.4%,$page_enter3.2%. - Stack handlers:
$th_push_r3.8%.$gs32is 8.0% incl, of which the SMC check$invalidate_code_writeis 5.3% incl, mostly from push/call. $code_page_test3.1%,$gl322.9%,$bx_hot_bump2.6%.- The run makes 424M API calls.
- SC.
$uop_code_write14.5%, reached through$invalidate_code_range←$gs8←$th_store8_ro.$invalidate_code_writeis 16.2% incl.- The cache counted 6.95M page invalidations in the window, and only 18,301 of them dropped a block.
$read_thread_word7.0%,$branch_end_at6.7% (17.2% incl),$uop_fast5.0%,$bx_hot_bump3.0%.
- MH3.
$handle_IDirectDrawSurface_BltFast60.1% self (61.4% incl via$win32_dispatch). Its SRCCOLORKEY path is a scalar per-pixel loop. It branches on bytes-per-pixel inside the loop and recomputesrow*pitchper pixel. - WC3g.
$th_x87_island10.1% incl.$g2w4.6%,$gl323.7%,$guest_page_translate1.2%.$read_thread_word4.8%,$branch_end_at10.5% incl,$bx_hot_bump2.8%.- GL is 0.2%.
- The time is on the Miles audio thread. See the next table.
Tier coverage and the threaded remainder (load-immune counts, window only):
| tier share of block entries | threaded ops/block | threaded remainder by handler class | |
|---|---|---|---|
| H3 main | 77.1% (81M uop vs 24.1M) | 9.74 | alu/mov 40.9%, mem 39.7%, branch 6.9%, stack/call 5.3% |
| H3 T1 (audio) | 35.2% | — | mem 31.8%, alu 26%, x87 24.7% |
| Diablo main | 28.4% (732M threaded entries) | 2.86 | stack/call/ret 43.4% (push_r 14.1, pop_r 9.3, call_rel 5.9, call_ind 4.3, ret 4.0, push_i32 3.9, ret_imm 2.0), alu 17.9%, mem 16.9%, branch 16.8% |
| SC main | 36.3% (T0xe1001 adds 37.8M uop blocks) | — | alu 39.3%, mem 27.1%, branch 13.3%, stack 6.6% |
| MH3 main | 66.5% | — | mem 30.5%, stack/call 25.4%, alu 21.9% |
| WC3g main | 47.0% (79.5M vs 89.7M) | 7.46 | alu 40.2%, mem 25.7%, stack/call 16.2% |
| WC3g audio thread (h=0xe1006) | ~5% (≈3.1M uop vs 60.4M per ⅓ window) | 7.11 | alu 46.3%, mem 31.1%, branch 13.8%, x87 4.1% |
The WC3g audio thread runs ~1.29G threaded dispatches over the window,
about twice the main thread's 669M. Two blocks make up 54% of its block
entries: Mss32.dll 0x2113c300/0x2113c334, the Miles resampling mixer
(mov eax,[moffs]; … imul; add [edi],eax; …; add edx,[moffs]; jnb head). The
MP3 decoder in mp3dec.dll (+0x38f6 and neighbours) is most of the rest.
Why the untaken entries were not taken (share of threaded entries, by the block's verdict as a head):
| no-backedge | head-unsupported | no-verdict (of which in a shared hot slot) | poor | live | |
|---|---|---|---|---|---|
| H3 T0 | 41.5% | 14.2% | 15% | 15% | 10.9% |
| H3 T1 | 74% | — | — | — | — |
| Diablo | 37.6% | 36.0% | 22.1% (13.7%) | 0.1% | 3.9% |
| SC | 25.9% | 9.5% | 31.9% (18%) | 14.7% | — |
| MH3 | 44.4% | — | 39.8% (30.4%) | — | — |
| WC3g main | 31.4% | 10.6% | 12.9% (7.3%) | 0.8% | 0.7% |
Notes on the verdicts:
- H3's biggest head-unsupported case is one switch loop,
jmp [0x472a9c+ecx*4]at exe+0x47227c. Its case blocks 0x472266, 0x472283, 0x47229d and 0x472320 are each 6.17% of threaded entries, about 31% of T0's remainder. - SC's poor heads. storm.dll+0x1502508b is poor (5.7%). exe+0x4b43ea and 0x4b43f6 (4.1% each) are poor only because their stores land on a code page.
- Diablo's remainder is short storm.dll functions whose heads are
mov eax,[esp+4],push eaxandret(2.2% each). That is call-heavy straight-line code with no back edge for the tier to key on. - Hot-table churn. The 512-slot table saw 51.9M (H3), 73.0M (Diablo), 11.2M (SC), 25.7M (MH3) and 120.6M (WC3g) slot takeovers in the run.
Decode is not a lever any more. Gameplay decodes in the window were:
| game | decodes |
|---|---|
| H3 | 5,033 (93.9% of batches decode-free) |
| Diablo | 858 |
| SC | 1,851 |
| MH3 | 306 |
| WC3g | 34,572 ($decode_block 1.3%) |
Batches overwhelmingly stop on "budget spent".
Machine-code sizes (SpiderMonkey Ion, arm64, tools/wasm-native.js) that
bear on the ideas below:
| function | instructions | notes |
|---|---|---|
$uop_fast |
1146 | |
$th_uop_enter |
280 | 2 indirect tail calls |
$fpu_exec_reg |
1415 | 10 indirect calls |
$fpu_exec_mem |
594 | |
$x87_island_body |
111 | a compare chain into the two above, per op |
$branch_end_at |
226 | |
$bx_hot_bump |
82 | |
$read_thread_word |
12 | not inlined, 237 call sites |
$uop_code_write |
111 | a linear scan over $uop_nranges |
$handle_IDirectDrawSurface_BltFast |
487 |
11.3 Ranked ideas
Saving = measured share × plausible speedup of that share, per game. Shares are self time unless marked incl.
- BltFast colour-key blit: specialise and vectorise.
- What: hoist the bytes-per-pixel switch out of the pixel loop, keep row
pointers, and do the key compare and select with v128 (
i8x16.eq/v128.bitselecton 8bpp,i16x8on 16bpp). - Games: MH3, plus every DirectDraw sprite game that blits with a colour key (unmeasured).
- Share: 60.1% of MH3.
- Saving: −45-50% MH3 CPU (at 4-6x on the blit).
- Cost/risk: low. One handler, and its output is checkable pixel-for-pixel against the scalar path.
- Evidence: MH3 cpu-prof,
$handle_IDirectDrawSurface_BltFast60.1% self.
- What: hoist the bytes-per-pixel switch out of the pixel loop, keep row
pointers, and do the key compare and select with v128 (
- Stop
$uop_code_writescanning every uop range on every code-page store.- What: gate it on a per-page "has uop range" bit, set at install and cleared at flush, or at least on a hull test over all ranges.
- Games: SC, and any game that writes data on pages it also executes.
- Share: 14.5% of SC.
- Saving: −13-14% SC.
- Cost/risk: low. The check must stay conservative, and test-uop-compiler's SMC cases cover it.
- Evidence: SC cpu-prof,
$uop_code_write←$invalidate_code_range←$gs8←$th_store8_ro.
- Make PeekMessage's empty-queue path O(1).
- What: keep dirty/non-client counts, or a summary bit, so
$paint_flag_first/any/select_next_dirtyand$nc_flags_scando not walk MAX_WINDOWS per poll when nothing is pending. - Games: Diablo, and every PeekMessage-polling game loop.
- Share: 13.8% self in Diablo (16.1% incl under
$handle_PeekMessageA). - Saving: −12% Diablo.
- Cost/risk: low-medium. The counters must stay exact, and the paint-order tests guard it.
- Evidence: Diablo cpu-prof.
- What: keep dirty/non-client counts, or a summary bit, so
- uop compiler: accept
mov eax,[moffs32]/mov [moffs32],eax(A1/A3).- What: 07e decodes the 88-8B forms but not the moffs encodings, so any loop whose body uses them is declined.
- Games: WC3g. Miles is also used by H3 and others, but not verified there.
- Share: WC3's Miles mixer head
0x2113c300has one in its second instruction and another in its tail. Its two blocks are 54% of the audio thread's block entries. The audio thread is ~⅔ of WC3g's threaded dispatches, and threaded is 43% of WC3g CPU, so the loop is ≈15% of WC3g CPU. - Saving: −7-10% WC3g at the tier's 2-3x.
- Cost/risk: trivial. It is a disp32 memory operand with no base.
- Evidence: WC3g guest-thread hist (
e07300/e07334= 27.1% each) plus disassembly. The verdict itself was not observed (§11.1).
- Cheaper x87 island body.
- What:
$x87_island_bodydispatches each op through a compare chain into$fpu_exec_mem(594) or$fpu_exec_reg(1415 instructions, 10 indirect calls). Pre-decode each island op to a direct small handler index at fold time, and keep ST(0)/ST(1) in locals across the island. Alternatively, give the uop tier an f64 register class for pure x87 islands inside loops. - Games: H3, WC3g, and H3's MP3 thread.
- Share: H3
$th_x87_island14.0% incl; WC3g 10.1% incl. - Saving: −6-7% H3, −4-5% WC3g at 2x.
- Cost/risk: medium. The fold's results must stay bit-exact, which test-x86-ops' x87 cases check.
- Evidence: H3 and WC3g cpu-prof, plus Ion sizes.
- What:
- Inline
$read_thread_word.- What: make it a
defmacro, as dispatch-next is. It is 12 instructions, called from 237 sites, and V8 does not inline it. - Games: all.
- Share: H3 4.3%, Diablo 3.5%, SC 7.0%, WC3g 4.8%.
- Saving: −2-3.5% everywhere, taking call overhead as about half of it.
- Cost/risk: trivial. The body grows, measured at the §8 dispatch-macro scale.
- Evidence: all five cpu-profs.
- What: make it a
- push/pop/call/ret in the tier, with shallow callee inlining.
- What: the lowering §7-§8 deferred. The tier still keys on back edges, so on its own this converts loops that call leaves, not Diablo's loopless call chains. The Diablo win needs call-inlined traces (a trace head at a hot call target).
- Games: Diablo; also MH3 (stack/call 25% of remainder) and WC3g main (16%).
- Share: Diablo stack/call/ret is 43.4% of threaded dispatches, the threaded group is 34.7% of CPU, and 72% of Diablo's entries are untaken.
- Saving: −10-20% Diablo if half the call chains convert; −3-5% MH3 and WC3g.
- Cost/risk: high. It needs ESP-relative guest stores under the SMC guard, and exact exceptions at every push.
- Evidence: Diablo census (no-backedge 37.6% + head-unsupported 36.0%, storm.dll leaf functions).
- Sub-page code-write granularity.
- What: split
$code_page_testinto 64-256 B code bits, or a per-page "code range" hull, so a data store beside code is not an invalidation. - Games: SC, Diablo.
- Share: SC's 6.95M invalidations dropped a block 0.26% of the time, and
two heads (8.1% of untaken entries) are "poor" only because of
same-page stores. Diablo's
$invalidate_code_writeis 5.3% incl from stack pushes, and$code_page_testis 3.1%. - Saving: −2-4% SC (beyond idea 2, plus un-poored heads); −3-4% Diablo.
- Cost/risk: medium. Correctness is central (a missed SMC is silent), and
--trace-code-writesis the check. - Evidence: SC and Diablo cpu-prof, SC cache counters, SC census.
- What: split
- Multiway branch in the tier (
jmp [tbl+r*4]).- What: lower an in-image jump table as a guarded br_table over its in-loop targets, and exit on any other target.
- Games: H3 (other switch loops unmeasured).
- Share: about 31% of H3 T0's untaken entries, with H3 main already at 77% coverage, is ≈6% of H3 CPU.
- Saving: −3% H3.
- Cost/risk: medium. Table bounds are read from guest memory, so there is SMC and table-write exposure.
- Evidence: H3 census, switch loop at exe+0x47227c.
- A bigger, or 2-way, hot table.
- What: grow the 512-slot table so hot heads stop evicting each other.
- Games: SC, MH3, Diablo.
- Share: no-verdict entries in a shared slot are 18% (SC), 30.4% (MH3) and 13.7% (Diablo) of untaken entries.
- Saving: −1-3%. Many of those blocks are bodies, not heads, so this is an upper bound.
- Cost/risk: trivial (a region size). Try it first, because it is cheapest to price.
- Evidence: census hot-table lines, with takeovers in the tens of millions.
- Skip
$bx_hot_bumpfor blocks that already have a verdict or an installed program.- Games: all five.
- Share: 1.3-3.0% self.
- Saving: −1-2%.
- Cost/risk: trivial.
- Headless only: the per-API
log+log_api_exithost calls.- What: two host calls per Win32 call even under
--quiet-api(88M of each in one H3 run). The browser no-ops them, so this is benchmark hygiene rather than product speed.--quiet-api-fastexists and should become what--quiet-apidoes. - Share: Diablo
h.logis 2.3%.
- What: two host calls per Win32 call even under
Not recommended:
- Exit chaining of uop windows. It takes <0.4% of entries (§9), and
block chaining priced +1.4% slower on the box
(docs/block-chaining-design.md §11.1). Block transfer
(
$branch_end_at+$page_resolve+$page_enter) is 13% of Diablo self time, but that time is paid on call/ret transfers, which idea 7 removes and chaining would not. - Decode or cache work. Decode rates are low on every game (the table above).
- String ops, adc, setcc in the tier. None shows as a measurable share of any threaded remainder in these windows.
Order of work:
- Build ideas 1, 2, 3, 4 and 6. Each is a day or less, with a large, single-game-proven share.
- Fix the handler-hist guard (§11.1).
- Then idea 5.
- Idea 7 is the large project, and the only one that moves Diablo's remaining two thirds.