A toyvm-style micro-op tier for the main emulator

Status: design + phase-0 prototype, 2026-09-26. Off by default; nothing ships from this document until phase 0 says the engine is worth building.

1. Why the block executor did not pay, measured in machine code

The H454/H458 block executor (src/07c-block-exec.wat, --block-exec) was timed on the quiet box for the first time in September 2026:

app executor off → on native share
Heroes III gameplay 13.52 → 13.10 ms/batch (−3%), boot +9%, total user CPU +7% 97.6%
MechWarrior 3 gameplay 553.8 → 519.1 ms/batch (inside a ~9% null band) 99.87%
MCM race 2.345 → 2.462 s (+5%) 95.6%
Diablo II to Act I load 21.99 → 23.01 s (+4.6%) —

Coverage was not the problem: 95–99.9% of ops ran natively. So the question is what each native op costs, and tools/wasm-native.js --func='$th_block_exec' (SpiderMonkey Ion, arm64) answers it directly.

A threaded handler is ~150 Ion instructions including its dispatch. The executor removed that dispatch and paid the same amount back in its own prefix. That is why it measured flat.

toyvm's E1 engine (tools/toyvm/uop-wasm.js) pays 11 instructions per transition on the same compiler. It has the same one-function br_table shape, but:

toyvm measured the full effect at ns per x86 instruction (V8/SM):

engine ns / x86 insn
L1 interpreter 3.71 / 2.82
E1 loop engine 0.64 / 0.92
straight wasm (not an engine, upper bound) 0.23 / 0.20

Engine shape and the optimizer passes are the lever, not op coverage. (Memories: project_block_exec_native_cost, project_toyvm_uop_tier.)

2. Constraints this design keeps

3. The engine (E1 shape)

One function, $uop_run(pc) -> exit code: a loop over one br_table on i32.load(pc).

Vregs live in memory.

Ops are a small RISC set, all 32-bit, each doing only its own work:

family ops
move movi, mov
ALU add sub and or xor + immediate forms, shli shri sari, mul, sx8 sx16 zx8 zx16, merge8l merge8h merge16, ext8h
branch compare-and-branch bcc cc,a,b,w,target (flag-forwarded cmp/test + jcc) and bccf (test one materialized flag bit)
memory ld{8,16,32}{s,u} and st{8,16,32}, addressed as window-relative base + (index << s) + disp
guard guard w, base, lo, hi, rw, exit
flags rec op,a,b,res,shift: writes the five lazy globals; emitted only on exit paths where flags are live
control jmp, clock n, exit (one per loop header), exit eip

Memory: the hoisted range guard (the user's rule). Guest virtual mappings do not change between iterations of a loop that makes no calls out of the program. So translation is proved once per window, not once per access:

What the window proves, and for how long:

Flags. The lowering starts naive, like toyvm: every flag producer is a rec. Two passes then work on it:

Surviving recs sit on exit stubs, so the steady-state loop writes no flag globals at all.

Clock. The threaded path charges $steps per block transfer. The program keeps the same accounting exactly: clock n at each loop header subtracts the iteration's block transfers, and exits with a precise EIP when the budget runs out.

4. The lowering and the passes

Input is the region the H458 machinery already discovers, a loop nest of decoded blocks. But the lowering reads the x86 instruction, not the handler index, for the reason tools/toyvm/uop-x86.js gives: a handler index hides its operands. The main emulator's decoder already decodes ModRM and SIB; the lowering calls into that decode rather than copying it.

Passes are ported from tools/toyvm/uop-opt.js, first those that toyvm measured as paying:

  1. promote: guest regs stay in their vregs; temps coalesce.
  2. constprop + addrfold: [esi+4] becomes one address expression.
  3. forwardFlags + flagLiveness: bcc, dead rec removal.
  4. guard hoisting: the per-access window checks above.
  5. licm: loop-invariant loads ([esp+0x40], a LUT base).
  6. clock: one check per header.
  7. dce.

Later, if the census says so: inlined call/ret (toyvm's call contexts), and load forwarding (toyvm measured it as nearly free: project_toyvm_uop_tier).

Where the optimizer runs: in WAT (src/07e-uop-compiler.wat, since 2026-09-27). It started as lib/uop-compiler.js behind a synchronous host import, which was the quickest route to a first measurement, and was ported so browser workers and every host get it without an RPC and without a host call per hot head. $uop_try (07d) calls $uop_compile directly; the compiler decodes the guest bytes, forms the loop, lowers it and writes the encoded program into $UOP_ARENA. That is data, not code, so the no-codegen rule holds. Its working memory is the $UOP_CSCRATCH region; nothing in it outlives one compile.

The port was checked word for word against the JS reference before the JS was deleted: every instruction address of every test-uop-compiler.js case under both clocks, and every head MW3, MCM, D2 and Heroes III compiled on their benchmark routes, produced identical programs (or the same decline reason).

Exports: uop_compile(eip) (what $uop_try calls; answers the program, 1 when another thread holds the scratch, or 0), uop_cstat(k) (compiled, declined, instructions, uops, flushes, words), uop_decline_count(reason).

5. Phases and the numbers that gate them

Phase 0: price the engine, not the compiler. Add the $uop_run engine plus a bench-loops.js arm that runs a hand-lowered program for existing shapes (lut, blk_mix512, and a Heroes III RLE-run shape) next to the threaded and block-executor arms, in one process with alternating reps. Hand-lowered means written the way the optimizer is expected to emit it, so phase 0 is toyvm's "what is the engine worth on this loop" before any compiler work.

Phase 1: JS lowering + passes on real regions (CLI). Install on the loops H458 would install on, and compare final state against threaded code.

Phase 2: move the lowering into WAT; browser and worker coverage.

Phase 3: calls. Inline short leaf calls, so regions stop ending at every call. The H458 census shows shortChain and unsafeOp as the biggest decline rows, and MW3's regions are almost all one block because of calls.

6. What phase 0 does NOT prove

7. Phase 0 results (2026-09-26)

Engine code (Ion, arm64, tools/wasm-native.js --func='$uop_run'). $uop_run is 2992 bytes of Ion code, against 11 KB for $th_block_exec.

Timing (node tools/uop-engine-bench.js --bytes=4m --reps=7, laptop at load ~5.5, V8). Every rep of every arm had identical registers, CF/ZF/SF/OF, EIP and destination hash; no guard failed.

shape threaded ns/insn block-exec uop ns/insn uop vs threaded
lut (byte LUT blit) 12.50 x0.96 3.28 x3.81
ckey (colour-key diamond) 12.71 x0.93 3.60 x3.53
h3shadow (Heroes III 0x471da6, 16-bit) 11.97 x0.96 3.19 x3.75

How to read it:

Not yet shown:

Phase 1 is the question now.

8. Games, and the call-free engine (2026-09-27)

Games (node tools/uop-game-ab.js, each game's own gameplay route, both arms --branch-clock, laptop at load ~2-4). Frames first: every game's final frame matches off vs uop, or differs by no more than off vs off.

game frames whole-run user CPU gameplay phase
Heroes III identical -7.7% -41%
Diablo (shareware) identical -6.2% -34%
StarCraft 1.48%, null band 1.43-1.51% ~-5.5% --
Warcraft III (menu, software GL) identical -3.4% flat (GL-bound)
Warcraft III gameplay (wc3g, headless GL) identical -8.8% -20% (4.4 -> 3.5s); map load -7%
Warcraft III gameplay, software GL identical -5.7 to -7.5% 0% and +13% in two concurrent pairs (noise; GL-bound)
Heroes II, Diablo demo menu identical flat flat
Diablo II nondeterministic off vs off too -3.5% --

The call-free split. With any call inside the dispatch loop, Ion kept $pc and $budget in stack slots: 110 [x20,#28] references in the old $uop_run, a store after every op and a reload on the pc chain every op depends on. The loop called $uop_reguard from 16 memory ops plus $uop_window_set, $get_cf (x2) and $eval_cc. Now $uop_fast makes no call at all and hands any op that needs one back to $uop_run (a window miss, GUARD, SAVECF, GETCF, BCC). A missed op is re-run from scratch after the re-guard; nothing in it changed before its window check. Result: pc lives in w0 with zero stack references in the loop, and a register MOV is 7 instructions plus an 8-instruction dispatch.

What it bought is small, and the reason is the finding:

So windows that survive across entries (an epoch bumped on every mapping change and code-page mark, shared across Worker instances) would buy under 1% of Heroes III. That is not worth a stale-window SMC hole. The remaining lever is coverage: the declines are no-backedge and head-unsupported, and Heroes III's hot threaded time is FPU code the tier does not lower.

(Section 9 built those windows anyway. The epoch closes the stale-window hole within a thread. Building them also found a real SMC hole in store windows.)

Where that FPU time is, measured (2026-09-28). It is not on the thread the tier works on. --handler-hist per guest thread over gameplay batches 4100-4300 (--handler-hist-thread=1,3,2; the main thread over 4100-5101):

thread x87 dispatches share of that thread's dispatches
T0 (game) 585 of 1.19G 0.00%
T1 (start 0x8414a0) 66.6M of 141M 47%
T2, T3 0 0%

So the ~18% x87 CPU in the profile is all T1, and "x87 in uops" would be aimed at the wrong thread. The existing semantic x87 folds (--x87-fusion: pipeline4, island, affine) already reach it. On T1 they cut unfused x87 dispatches from 66.6M to 14.1M (plus 4.2M fused: H449 pipeline4 3.08M and H451 1.13M). That is 79% of x87 dispatches absorbed, and T1's total dispatches fall from 141M to 93M. With --uop as well, T1 reads 15.3M raw and 4.6M fused, so the fold and the tier compose: the tier is on T0, the fold on T1. The remainder is mostly $th_fpu_mem_ro (8.4M).

The fold is off for Heroes III: x87Fusion: true is set only for ut2003_demo in lib/apps.js. Fewer dispatches is not a time win by itself (on MCM the fold was flat until the x87 file moved to memory; see project_x87_fusion_mcm), so the time A/B is the fold / uopfold arms of tools/uop-game-ab.js.

Bench box, 2026-09-28 (x86_64 V8, node 20, Ryzen 9 9950X, 4 vCPU, idle, serial arms, HEAD 84e79bb4). gameplay is the --slice-split main-thread guest slice, so guest-thread (T1) work shows only in user CPU.

game frames uop vs off: gameplay uop vs off: user CPU null band (user / gameplay)
Heroes III identical -52.9% -5.2% 1.8% / 3.3%
Diablo shareware identical -43.2% -4.1% 0.2% / 0%
Warcraft III, software GL (wc3g)* identical -13.0% -6.0% 0.3% / 0%
StarCraft nondeterministic (off~off2 differ too) -5.9% (0.1 s resolution) -2.3% 0.9% / 0%

*wc3g needs the then-uncommitted Game.dll ordinal-import linking from the working tree; on bare HEAD, Game.dll's DllMain stops at KERNEL32.#00001.

Heroes III x87 fold, same box, frames identical in every pair: uopfold vs uop user CPU -4.3% (86.58/86.53 s vs 90.46/90.49 s, null band 0.03%), gameplay slice 0.0%. fold vs off is -5.3% (null 0.9%). All of the saving is guest-thread time, as the dispatch counts predicted.

Moorhuhn, same box. These ran on the working tree as of 2026-09-28 morning; routes are mh1/mh2/mhw/mh3 in tools/uop-game-ab.js, and every final frame was looked at and shows a live round:

game frames uop vs off: gameplay uop vs off: user CPU null band (user)
Moorhuhn nondeterministic (off~off2 1.4%) -14.3% -30.0% 1.2%
Moorhuhn 2 identical -4.5% (0.1 s resolution) -14.0% 0.5%
Moorhuhn Winter nondeterministic (off~off2 17.6%) -13.2% -13.9% 1.1%
Moorhuhn 3 nondeterministic (off~off2 67%) -25.3% -32.4% 0.2-2.5%

On Moorhuhn 3 the x87 fold (uopfold vs uop) is inside its 3.5% null band, so it has no measurable effect, even though an FPU MP3 filter is its hottest gameplay block.

9. Entry cost: chaining measured, windows kept (2026-09-28)

Chaining would let a program exit that lands on another installed program's head jump straight into that program. It was measured before anything was built, with uop-game-ab.js --arms=uop on the gameplay routes under --branch-clock:

game enters exit lands on a live head same program as the last entry
Heroes III 7.81M 2,188 (0.03%) 83%
Diablo 115.4M 3,779 (0.003%) 70%
StarCraft 3.52M 12,295 (0.36%) 68%

Programs exit into threaded code that is not another loop head, so chaining would remove at most 0.36% of entries. It was not built.

Where an entry's cost went. Every entry poisoned all of its windows, 2.1 to 3.7 per entry on average. The first access through each window then missed and paid $uop_run → $uop_reguard → $uop_window_set → $g2w_affine_span, plus the code-page walk for a written window. These first touches after poisoning, rather than streams walking off a page, were:

game first-touch re-guards share of all re-guards per entry
Heroes III 20.5M 90% 2.6
Diablo 168.8M of 256.6M 66% 1.5
StarCraft 7.9M 86% 2.3

Windows now outlive a run. Header +28 holds the $UOP_WIN_EPOCH under which the windows were last poisoned. The enter op poisons only when the shared epoch has moved since then. The epoch is bumped atomically, after the change, in four places:

Each program now owns its window slots, placed after its code, instead of all programs sharing one set. Interleaved programs therefore keep their windows too.

The SMC hole this closed. It predates this change:

The fix:

What the fix cost, and the retirement that pays for it. With rw honoured, StarCraft's head exits went from 283 to about 770K. Four programs (0x4b4417, 0x4c789f, 0x4b57a1, 0x4b43f6) store into pages that hold decoded code on their first access. Each entry therefore exited at its own head with zero blocks run: 720K entries that did no work at all. The poor-retirement check only ran on non-head exits, so these programs were never retired. It now also runs on a head exit that spent no block. Such an exit made no progress, so the same enters >= 256 && blocks < 2*enters rule applies to it unchanged. After the change, StarCraft has 2,891 head exits and 35 programs retired as poor (15 before).

test-uop-compiler.js window-keep covers all three parts:

Measured with uop-game-ab.js --arms=off,off2,uop,refuop, where refuop is the base build. Load was 9-19, so timings are noisy.

game frames off~uop re-guards base → new windows kept gameplay phase uop vs refuop null band (off2/off)
Heroes III IDENTICAL 22.76M → 5.51M (-76%) 7,807,899 of 7,808,452 15.2s vs 15.4s 1.2%
Diablo IDENTICAL 256.6M → 147.3M (-43%) 115,365,124 of 115,365,490 6.3s vs 6.4s 0.0%
StarCraft 0.94% (off~off2 1.41%) 8.87M → 2.08M (-77%) 3,609,198 of 3,689,426 1.7s vs 1.6s 5.6%

For Heroes III, enters, blocks and installs are identical between the two builds. Diablo differs by a few installs, and the base build alone varies by that much from run to run (179, then 177). StarCraft has 7.7% more enters. Those are programs whose stores now correctly exit partway through a trip on a code page.

The re-guard counters fall a great deal, but the time saved is inside the noise: -1.3% on Heroes III gameplay against a 1.2% band, and -0.9% on Diablo's whole-run user CPU. The per-entry saving is real but small next to what an entry already costs. It needs the quiet box to price.

What remains unguarded is the cross-thread case §3 already documents: another instance decoding or remapping during this instance's run. Between runs the epoch catches that case too, because the epoch is shared.

10. Coverage: what declined scans actually stop on (2026-09-28)

--uop-census now also records the instructions that end each declined scan: kind-9 records, keyed by an opcode signature. tools/uop-census.js prints them as "unsupported instructions in declined scans, by form", weighted by the threaded block entries at the declined head in the histogram window. Before this change, the top non-FPU forms were:

game top unsupported forms (weight = block entries at head)
StarCraft mov eax,moffs 14.5M, sbb 8.6M, push imm8 7.5M, call 6.1M, setcc 5.3M, shr r,cl 4.3M
Diablo mov eax,moffs 67M, mov al,moffs 32M, setcc 16M, mov moffs,eax 12M

New lowerings in 07e, all without any call on the fast path:

Results. hu = head-unsupported, nb = no-backedge. Each census is one run of the game's route. The hot-window share is the share of threaded entries the tier did not take, by verdict.

game installs before → after declines before → after hot-window verdicts after
StarCraft 578 → 711 nb 1111 → 899, hu 248 → 205 nb 15.4% → 3.9%, hu 6.2% → 1.2%, poor 4.3% → 0.2%
Heroes III 290 → 432 nb 1370 → 1205, hu 418 → 364 unchanged (its hot code is x87)
Diablo 249 → 254 nb 835 → 804, hu 1663 → 1535 hu 26.1% → 22.1%

Every targeted form is gone from all three censuses. What still stops a scan is almost entirely stack and control transfer: push r, call, push imm, pop r, ret, ret imm, loop, jmp (as a head). After those come div, rep movsd/rep cmpsb and lodsb. Push, pop and call cover the most weight, well ahead of everything else. That is the next coverage step: straight-line stack traffic, and possibly inlining a callee that returns. The FPU is a separate step.

Speed (bench box, 2026-09-28). Candidate = 38144a80 + this section's lowerings; reference = 38144a80. uop-game-ab.js --arms=uop,refuop,uop2,refuop2 --ref-wasm, serial, idle box (load 0.0), mean of two repeats per arm.

game frames user CPU, candidate vs main gameplay slice null band (user)
StarCraft nondeterministic (repeats of one build differ 1.2-1.5%) 19.66 vs 31.46 s, -37.5% 1.2 vs 1.75 s 0.1% / 0.8%
Diablo shareware identical 195.5 vs 198.2 s, -1.35% 5.15 vs 5.3 s 0.4% / 0.5%
Heroes III identical 80.2 vs 85.0 s, -5.6% 5.65 vs 5.7 s 1.3% / 0.4%
Warcraft III, software GL identical 114.1 vs 137.7 s, -17.2% 1.8 vs 2.0 s 0.1% / 0.1%

Most of the win is guest-thread time, which the main-thread gameplay slice does not see: StarCraft's thread 1 runs 340M blocks in the tier against 84M, and Warcraft III's Miles audio thread 432M against 90M (the moffs forms were what kept its mixer loop out). Installs rise on every game (StarCraft 344 -> 425, Warcraft III 743 -> 791); StarCraft's head-unsupported declines fall from 841 to 206.

11. Remaining bottlenecks (2026-09-28)

§8 and §9 asked whether the tier pays. This section asks the next question: with the tier and the x87 fold both on by default (89c6890c), where does gameplay CPU go now, and what should be built next? Every number is from the gameplay window of a tools/uop-game-ab.js route (batches split..max-1), not the whole run.

11.1 Method, and a tooling bug that invalidated earlier histograms

11.2 Where the time goes, per game

Share of window user CPU, by self time, grouped by subsystem:

group H3 Diablo SC MH3 WC3g
window user CPU 12.6s 45.8s 14.7s 6.9s 26.4s
threaded dispatch/handlers 30.9% 34.7% 49.9% 16.5% 43.4%
uop tier ($uop_fast…) 19.3% 6.1% 20.4%¹ 8.4% 4.1%
x87 (fold + handlers) 21.3% ~0 ~0 0.3% 13.3%
wasm other ($read_thread_word, $bx_hot_bump, paint scans…) 10% 26% 15.1% — 16.6%
memory translation ($g2w/$gl32/$gs32) 8.4% 8.1% 7.3% — 12.4%
decode/cache ($page_*, $code_page_test, decode) 5.2% 13.9% 6.1% — 8.3%
DirectDraw handler — — — 60.1% —
JS (harness canvas + h.log) 4.7% 7.1% — — —

¹ SC's uop group is dominated by $uop_code_write (14.5%). $uop_fast itself is 5.0%.

Named hot spots (self time unless marked incl):

Tier coverage and the threaded remainder (load-immune counts, window only):

tier share of block entries threaded ops/block threaded remainder by handler class
H3 main 77.1% (81M uop vs 24.1M) 9.74 alu/mov 40.9%, mem 39.7%, branch 6.9%, stack/call 5.3%
H3 T1 (audio) 35.2% — mem 31.8%, alu 26%, x87 24.7%
Diablo main 28.4% (732M threaded entries) 2.86 stack/call/ret 43.4% (push_r 14.1, pop_r 9.3, call_rel 5.9, call_ind 4.3, ret 4.0, push_i32 3.9, ret_imm 2.0), alu 17.9%, mem 16.9%, branch 16.8%
SC main 36.3% (T0xe1001 adds 37.8M uop blocks) — alu 39.3%, mem 27.1%, branch 13.3%, stack 6.6%
MH3 main 66.5% — mem 30.5%, stack/call 25.4%, alu 21.9%
WC3g main 47.0% (79.5M vs 89.7M) 7.46 alu 40.2%, mem 25.7%, stack/call 16.2%
WC3g audio thread (h=0xe1006) ~5% (≈3.1M uop vs 60.4M per ⅓ window) 7.11 alu 46.3%, mem 31.1%, branch 13.8%, x87 4.1%

The WC3g audio thread runs ~1.29G threaded dispatches over the window, about twice the main thread's 669M. Two blocks make up 54% of its block entries: Mss32.dll 0x2113c300/0x2113c334, the Miles resampling mixer (mov eax,[moffs]; … imul; add [edi],eax; …; add edx,[moffs]; jnb head). The MP3 decoder in mp3dec.dll (+0x38f6 and neighbours) is most of the rest.

Why the untaken entries were not taken (share of threaded entries, by the block's verdict as a head):

no-backedge head-unsupported no-verdict (of which in a shared hot slot) poor live
H3 T0 41.5% 14.2% 15% 15% 10.9%
H3 T1 74% — — — —
Diablo 37.6% 36.0% 22.1% (13.7%) 0.1% 3.9%
SC 25.9% 9.5% 31.9% (18%) 14.7% —
MH3 44.4% — 39.8% (30.4%) — —
WC3g main 31.4% 10.6% 12.9% (7.3%) 0.8% 0.7%

Notes on the verdicts:

Decode is not a lever any more. Gameplay decodes in the window were:

game decodes
H3 5,033 (93.9% of batches decode-free)
Diablo 858
SC 1,851
MH3 306
WC3g 34,572 ($decode_block 1.3%)

Batches overwhelmingly stop on "budget spent".

Machine-code sizes (SpiderMonkey Ion, arm64, tools/wasm-native.js) that bear on the ideas below:

function instructions notes
$uop_fast 1146
$th_uop_enter 280 2 indirect tail calls
$fpu_exec_reg 1415 10 indirect calls
$fpu_exec_mem 594
$x87_island_body 111 a compare chain into the two above, per op
$branch_end_at 226
$bx_hot_bump 82
$read_thread_word 12 not inlined, 237 call sites
$uop_code_write 111 a linear scan over $uop_nranges
$handle_IDirectDrawSurface_BltFast 487

11.3 Ranked ideas

Saving = measured share × plausible speedup of that share, per game. Shares are self time unless marked incl.

  1. BltFast colour-key blit: specialise and vectorise.
    • What: hoist the bytes-per-pixel switch out of the pixel loop, keep row pointers, and do the key compare and select with v128 (i8x16.eq / v128.bitselect on 8bpp, i16x8 on 16bpp).
    • Games: MH3, plus every DirectDraw sprite game that blits with a colour key (unmeasured).
    • Share: 60.1% of MH3.
    • Saving: −45-50% MH3 CPU (at 4-6x on the blit).
    • Cost/risk: low. One handler, and its output is checkable pixel-for-pixel against the scalar path.
    • Evidence: MH3 cpu-prof, $handle_IDirectDrawSurface_BltFast 60.1% self.
  2. Stop $uop_code_write scanning every uop range on every code-page store.
    • What: gate it on a per-page "has uop range" bit, set at install and cleared at flush, or at least on a hull test over all ranges.
    • Games: SC, and any game that writes data on pages it also executes.
    • Share: 14.5% of SC.
    • Saving: −13-14% SC.
    • Cost/risk: low. The check must stay conservative, and test-uop-compiler's SMC cases cover it.
    • Evidence: SC cpu-prof, $uop_code_write ← $invalidate_code_range ← $gs8 ← $th_store8_ro.
  3. Make PeekMessage's empty-queue path O(1).
    • What: keep dirty/non-client counts, or a summary bit, so $paint_flag_first/any/select_next_dirty and $nc_flags_scan do not walk MAX_WINDOWS per poll when nothing is pending.
    • Games: Diablo, and every PeekMessage-polling game loop.
    • Share: 13.8% self in Diablo (16.1% incl under $handle_PeekMessageA).
    • Saving: −12% Diablo.
    • Cost/risk: low-medium. The counters must stay exact, and the paint-order tests guard it.
    • Evidence: Diablo cpu-prof.
  4. uop compiler: accept mov eax,[moffs32] / mov [moffs32],eax (A1/A3).
    • What: 07e decodes the 88-8B forms but not the moffs encodings, so any loop whose body uses them is declined.
    • Games: WC3g. Miles is also used by H3 and others, but not verified there.
    • Share: WC3's Miles mixer head 0x2113c300 has one in its second instruction and another in its tail. Its two blocks are 54% of the audio thread's block entries. The audio thread is ~⅔ of WC3g's threaded dispatches, and threaded is 43% of WC3g CPU, so the loop is ≈15% of WC3g CPU.
    • Saving: −7-10% WC3g at the tier's 2-3x.
    • Cost/risk: trivial. It is a disp32 memory operand with no base.
    • Evidence: WC3g guest-thread hist (e07300/e07334 = 27.1% each) plus disassembly. The verdict itself was not observed (§11.1).
  5. Cheaper x87 island body.
    • What: $x87_island_body dispatches each op through a compare chain into $fpu_exec_mem (594) or $fpu_exec_reg (1415 instructions, 10 indirect calls). Pre-decode each island op to a direct small handler index at fold time, and keep ST(0)/ST(1) in locals across the island. Alternatively, give the uop tier an f64 register class for pure x87 islands inside loops.
    • Games: H3, WC3g, and H3's MP3 thread.
    • Share: H3 $th_x87_island 14.0% incl; WC3g 10.1% incl.
    • Saving: −6-7% H3, −4-5% WC3g at 2x.
    • Cost/risk: medium. The fold's results must stay bit-exact, which test-x86-ops' x87 cases check.
    • Evidence: H3 and WC3g cpu-prof, plus Ion sizes.
  6. Inline $read_thread_word.
    • What: make it a defmacro, as dispatch-next is. It is 12 instructions, called from 237 sites, and V8 does not inline it.
    • Games: all.
    • Share: H3 4.3%, Diablo 3.5%, SC 7.0%, WC3g 4.8%.
    • Saving: −2-3.5% everywhere, taking call overhead as about half of it.
    • Cost/risk: trivial. The body grows, measured at the §8 dispatch-macro scale.
    • Evidence: all five cpu-profs.
  7. push/pop/call/ret in the tier, with shallow callee inlining.
    • What: the lowering §7-§8 deferred. The tier still keys on back edges, so on its own this converts loops that call leaves, not Diablo's loopless call chains. The Diablo win needs call-inlined traces (a trace head at a hot call target).
    • Games: Diablo; also MH3 (stack/call 25% of remainder) and WC3g main (16%).
    • Share: Diablo stack/call/ret is 43.4% of threaded dispatches, the threaded group is 34.7% of CPU, and 72% of Diablo's entries are untaken.
    • Saving: −10-20% Diablo if half the call chains convert; −3-5% MH3 and WC3g.
    • Cost/risk: high. It needs ESP-relative guest stores under the SMC guard, and exact exceptions at every push.
    • Evidence: Diablo census (no-backedge 37.6% + head-unsupported 36.0%, storm.dll leaf functions).
  8. Sub-page code-write granularity.
    • What: split $code_page_test into 64-256 B code bits, or a per-page "code range" hull, so a data store beside code is not an invalidation.
    • Games: SC, Diablo.
    • Share: SC's 6.95M invalidations dropped a block 0.26% of the time, and two heads (8.1% of untaken entries) are "poor" only because of same-page stores. Diablo's $invalidate_code_write is 5.3% incl from stack pushes, and $code_page_test is 3.1%.
    • Saving: −2-4% SC (beyond idea 2, plus un-poored heads); −3-4% Diablo.
    • Cost/risk: medium. Correctness is central (a missed SMC is silent), and --trace-code-writes is the check.
    • Evidence: SC and Diablo cpu-prof, SC cache counters, SC census.
  9. Multiway branch in the tier (jmp [tbl+r*4]).
    • What: lower an in-image jump table as a guarded br_table over its in-loop targets, and exit on any other target.
    • Games: H3 (other switch loops unmeasured).
    • Share: about 31% of H3 T0's untaken entries, with H3 main already at 77% coverage, is ≈6% of H3 CPU.
    • Saving: −3% H3.
    • Cost/risk: medium. Table bounds are read from guest memory, so there is SMC and table-write exposure.
    • Evidence: H3 census, switch loop at exe+0x47227c.
  10. A bigger, or 2-way, hot table.
    • What: grow the 512-slot table so hot heads stop evicting each other.
    • Games: SC, MH3, Diablo.
    • Share: no-verdict entries in a shared slot are 18% (SC), 30.4% (MH3) and 13.7% (Diablo) of untaken entries.
    • Saving: −1-3%. Many of those blocks are bodies, not heads, so this is an upper bound.
    • Cost/risk: trivial (a region size). Try it first, because it is cheapest to price.
    • Evidence: census hot-table lines, with takeovers in the tens of millions.
  11. Skip $bx_hot_bump for blocks that already have a verdict or an installed program.
    • Games: all five.
    • Share: 1.3-3.0% self.
    • Saving: −1-2%.
    • Cost/risk: trivial.
  12. Headless only: the per-API log + log_api_exit host calls.
    • What: two host calls per Win32 call even under --quiet-api (88M of each in one H3 run). The browser no-ops them, so this is benchmark hygiene rather than product speed. --quiet-api-fast exists and should become what --quiet-api does.
    • Share: Diablo h.log is 2.3%.

Not recommended:

Order of work:

  1. Build ideas 1, 2, 3, 4 and 6. Each is a day or less, with a large, single-game-proven share.
  2. Fix the handler-hist guard (§11.1).
  3. Then idea 5.
  4. Idea 7 is the large project, and the only one that moves Diablo's remaining two thirds.