AoE Performance Optimization Notes

Current Measurement

Latest AoE campaign gameplay profile:

Timing from the gameplay window:

profile elapsed:       10011.9 ms
main.runSlice:          6580.6 ms
wrapped host imports:    362.0 ms
guest/interpreter:      6218.6 ms
present fps:              21.5

The bottleneck is the WAT x86 interpreter, not canvas repaint and not mostly DIB conversion.

Repeat timing command:

LABEL=baseline RUNS=3 HANDLER_HIST=0 node tools/profile-aoe-repeat.js

This writes /private/tmp/aoe-repeat-<label>-{1,2,3}.json plus a summary JSON and prints mean/min/max/stddev. Use the histogram profiles to choose candidates, then use this repeated no-hist timing before keeping a small optimization.

Hot-loop report command:

node tools/aoe-hot-block-report.js --top=8 --disasm=3

This reads the newest nonempty AoE handler-histogram profile in /private/tmp, prints top handlers/pairs/SIB consumers/branch operands, clusters hot guest VAs, and disassembles the hottest blocks from Empires.exe. Pass an explicit profile path as the first argument when comparing older histogram runs.

Player-name EDIT composition (2026-08-26)

AoE I and II both subclass a native EDIT for the new-player name. The control did receive every character and painted the growing string into its top-level window's canonical GDI surface, but an exclusive DirectDraw frame layer was then composited over that surface. The field looked blank while typing even though the game accepted the hidden name.

lib/renderer.js now retains the suppressed top-level GDI surface as a sparse native-child overlay. Only the rectangles of visible children are copied above DirectDraw, so menu/chrome painting cannot displace the game frame. The focused regression is test/test-directdraw-native-child-overlay.js; real captures show live AOE in AoE I and Codex in AoE II.

Measured Experiments

See also interpreter-dispatch-perf.md, which carries the app-independent dispatch findings (two generic $next changes, both measured zero) and generalises the SIB/br_table reverts below into a rule: handler-op count proves two builds did equal WORK, never that one is faster. It also records the load-independent --handler-hist-thread=N method, which is more reliable than single-run Chrome profiles on a busy machine.

Single-run Chrome campaign profiles are noisy, but the direction has been consistent enough to prune several ideas:

experiment                         profile file                                           main.runSlice   delta vs baseline
baseline                           /private/tmp/aoe-web-profile.json                          6580.6 ms      0.0%
Jcc specialized                    /private/tmp/aoe-web-profile-jcc-specialized.json          6571.9 ms     -0.1%
Jcc + push/pop specialized         /private/tmp/aoe-web-profile-jcc-pushpop.json              6562.3 ms     -0.3%
Jcc + push/pop + base load/store   /private/tmp/aoe-web-profile-jcc-pushpop-loadstorebase.json 6524.4 ms     -0.9%
br_table register helpers          /private/tmp/aoe-web-profile-reg-brtable.json              6618.1 ms     +0.6%
broad SIB fused handlers           /private/tmp/aoe-web-profile-sib-fused.json                6593.9 ms     +0.2%
direct stack load/store handlers   /private/tmp/aoe-web-profile-stackfast.json                6571.6 ms     -0.1%

Disposition:

Generic AoE I/II narrow handlers (2026-08-27)

Two exact instruction-shape matchers now lower the hottest narrow routines without testing a literal executable address. The span executor is shared by both games; its mode describes only the compiler's register allocation and relative continuation offsets:

AoE I bytes  --\
               matcher -> { mode, start EIP } -> shared sort/clip/row-lookup policy
AoE II bytes --/                                -> mode-specific register writeback

AoE I:  x0=EAX x1=EBP min=ECX max=EDX row4=EDI rowHead=EBX
AoE II: x0=EBX x1=EBP min=EAX max=ECX row =EDI rowHead=EAX

The byte-grid handler recognizes the common six-instruction row-fill loop and uses one affine translation, one code-range invalidation, and memory.fill. Aliasing or non-affine mappings take an instruction-equivalent slow arm.

Repeated direct measurements after warmup:

region                         ordinary decoder       narrow handler       speedup
AoE II span prefix, 200k       100.0-101.3 ms          23.1 ms             4.34-4.38x
grid fill, 300k bytes           22.5-22.8 ms             0.1 ms             ~205x

The fixed 1,600-batch full-game route was neutral within run noise once caches were warm (8.98s baseline versus 9.19s both, user CPU), so it does not support a whole-game percentage claim. It did prove the regions are live: 352,912 span executions and 535 grid fills/41,472 bytes. Focused tests compare ordinary and lowered register, stack, observable-flag, memory, and continuation state; pass --bench to repeat the direct measurements.

Current Stack-Packet Prototype

There are disabled-by-default executable packets for exact AoE hot addresses. They are intentionally narrow:

src/07-decoder.wat emits handler 356 only when all are true:
  stack_packet_enabled != 0
  start_eip == stack_packet_addr
  start_eip is one of the hardcoded packet addresses

src/05-alu.wat:$th_stack_packet op=1:
  executes the whole block in Wasm locals
  flushes EAX/ECX/EDX/EDI once at the exit
  preserves lazy CMP flags for the block tail
  sets EIP to 0x0049d9f9 or 0x0049da1a

src/05-alu.wat:$th_stack_packet op=2:
  executes the 0x0049dd20 span-builder prefix through the row-head test
  implements prologue pushes, x0/x1 sorting, row/x clipping, and lazy flags
  side-exits to 0x0049dd8b, 0x0049ddc7, or 0x0049e0ad

Exports:

set_stack_packet_enabled(flag, addr)
reset_stack_packet_counters()
get_stack_packet_entries()
get_stack_packet_0049d9d1_entries()
get_stack_packet_0049dd20_entries()
get_stack_packet_0049dd20_to_dd8b_entries()
get_stack_packet_0049dd20_to_ddc7_entries()
get_stack_packet_0049dd20_to_e0ad_entries()

Coverage:

node test/test-aoe-stack-packet-handler.js
node test/test-aoe-span-trace-handler.js

Gameplay measurement:

LABEL=stack-packet RUNS=3 HANDLER_HIST=0 STACK_PACKET=1 node tools/profile-aoe-repeat.js
LABEL=span-trace RUNS=3 HANDLER_HIST=0 STACK_PACKET=1 STACK_PACKET_ADDR=0x0049dd20 node tools/profile-aoe-repeat.js

This test drives both exits with synthetic guest memory and verifies register state, memory stores, lazy flags, EIP, and counters.

10s repeated gameplay result:

profile                                  runSlice mean   guest mean   present fps   packet hits
/private/tmp/aoe-repeat-baseline-10s      6253.8 ms      5881.5 ms       22.98              0
/private/tmp/aoe-repeat-stack-packet-10s  6294.7 ms      5927.1 ms       21.73        498,890

delta: +40.9 ms runSlice (+0.65%), +45.6 ms guest (+0.78%)

Conclusion: the first handwritten stack packet is a correctness and measurement prototype, not a performance win. It proves the allowlisted packet can execute real AoE gameplay about 500k times per 10s window, but saving this block's threaded dispatch does not offset the cost of the larger WAT handler plus its remaining gl32/gs32 and flag work. Do not expand this backend until the next packet either removes more memory translation/flag work or targets a larger trace where local-state savings are big enough to measure.

0x0049dd20 prefix trace result, same 3x10s campaign gameplay harness:

profile                                      runSlice mean   guest mean   present fps   packet hits
/private/tmp/aoe-repeat-span-baseline-10s      6239.7 ms      5824.6 ms       27.91              0
/private/tmp/aoe-repeat-span-trace-10s         6437.8 ms      6036.4 ms       24.06        939,706

delta: +198.1 ms runSlice (+3.18%), +211.8 ms guest (+3.64%)

Trace exit mix:

to 0x0049dd8b empty-row allocator path: 420k-431k/run
to 0x0049ddc7 non-empty row-list path: 497k-534k/run
to 0x0049e0ad reject path:              1.4k-1.6k/run

Optimization passes tried inside this prefix trace:

opt1:
  direct stack i32.load/store through g2w, not generic gs32
  lazy flags only on real side exits
  result: 6426.0 ms runSlice, 6012.6 ms guest

opt2:
  reuse translated stack base and this-object base inside the trace
  result: 6282.1 ms runSlice, 5870.8 ms guest

opt3:
  delay register-global flushes until side exits
  make packet counters optional; speed run used STACK_PACKET_COUNT=0
  result: 6269.8 ms runSlice, 5888.7 ms guest

empty-inline:
  inline the empty-row insertion path when the local allocator can satisfy it
  counted run: 454k-460k empty-row completions/run, toDd8b=0
  best speed run: 6266.0 ms runSlice, 5809.2 ms guest
  final current-code speed run: 6341.3 ms runSlice, 5928.2 ms guest

append-inline:
  tried append-after-tail non-empty insertion
  counted run: only 38k-41k append completions/run
  result: slower, because every non-empty row paid extra branch checks
  action: backed out append inline code; keep only empty-row inline

Fresh current-build baseline for the same harness was noisy:

/private/tmp/aoe-repeat-span-baseline-current-10s:
  6374.8 ms runSlice, 5887.6 ms guest

Conclusion: after local optimizations and empty-row insertion inlining, the trace is still roughly neutral rather than a reliable speed win. It does prove that avoiding helper calls, delaying register flushes, and inlining a genuinely hot side path matter. The attempted append-after-tail path is too rare to justify checks on every non-empty row. The next useful trace work should target one of the hotter non-empty merge/delete paths only if profiling can prove a cheap selector for that path.

Handler Histogram

The handler histogram is enabled only during the gameplay measurement window. Pair counts reset at decoded-block boundaries, so pair data is useful for superinstruction selection.

total handlers:        368,967,847
intra-block pairs:     302,370,952

Top handlers:

13.934%  $th_load32_ro
12.989%  $th_jcc
 6.648%  $th_push_r
 4.338%  $th_compute_ea_sib
 4.170%  $th_load32
 4.065%  $th_mov_r_r
 3.586%  $th_pop_r
 3.421%  $th_cmp_r_r
 2.979%  $th_store32_ro
 2.747%  $th_alu_r_m32_ro

Top intra-block pairs:

3.861%  $th_load32_ro       -> $th_load32_ro
3.725%  $th_cmp_r_r         -> $th_jcc
3.061%  $th_alu_r_m32_ro    -> $th_jcc
2.976%  $th_push_r          -> $th_push_r
2.520%  $th_pop_r           -> $th_pop_r
2.401%  $th_compute_ea_sib  -> $th_load32
2.135%  $th_test_r_r        -> $th_jcc
1.856%  $th_push_r          -> $th_mov_r_r
1.702%  $th_load32_ro       -> $th_compute_ea_sib
1.632%  $th_push_r          -> $th_load32_ro

Category-level read:

RO base+disp/SIB:   ~33%
loads/stores:       ~30%
branches/calls:     ~18%
stack:              ~12%
ALU/test/cmp:       ~62%
FPU:                 ~3%

Categories overlap because a handler can be both ALU and RO, or load/store and RO.

Branch Operand Histogram

Profile output: /private/tmp/aoe-web-profile-branch-operands.json.

The branch operand histogram records producer operands only when the next threaded handler is a Jcc. This is more useful for choosing fusions than raw handler-pair counts.

cmp_r_r -> Jcc total:          11,398,509
test_r_r -> Jcc total:          6,372,538
alu_r_m32_ro -> Jcc total:      9,373,017

Top cmp_r_r -> Jcc shapes:

889,989  cmp ebp,ecx -> jl
876,525  cmp eax,ecx -> jge
847,520  cmp eax,edx -> jg
847,409  cmp ebp,edx -> jle
844,095  cmp eax,ebp -> jle
746,196  cmp ebx,ebp -> jl
461,092  cmp ecx,eax -> jg
446,390  cmp edi,ecx -> jl

Top test_r_r -> Jcc shapes:

1,277,128  test eax,eax -> jz
  951,533  test eax,eax -> jnz
  857,627  test ebx,ebx -> jnz
  572,725  test esi,esi -> jz
  342,707  test edx,edx -> jge
  293,793  test eax,eax -> jl

Top alu_r_m32_ro -> Jcc shapes are all CMP forms:

1,208,861  cmp ecx,[edi+disp] -> jl
1,190,395  cmp edx,[edi+disp] -> jle
  874,012  cmp eax,[ebp+disp] -> jl
  862,752  cmp edi,[esi+disp] -> jg
  842,302  cmp edi,[esi+disp] -> jl
  798,429  cmp edx,[ebx+disp] -> jl

Implication:

Branch Fusion Probes

Several test r,r -> jz/jnz fusion variants were tried after the operand histogram. They reduced dispatch count, but did not produce a durable speedup.

hist-enabled profile                                         main.runSlice   result
/private/tmp/aoe-web-profile-branch-operands.json              6539.2 ms     baseline
/private/tmp/aoe-web-profile-branch-fused.json                 6539.4 ms     flat
/private/tmp/aoe-web-profile-test-jznz-fused.json              6563.0 ms     slower by 23.8 ms
/private/tmp/aoe-web-profile-test-jcc-decode-fused.json        6621.9 ms     slower by 82.7 ms
/private/tmp/aoe-web-profile-test-jcc-specialized-fused.json   6592.3 ms     slower by 53.1 ms

Production-style no-hist timing was slightly more favorable for the shape-specific decode fusion, but too small to justify the code complexity without repeated-run confirmation:

profile                                                       main.runSlice   guest/unwrapped
/private/tmp/aoe-web-profile-baseline-nohist.json               6241.8 ms       5865.6 ms
/private/tmp/aoe-web-profile-test-jcc-specialized-fused-nohist.json
                                                                 6222.6 ms       5842.6 ms

An even narrower cmp r32,[base+disp] + signed Jcc fusion was also tested after the exact SIB jump win. It preserved CMP flags, then branched with direct signed comparisons for JL/JGE/JLE/JG, so it avoided $do_alu32 and the separate Jcc flag-reader path but did not skip lazy flag writes.

variant                                runSlice mean   runSlice sd   guest mean
jmp [disp+eax*4] kept baseline             6243.9 ms      10.6 ms     5876.1 ms
cmp r,[base+disp] + signed Jcc fused       6575.3 ms      36.8 ms     6159.4 ms

Profile files:

/private/tmp/aoe-repeat-cmpmem-jcc-{1,2,3}.json

Conclusion:

Flag-Dead Branch Selector

Tool:

node tools/aoe-branch-fusion-candidates.js \
  --profile=/private/tmp/aoe-web-profile-hot-blocks-32k.json \
  --hot-limit=160 \
  --top-rows=8

This ranks only hot producer/Jcc tails where both branch exits overwrite flags before reading them. That is the missing condition from the failed flag-preserving fusion probes above.

10s hot-block result:

branch-tail hot block entries:             28,425,669   7.6% vs all handlers
flag-dead branch entries:                  15,917,674   42.5% of covered blocks
fusable flag-dead producer entries:        13,871,998    3.7% vs all handlers
conservative branch-dispatch saves:        13,871,998    3.7% vs all handlers
producer flag writes skipped:              13,871,998    3.7% vs all handlers

both exits overwrite flags immediately:       934,982    6.7% of fusable
both exits overwrite flags within 2 insns:  5,025,218   36.2% of fusable
both exits overwrite flags within 4 insns: 11,056,989   79.7% of fusable

cmp r32,r32 signed dead:                    4,362,163    1.2% vs all handlers
cmp r32,[base+disp] signed dead:            4,243,248    1.1% vs all handlers
cmp r32,[base+disp] immediate-dead:           133,086    0.0% vs all handlers
test r32,r32 dead:                          1,078,046    0.3% vs all handlers
test r32,r32 self dead:                       938,172    0.3% vs all handlers

Implications:

Runtime prototype, not kept:

variant                         runSlice mean   runSlice sd   guest mean
cmp-rr safe disabled control        6527.2 ms     105.5 ms      6112.4 ms
cmp-rr signed dead safe fused       6538.8 ms      34.3 ms      6130.4 ms
delta                                +11.6 ms                   +18.0 ms

The safe prototype did not produce a durable win. Keep the selector, but do not keep this runtime peephole. Future work should use the same liveness data inside a broader threaded-IR/block-local compiler stage where it can also remove register writes and reuse memory address work.

Second runtime prototype, also not kept:

variant                                  runSlice mean   runSlice sd   guest mean
block-local prototype disabled control       6301.7 ms      24.6 ms      5925.2 ms
mov/add/dec/cmp/Jcc dead prototype           6281.7 ms      22.3 ms      5918.3 ms
delta                                         -20.1 ms                    -6.9 ms

Profile files:

/private/tmp/aoe-repeat-blocklocal-disabled-control-{1,2,3}.json
/private/tmp/aoe-repeat-blocklocal-mov-add-dec-cmp-jcc-{1,2,3}.json

Conclusion:

do not keep this standalone runtime handler.
it is mechanically valid, but the measured CPU delta is below browser noise.
the block has real local savings, but one exact shape only covers 0.33%.
the next implementation should be a compiler-selected block-local lowering,
not another hand-added exact peephole.

Threaded IR Liveness Report

This is not a generated-Wasm direction. The scalable path is still generated threaded code, but with a small IR pass after x86 parsing:

x86 decode -> normalized IR -> local liveness/specialization -> threaded words

Compiler pipeline sketch:

guest EIP
  |
  v
+-------------------+
| decode x86 bytes  |
| op/modrm/sib/imm  |
+-------------------+
  |
  v
+-------------------+      examples:
| normalized IR     | ---> cmp r32,[base+disp]
| explicit effects  |      test r,r
| explicit operands |      load32 dst,[base+disp]
+-------------------+      jcc cc,target
  |
  v
+------------------------------+
| local analysis per block     |
| - flag liveness              |
| - branch exits               |
| - concrete reg/base/index    |
| - hot profile weight         |
+------------------------------+
  |
  v
+----------------------------------------+
| threaded-code selection                |
|                                        |
| flags live?  -> existing safe handlers |
| flags dead?  -> *_jcc_dead primitive   |
| hot SIB?     -> exact SIB primitive    |
| generic case -> existing generic path  |
+----------------------------------------+
  |
  v
+-------------------------+
| emitted threaded words  |
| handler id + operands   |
+-------------------------+
  |
  v
current tail-call threaded interpreter

What the compiler can optimize before emitting threaded words:

flag-dead branch producers
  cmp/test result used only by immediate Jcc, both exits overwrite flags first
  => skip set_flags_sub/set_flags_logic and branch directly

operand-specific memory ops
  [base+disp], [disp+eax*4], common SIB shapes
  => avoid generic EA helpers and sentinel paths only for measured hot shapes

dispatch reduction
  combine multiple IR ops only when the fused primitive removes real helper work
  => avoid one-dispatch fusions that just make bigger slower handlers

profile-directed selection
  hot block histogram chooses which primitives are worth adding
  => no broad "optimize every possible shape" handler bloat

RISC-like micro-IR model:

x86 instruction:
  cmp edx, [ebx+0x8]
  jl target

micro-IR:
  t0 = add ebx, 0x8
  t1 = load32 t0
  t2 = sub edx, t1        ; produces virtual flags
  br_lt t2, target, fall

threaded output choices:
  if flags live:
    th_alu_r_m32_ro(cmp edx,[ebx+8])
    th_jcc_l(fall,target)

  if flags dead:
    th_cmp_r_m32_ro_jl_dead(edx, ebx, 8, fall, target)

Important constraint:

Do not execute every micro-op as its own threaded handler.
That would increase dispatch count and likely regress.

Use micro-ops to reason, then pack them back into coarse threaded primitives.

Useful micro-IR nodes:

EA(base,index,scale,disp)     address calculation, no memory side effects
LOAD(width, ea)               memory read
STORE(width, ea, value)       memory write
ALU(op, width, a, b)          arithmetic/logical value op
FLAGS(op, width, a, b, res)   virtual flags, can be dead
BR(cc, flags, fall, target)   conditional branch
MOV(dst, src)                 register transfer
CALL/JMP/RET/YIELD            hard block boundaries

What this buys:

flag SSA/liveness
  flags become virtual values, not mandatory global writes.
  dead flags can disappear before threaded-code emission.

address specialization
  EA shape is explicit, so [base+disp] and [disp+eax*4] are easy to match.

primitive packing
  many x86 forms map to the same IR, then choose one measured threaded handler.

profile-guided code size
  only hot IR shapes get new primitives; cold shapes stay generic.

Register global-write estimate:

The same IR idea applies to general registers, not just flags. Flags are often consumed by the immediate Jcc, while general registers are visible machine state at block exits. A block compiler can still keep register values virtual inside a block and flush each changed register once at exit.

Toy estimator command:

node tools/aoe-reg-liveness-estimate.js \
  --profile=/private/tmp/aoe-web-profile-hot-blocks-32k.json \
  --hot-limit=120

The estimator is offline only. It decodes hot block entries from the existing profile, excludes ESP by default, treats all non-ESP registers as live at block exit, and counts current full-register writes versus final block-exit flushes.

10s profile result:

covered block entries:        37,475,452 / 66,254,912 = 56.6%
full register writes seen:    90,064,148
final block-exit flushes:     63,194,255
avoidable global writes:      26,869,893   29.8% of seen writes
avoidable vs dispatch count:        7.2%
identity writes:               2,072,787    2.3% of seen writes, 0.6% of dispatches
partial-byte writes ignored:  13,779,133
dead overwritten writes:               0

Naive timing scale from the same profile:

main.runSlice:                          6,559.7 ms
guest/unwrapped:                        6,226.0 ms
block-compiler saved-write upper scale:   474.3 ms, 7.2% of runSlice
identity-write-only scale:                 36.6 ms, 0.6% of runSlice

This is not a benchmark. The scale assumes a saved register write is comparable to one average handler-dispatch share, which is only a rough yardstick. The useful conclusion is relative:

identity-only handlers are small:       about 0.5-0.6%
true block-local virtual regs are real: about 7% dispatch-scale opportunity
trivial dead overwritten writes:        basically absent in the top AoE blocks

Example hot block:

0x00535e08 count=747,413
  mov edx, edi
  add edx, ecx
  dec edx
  cmp edx, [ebx+0x8]
  jl 0x00535e7c

current threaded handlers write EDX three times.
block-local virtual-reg output writes EDX once at exit or side exit.

Toy block compiler printer:

node tools/aoe-block-compiler-printer.js \
  --profile=/private/tmp/aoe-web-profile-hot-blocks-32k.json \
  --addr=0x00535e08

This prints original instructions, approximate current threaded handlers, and a block-local virtual-register/virtual-flag lowering. It is not executable code; it is a shape report before runtime compiler work.

Stack-packet compiler printer:

node tools/aoe-stack-packet-compiler.js \
  --profile=/private/tmp/aoe-web-profile-hot-blocks-32k.json \
  --top=20 \
  --details=5

This is the first offline compiler artifact for the Wasm-friendly stack-packet backend described in docs/wasm-stack-threaded-code.md. It still emits no runtime code. It accepts conservative clean 32-bit blocks, prints packet ops, exit flushes, flag policy, EA/g2w grouping, and weighted bailout reasons.

10s hot-block top-20 result:

rows compiled:                14/20
block-entry coverage:         10,843,970 / 66,254,912 = 16.4%
current-dispatch estimate:    37,289,326 = 10.0% vs all handlers
packet op estimate:          156,445,731
register writes saved:         4,407,901 = 28.2% of compiled current reg writes
flag writes skipped:          10,398,327 = 65.9% of compiled flag writes
virtual flag ops skipped:      4,923,677
exact EA reuses:                 706,030
page-window g2w save est:        689,248

For the larger clean candidate 0x0049d9d1, the packet plan estimates 18 current dispatches, 92 packet micro-ops, register writes 11 -> 4, flag writes 5 -> 1, four skipped virtual flag ops, one exact EA reuse, and five page-window g2w saves across a six-access [esi+disp] memory group. This is a better compiler target than one-off branch or pair fusions because several savings stack in the same block.

Example output shape:

0x00535e08
  original:
    mov edx, edi
    add edx, ecx
    dec edx
    cmp edx, [ebx+0x8]
    jl 0x00535e7c

  current:
    5 threaded dispatches
    3 EDX global writes
    1 g2w call

  toy optimized:
    v_edx = edi
    v_edx = edi + ecx
    v_edx = edi + ecx - 1
    wa0 = g2w(ebx + 0x8)
    branch jl using sub(v_edx, load32(wa0))
    exit flush: edx
    flag exits: fall=dead target=dead, skip flag global write

Top hot blocks show two different classes:

compare/Jcc-only blocks
  no register-write savings, but flag-dead direct branch is useful.

multi-op blocks like 0x00535e08
  register writes coalesce well inside a block.

stack/setup blocks
  often need register flushes and do not benefit much from block-local
  virtual registers unless we also model ESP/stack effects.

Another register-write example:

0x00535aa0
  mov edi, [esi+eax*4]
  or edi, edi
  jz 0x005362e0

`or edi,edi` is an identity register write. It only needs flags, so it can be
selected as a flag-only/test-like primitive without a register global write.

Block-shape classifier:

node tools/aoe-block-shape-census.js \
  --profile=/private/tmp/aoe-web-profile-hot-blocks-32k.json \
  --hot-limit=120 \
  --top-rows=10 \
  --page-size=4096

This aggregates the printer/estimator view across hot blocks. It reports primary buckets, overlapping optimization signals, exact branch candidates, register-coalescing block signatures, EA/g2w access shapes, concrete memory op shapes, and combined block-local candidates.

10s hot-block result:

covered block entries:            37,475,452 / 66,254,912 = 56.6%
current dispatches in covered:   208,333,276   56.1% vs all handlers
branch fusion dispatch saves:     28,425,669    7.6% vs all handlers
flag-dead branch opportunities:   14,881,722   39.7% of covered blocks
coalescible register writes:      26,869,893    7.2% vs all handlers
identity register writes:          2,072,787    0.6% vs all handlers
unknown decode boundary weight:            0    0.0% of covered blocks

Compiler feature buckets after the decoder-coverage pass:

10,268,028  27.4%  clean 32-bit base+disp memory
 8,129,020  21.7%  blocked: non-Jcc control
 7,308,294  19.5%  needs stack/ESP model
 5,160,909  13.8%  clean scalar/branch
 1,898,077   5.1%  clean 32-bit SIB memory
 1,433,003   3.8%  needs partial-width model
 1,347,163   3.6%  clean 32-bit memory writes
 1,010,195   2.7%  clean 32-bit absolute memory
   772,956   2.1%  blocked: FPU model
   147,807   0.4%  blocked: string/rep model

Decoder coverage note:

top-120 hot-block unknown decode boundary is now zero.
earlier "unknown-heavy" buckets were mostly control, FPU/string, byte/word ops,
and immediate forms that the offline decoder did not model yet.
the next compiler work is no longer blocked by decoder visibility for this
profile; it is blocked by choosing which explicit buckets to support first.

EA/g2w result:

blocks with any EA/g2w access:        28,865,060   77.0% of covered blocks
blocks with multiple EA/g2w accesses: 15,230,200   40.6% of covered blocks
blocks with repeated same EA/g2w:      1,326,604    3.5% of covered blocks
blocks with related EA families:       8,486,979   22.6% of covered blocks
current g2w-like memory accesses:     72,938,433   19.6% vs all handlers
repeated-EA g2w saves in block:        2,227,338    0.6% vs all handlers
related-EA memory accesses:           30,321,764   41.6% of current g2w accesses
related-EA adjacent pairs:            19,023,392    5.1% vs all handlers
related-EA delta 4 pairs:              9,395,000    2.5% vs all handlers
related-EA delta <=16 pairs:          14,513,678    3.9% vs all handlers
stable same-base multi-disp blocks:    5,639,724   15.0% of covered blocks
stable same-base accesses:            19,552,195   26.8% of current g2w accesses
stable same-base page-range blocks:    5,639,724   15.0% of covered blocks
stable same-base page-range groups:    6,299,831
stable same-base page-range accesses: 19,552,195   26.8% of current g2w accesses
stable same-base page g2w saves:      13,252,364    3.6% vs all handlers
consecutive same-base page blocks:     4,821,243   12.9% of covered blocks
consecutive same-base page pairs:      7,243,349    1.9% vs all handlers

Interpretation:

repeated same-EA reuse is small on its own.
stable same-base small-window reuse is much larger.
one page-window translation per stable group is a 13.3M g2w-save estimate,
about 5.7x the exact repeated-EA CSE estimate.
all EA/g2w access is large enough to keep high priority.
best g2w work is likely fast-path mapping or base-window prechecks,
not only exact common-subexpression reuse inside a block.

Top EA/g2w access shapes:

13,827,370  [disp]
13,153,396  [esi+disp]
12,481,226  [esp+disp]
 5,307,449  [ebp+disp]
 3,525,440  [edi+disp]

Top concrete memory op shapes:

3,086,374  load32 edx <- [esi+disp]
3,058,963  load32 eax <- [esi+disp]
2,833,582  load32 ecx <- [esi+disp]
2,596,363  load32 ecx <- [esp+disp]
2,592,216  load32 edx <- [esp+disp]
2,416,002  load32 eax <- [esp+disp]
2,027,562  load32 eax <- [ebp+disp]
2,018,750  load8  al  <- [esi]
1,972,275  store32 [disp] <- eax
1,967,088  load32 esi <- [disp]
1,892,237  jmp [eax*4+disp]
1,699,793  cmp edi,[esi+disp]

Destination+base load probe, not kept:

prototype handlers:
  load32 eax/ecx/edx <- [esi+disp]
  load32 eax/ecx/edx <- [esp+disp]

1s hist smoke:
  new handlers total: 2,627,688 / 33,246,383 dispatches = 7.90%

10s no-hist timing:

variant                                  runSlice mean   runSlice sd   guest mean
dst+base load prototype                      6796.0 ms      92.2 ms      6375.4 ms
dst+base load disabled control               7330.1 ms      85.5 ms      6861.5 ms
normal 356-handler reference band            ~6280 ms       ~20 ms       ~5920 ms

Conclusion:

do not keep the destination+base load handlers.
they have real surface, but this session did not produce a safe win.
removing only set_reg from many loads is still too small or too sensitive to
code layout/browser state.
memory work should stay tied to the block-local compiler path, where the same
address/value can be reused across several instructions.

Top stable same-base page-range families:

874,688  [esp+disp] range=0x4  disps=+0x4,+0x8
849,566  [esp+disp] range=0x4  disps=+0x14,+0x18
480,494  [esi+disp] range=0x20 disps=+0x0,+0x4,+0x14,+0x18,+0x20
468,091  [esi+disp] range=0x4  disps=+0x40,+0x44
380,359  [eax+disp] range=0xc  disps=+0x0,+0x4,+0x8,+0xc
380,359  [esi+disp] range=0x10 disps=+0x3c,+0x40,+0x44,+0x48,+0x4c
355,685  [esp+disp] range=0x4  disps=+0x18,+0x1c
327,536  [esi+disp] range=0xa4 disps=+0x28,+0x30,+0xcc

Top combined block-local compiler candidates:

addr        weight   score dispatch reg-save branch-pack page-g2w-save repeat-g2w-save
0x0049d9d1  480,494     13       18        7           0             6               2
0x0049dd92  380,359     11       25        4           0             7               0
0x005086c4  139,874     26       31       16           0            10               0
0x00535b13  314,989      9       20        7           1             1               0
0x00535e08  747,413      3        5        2           1             0               0

Read:

0x00535e08 was the narrow prototype and is too small alone.
0x0049d9d1 is the better clean block-compiler target:
18 current dispatches, 7 saved register writes, and 6 same-page g2w saves.
0x005086c4 has the highest per-block score but is stack-heavy.
0x00535b13 became visible after partial-width decode coverage and is now a
good branch+register+memory mixed target, but it needs 16-bit register support.

Potential base-window primitive:

block has a stable base register feeding [base+d0], [base+d1], ...
and max(d) - min(d) < 4096

guest_min = base + min_disp
guest_max = base + max_disp + access_width - 1

if guest_min and guest_max are in the same mapped/direct page:
  wasm_base = g2w_page_base(guest_min)
  access wasm_base + page_offset(guest_min) + (disp - min_disp)
else:
  fallback to normal per-access g2w

Important caveat:

offset range < page size is not enough by itself.
the base value must stay unchanged between the grouped accesses.
base + min_disp can still straddle a guest page or sparse allocation boundary.
the fast path needs a runtime same-page/contiguous-mapping guard.

Broad g2w page-cache probe:

variant                    runSlice mean   guest mean   result
baseline no-hist             6243.9 ms     5876.1 ms   prior measured baseline
g2w page cache                6598.9 ms     6208.8 ms   slower by 355.0 ms

Profile files:

/private/tmp/aoe-repeat-g2w-page-cache-{1,2,3}.json
/private/tmp/aoe-repeat-g2w-page-cache-summary.json

Conclusion:

do not add a generic g2w page cache.
the added branch/global traffic in every g2w call costs more than it saves.
same-base page-window work still needs compiler-selected grouped primitives,
not a broad cache inside every memory translation.

Top consecutive same-base page pairs:

pair kind totals:
4,925,238  load->load
1,564,657  store->store
  480,494  load->store
  139,874  store->load

849,566  load->load  [esp+disp] +0x18 -> +0x14 delta=-0x4
731,268  load->load  [esp+disp] +0x8  -> +0x4  delta=-0x4
480,494  load->load  [esi+disp] +0x18 -> +0x14 delta=-0x4
480,494  load->load  [esi+disp] +0x14 -> +0x0  delta=-0x14
480,494  load->store [esi+disp] +0x4  -> +0x18 delta=+0x14
468,091  load->load  [esi+disp] +0x44 -> +0x40 delta=-0x4
380,359  store->store [eax+disp] +0x4 -> +0x0  delta=-0x4
380,359  store->store [eax+disp] +0x0 -> +0x8  delta=+0x8

Interpretation:

a tiny consecutive pair primitive has a 7.1M-pair surface, about 1.9%.
that is smaller than full stable grouped reuse, but easier to emit.
the obvious adjacent load->load handler was tested and did not win.
future work should target compiler-local memory lowering, not a generic pair op.

Runtime load-pair probes:

variant                    runSlice mean   guest mean   result
baseline no-hist             6243.9 ms     5876.1 ms   prior measured baseline
loadpair page/direct guard    n/a           n/a         run1 6744.9 ms, later run hit drag timeout
loadpair direct guard         6555.5 ms     6150.7 ms   slower by 311.6 ms runSlice
loadpair dispatch-only        6248.0 ms     5869.0 ms   neutral, +4.1 ms runSlice

Profile files:

/private/tmp/aoe-repeat-loadpair-1.json
/private/tmp/aoe-repeat-loadpair-direct-{1,2,3}.json
/private/tmp/aoe-repeat-loadpair-direct-summary.json
/private/tmp/aoe-repeat-loadpair-dispatch-{1,2,3}.json
/private/tmp/aoe-repeat-loadpair-dispatch-summary.json
/private/tmp/aoe-web-profile-loadpair-hist-smoke.json

Histogram confirmation:

1s gameplay with handler histogram:
handler 356 load32_pair_ro_samebase count=650,741
handler 356 share=1.932% of 33,690,499 handler dispatches

Conclusion:

simple adjacent load-pair fusion has real surface but no measured speedup.
direct-window guards inside the fused handler are actively bad.
dispatch-only fusion proves removing about 1.9% of dispatches is not enough.
memory optimization needs a compiler stage that keeps base/address temporaries
live across several ops and avoids repeated get/set/g2w work without adding
branches to every fused handler.

Current priority from the classifier:

1. implement a block-local threaded-IR compiler skeleton for clean scalar/memory blocks
2. use exact flag-dead branch data inside that compiler, not standalone peepholes
3. keep non-ESP registers virtual inside a block and flush once at exits
4. add compiler-local grouped memory lowering for stable [base+disp] groups
5. add partial-width/stack models after clean 32-bit blocks prove a win

The report command now supports optional hot-block weighting:

node tools/superinstruction-census.js \
  --profile=/private/tmp/aoe-web-profile-hot-blocks-32k.json \
  --hot-limit=120 \
  test/binaries/shareware/aoe/aoe_ex/Empires.exe

Static executable-wide result:

cmp r,[mem]; jcc                  919 sites
cmp r,[base+disp]; jcc            858 sites
cmp r,[mem]; signed jcc           457 sites
cmp r,[mem]; jcc flags dead       475 sites
cmp r,[base]; signed dead         267 sites
test r,r; jcc                    9027 sites
test r,r; jcc flags dead         2350 sites
test r,r self; dead              2343 sites

Hot-block weighted result from the 10s profile, top 120 hot block entries:

covered block entries:        37,475,452 / 66,254,912 = 56.6%
cmp r,[mem]; jcc               8,776,877   23.4% of covered
cmp r,[base+disp]; jcc         7,248,406   19.3% of covered
cmp r,[mem]; signed jcc        8,296,383   22.1% of covered
cmp r,[mem]; jcc flags dead    4,305,364   11.5% of covered
cmp r,[base]; signed dead      3,957,105   10.6% of covered
test r,r; jcc                  2,932,946    7.8% of covered
test r,r; jcc flags dead         700,595    1.9% of covered

ASCII TLDR:

Do not make more flag-preserving branch fusions.
Do make decode-time IR/liveness decide when flags are dead.
Standalone cmp r32,r32 + signed Jcc peephole was safe but not measurably faster.
Next useful path: block-local threaded IR, then flag-dead branch packing.
Memory branch packing should include [base+disp] page/prework reuse.
Use a local flag-dead scan, not one-instruction lookahead.
Keep output as threaded code, not generated Wasm.

Threaded primitive direction:

Conclusion:

Hot Block And SIB Consumer Histograms

The handler histogram now also records guest block-entry EIPs and the concrete consumer of $th_compute_ea_sib. This helps separate "which emulator handlers are hot" from "which AoE routines and SIB shapes are hot."

Profile outputs:

/private/tmp/aoe-web-profile-hot-blocks-32k.json   10s hot-block exploration
/private/tmp/aoe-web-profile-sib-fixed3.json       1s SIB consumer validation

Hot-block 32K table result:

block entries:    66,254,912
occupied slots:        8,071
collisions:       1,145,088   1.73%

Top hot block entries from the 10s run:

857,843  0x00535b56  rva 0x00135b56
850,107  0x0049dd20  rva 0x0009dd20
~846K    0x0049dd33..0x0049dd7e cluster
830,965  0x00536420  rva 0x00136420

Disassembly read:

SIB consumer validation from the 1s run:

recorded SIB events:  1,523,029
table total:          1,504,774
occupied keys:              202
collisions/lost:         18,255   1.20%

Top SIB consumer shapes:

149,324  jmp_ind   op=0x0  [none+eax*4+disp]
113,781  load32    op=0x7  [esi+eax*4+disp]
 89,552  load32    op=0x0  [eax+edx*4+disp]
 63,834  load32    op=0x3  [edi+eax*1+disp]
 62,488  load8     op=0x3  [ecx+eax*8+disp]
 62,350  alu_r_m32 op=0x71 [edi+edx*4+disp]
 58,718  load32    op=0x6  [eax+ecx*4+disp]
 54,015  mov_r16_m16 op=0x0 [esi+ebx*4+disp]

Hot-loop split from the 10s hot-block run:

profile:              /private/tmp/aoe-web-profile-hot-blocks-32k.json
elapsed:              10019 ms
main.runSlice:         6559.7 ms
wrapped host imports:   333.7 ms
guest/unwrapped:       6226.0 ms

The split below is block-count based, not exact wall-clock time. It is still useful because the top 120 recorded hot block entries cover 56.6% of all guest block entries in the sample.

bucket                    guest block share   est. runSlice time
span/region builder              20.8%              1363 ms
blitter command decoder          19.7%              1292 ms
tile bitfield updates             2.2%               145 ms
entity render glue                2.0%               134 ms
40-slot linear search             1.8%               116 ms
x87 float->int helper             1.2%                77 ms
negative-id lookup                0.9%                62 ms
other top-120 blocks              8.0%               522 ms
long tail / not top-120          43.4%              2846 ms

Disassembly interpretation of the two largest buckets:

Hot Loop Disassembly Read

This is the current best read of what the hottest AoE loops are doing. Address names are descriptive, not recovered symbols.

0x0049dd20: per-row span interval builder

Hot-block share in the 10s profile:

0x0049d900..0x0049e300 cluster: 13,115,005 block entries
top entries: 0x0049dd20, 0x0049dd33, 0x0049dd3c, 0x0049dd56,
             0x0049dd61, 0x0049dd6c, 0x0049dd74, 0x0049dd7e

The function at 0x0049dd20 is called with ecx=this and three stack arguments that behave like x0, x1, and row. It:

1. rejects rows outside [this+0x60, this+0x64]
2. sorts x0/x1
3. rejects ranges outside [this+0x58, this+0x5c]
4. clamps the x range to [this+0x58, this+0x5c]
5. indexes row arrays by row*4:
     [this+0x3c] row head/list slot
     [this+0x40] row endpoint/list slot
     [this+0x44] row min-x slot
     [this+0x48] row max-x slot
     [this+0x4c] per-row segment count slot
6. inserts, extends, merges, or deletes interval nodes

Node shape inferred from stores:

node+0x00 = next
node+0x04 = prev
node+0x08 = x0
node+0x0c = x1

Important paths:

0x0049dd20 fast reject/sort/clip
0x0049dd8b allocate first node for an empty row
0x0049ddc7 handle non-empty row list
0x0049dfb9 walk nodes until node.x1 >= new_x0-1
0x0049e005 merge/delete covered following nodes
0x0049e273 caller-side loop emitting spans into 0x0049dd20

Pseudocode:

void add_span(builder, x0, x1, row) {
  if (row < builder->minRow || row > builder->maxRow) return;
  if (x0 > x1) swap(x0, x1);
  if (x1 < builder->minX || x0 > builder->maxX) return;
  x0 = max(x0, builder->minX);
  x1 = min(x1, builder->maxX);

  list = rowLists[row];
  if (!list) {
    rowLists[row] = alloc_node(x0, x1);
    rowMin[row] = x0;
    rowMax[row] = x1;
    rowCount[row]++;
    return;
  }

  find the interval adjacent to or overlapping [x0, x1];
  insert a new node if it is separated by a gap;
  otherwise widen the existing interval;
  remove later intervals fully covered by the widened interval;
}

This is why small single-block packets did not help much. The hot cost is not one arithmetic sequence; it is repeated interval-list control flow, many short branches, allocator/free-list calls (0x0049d9a0, 0x0049da20, 0x0049da40), and row-array memory traffic.

0x005357e0..0x00536b40: clipped software blitter / command decoder

Hot-block share in the 10s profile:

0x00535700..0x00536c00 cluster: 13,044,401 block entries
top entries: 0x00535b56, 0x00536420, 0x00535b5b, 0x00535c20,
             0x00535e00, 0x00535e08, 0x005362e0, 0x00535e7c

The entry around 0x005357e0 clips a requested row/range against global clip bounds at 0x775080..0x77508c, looks up a row span list from [0x775004 + y*4], and initializes scratch globals:

0x775014  destination pointer/current row base plus x offset
0x775018  source/pattern pointer plus caller x offset
0x775020  row base used to convert absolute x to local x
0x775038  current blit line/index
0x77503c  current span-list node
0x775044  continuation target after skip/advance ops
0x775048  draw flags / alignment flags
0x77504c  left clip amount for current run
0x775050  saved command stream pointer
0x775054  saved destination x/pointer

0x00535aa0 is the per-line setup loop:

line = 0x775038
spanNode = rowSpanLists[y + line]
if no spanNode: goto next_line
load per-line source offsets from [0x775010 + line*4]
skip if high-bit sentinel is set
compute destination pointer and source pointer
walk spanNode linked list until [runLeft, runRight] intersects [node+8,node+c]

The inner command decoder is jump-table based:

0x00535c20  forward decoder: read opcode byte, dispatch low nibble via 0x534400
0x00536420  reverse/mirrored decoder: dispatch low nibble via 0x534540
0x00535c80  run with length = high-nibble / 4-ish encoding
0x00535d20  skip/advance by length, then jump to continuation [0x775044]
0x00535e00  clipped draw run, masked blend path using [0x77502c/30]
0x00535e7c  run before current span: skip source bytes and restart decoder
0x00535e8a  run past current span: advance to next span-list node
0x005362e0  next line; loops back to 0x00535aa0 until line count exhausted
0x00536b40  same shape for the reverse/mirrored path

The table at 0x534300 is not a code pointer table; it is a packed byte lookup used to combine run length/alignment/flags into indices for writer tables. The actual low-nibble dispatch table for 0x00535c20 is 0x534400:

0 -> 0x00535c80   1 -> 0x00535d20   2 -> 0x00535c85   3 -> 0x00535d31
4 -> 0x00535c80   5 -> 0x00535d20   6 -> 0x00535d60   7 -> 0x00535e00
8 -> 0x00535c80   9 -> 0x00535d20  10 -> 0x00535ea0  11 -> 0x00535f40
12 -> 0x00535c80 13 -> 0x00535d20  14 -> 0x00535c60  15 -> 0x005362e0

Pseudocode:

for each blit line {
  setup source/destination pointers for this line;
  for each row span that intersects the requested x range {
    while (true) {
      op = *cmd++;
      switch (op & 15) {
        case draw_run:
          len = decode_length(op, cmd);
          clip run to current span node;
          dispatch writer by dest alignment, length class, mask flags;
          break;
        case skip:
          advance destination/source by decoded length;
          goto continuation;
        case end_line:
          next line;
          break;
      }
      if (run is before current span) skip source and continue;
      if (run is after current span) advance span node and retry;
    }
  }
}

This is a software sprite/tile blitter with span clipping. It is hostile to tiny generic fusions because control flow is dominated by indirect jump tables and clip/list branches. The best emulator-side opportunities are exact jmp [disp+eax*4] style dispatch shortcuts, local memory-address reuse inside larger trace packets, and possibly recognizing these jump-table decoder blocks as a trace family rather than optimizing one block at a time.

Smaller hot loops

0x0045c280 and 0x0045c2e0 update packed 2-bit tile state:

index byte = base + (x*25)*8 + (y >> 2) + 0x7ef54
bit shift  = table[y & 3]
field      = (byte >> shift) & 3
0x0045c280 increments the field if it is below 3
0x0045c2e0 decrements the field if it is above 0

0x004c5fa0 is a fixed 40-slot linear search:

for (i = 0, p = this + 0x1c; i < 0x28; i++, p += 4)
  if (*p == needle) return 1;
return 0;

0x00508686..0x00509710 is higher-level render glue around the blitter. It walks visible object rows/lists, tests visibility and state bits, calls into span/blitter setup (0x535720, 0x535760, 0x5357a0, 0x5359a0), and has nested linked-list compare loops around 0x00509564..0x00509710.

Exact SIB Jump Probe

An exact handler for the hottest SIB shape was tested:

candidate: jmp [disp+eax*4]
profile:   /private/tmp/aoe-web-profile-jmp-sib-eax4-1s.json

The handler worked mechanically. In the 1s hist-enabled profile it ran 167,340 times, and the old generic SIB/jump path dropped accordingly:

metric                         previous 1s      exact-SIB 1s
compute_ea_sib count             1,523,029        1,355,298
jmp_ind count                      210,392           74,394
new exact handler count                  0          167,340

A single no-hist run was misleading, so the timing comparison was repeated with three full 10s gameplay profiles per variant:

variant          runSlice mean   runSlice sd   guest mean   present fps mean
baseline             6312.8 ms      48.6 ms     5940.6 ms          22.45
jmp [disp+eax*4]     6243.9 ms      10.6 ms     5876.1 ms          23.01
jmp + load SIB       6242.0 ms       1.3 ms     5874.6 ms          23.08

Profile files:

/private/tmp/aoe-repeat-baseline-{1,2,3}.json
/private/tmp/aoe-repeat-jmp-sib-eax4-{1,2,3}.json
/private/tmp/aoe-repeat-jmp-load-sib-{1,2,3}.json

Conclusion:

Implication:

Optimization Ideas

1. Specialize Jcc

$th_jcc is the second hottest handler. It calls $eval_cc, which is generic over all x86 condition codes. The hottest pairs are cmp/test/alu -> jcc, so branch evaluation is a good first target.

Candidate work:

Expected benefit:

Risk:

2. Fuse SIB Compute With Memory Operation

$th_compute_ea_sib -> $th_load32 is a top pair. The current design computes EA into global $ea_temp, then the following handler reads through the sentinel path. That costs one extra handler dispatch and a global temporary.

Candidate work:

Expected benefit:

Measured result:

Risk:

Next version, if revisited:

3. Stack Fast Paths

push_r -> push_r, pop_r -> pop_r, and pop_r -> ret_imm are very hot. Current push/pop uses generic register helpers and gs32/gl32.

Candidate work:

Expected benefit:

Risk:

Measured result:

4. Register Specialization

CPU sampling shows $get_reg and $set_reg are expensive. Handler histogram confirms generic register-heavy handlers dominate.

Candidate work:

Expected benefit:

Risk:

Measured result:

5. Memory Translation Fast Paths

CPU sampling showed $g2w hot. The direct mapping path is already simple arithmetic, but every load/store still calls it.

Candidate work:

Expected benefit:

Risk:

6. Superinstructions

Top pairs suggest concrete fusions:

load32_ro + load32_ro
cmp_r_r + jcc
alu_r_m32_ro + jcc
push_r + push_r
pop_r + pop_r
compute_ea_sib + load32
test_r_r + jcc
push_r + mov_r_r
load32_ro + cmp_r_r
load32_ro + test_r_r

Good first superinstructions:

Expected benefit:

Risk:

Future Directions To Explore

Longer N-Gram Profiling

Pair histograms are enough for first superinstructions, but triples could reveal stronger patterns:

Use a bounded top-K hash table instead of a dense matrix for triples.

Guest EIP Hot Blocks

Handler-level data tells what kind of emulator work is hot; hot-block data now shows which AoE guest routines drive it.

Next measurements:

Block-Local Native-ish Compilation

The current threaded interpreter emits handler IDs. A next step could emit block-specific WAT-like sequences or JS-generated wasm for hot blocks.

Options:

This is much larger work, but it is the most direct path beyond superinstructions.

Lazy Flag Refinement

Many hot branches need only ZF, SF, OF, or CF. Current lazy flag state is general.

Ideas:

Stack/Code Invalidation Separation

Stores currently guard against self-modifying code. Stack writes are frequent and should almost always skip that path.

Ideas:

SIMD/MMX

MMX/SIMD is probably not a first-order CPU win for AoE right now. The hot data is scalar interpreter dispatch, register muxing, memory translation, and branch evaluation. SIMD may still help DIB conversion or bulk memory/string ops, but that is not where the 10s gameplay profile spends most time.

Browser/Runtime Sensitivity

Repeat these profiles in:

The histogram itself adds overhead, so use it to pick changes, then verify speed without it.

External WASM-Centric References

These sources are specifically about WebAssembly emulators, x86-to-WASM translation, or WASM runtime benchmarking. They are more relevant than generic JavaScript emulator advice.

x86-to-WASM Systems

Emulator Benchmark Sources

WASM Runtime And Measurement Papers

Dispatch Strategy Comparisons

There is no universal answer that br_table, call_indirect, or tail-call dispatch is best for an interpreter written in WASM. The result depends on engine lowering, handler count, state passing, and whether the runtime can keep interpreter state in machine registers.

Approaches:

Practical implication for this project:

WASM Features Relevant To Future Work

Suggested Next Order

  1. Keep and commit the measured 355-handler specialization set.
  2. Use operand histograms for candidate selection, then verify with HANDLER_HIST=0.
  3. Add hot block/EIP sequence profiling so superinstructions can be tied to actual AoE routines, not only global pair counts.
  4. Revisit branch fusion only as generated trace-specific blocks or larger multi-op superinstructions.
  5. Revisit SIB only as shape-specific handlers after operand/addressing histograms identify exact forms.
  6. Revisit stack only as batch prolog/epilog or call/ret forms, not as generic direct stack load/store replacement.
  7. Explore triple histograms once pair-driven wins flatten out.