NFS III emulation profile — 2026-09-29

General unprefixed MOVSD now lowers into the micro-op tier. The initial profile below identified two such instructions blocking an otherwise supported integer loop. Remote validation subsequently removed over 99.97% of its three hot residual block entries per frame. A fixed-work reproduction reduced CPU time by 47–56%; whole-game measurements did not establish an FPS gain. The initial investigation and subsequent implementation results are separated below because their machines, instrumentation and rendering costs differ.

Capture

Initial Glide worktree at 32ef6f09, original NFS III demo, seed 12345, 640×480, AI off, rain on, idle cockpit at the race start. Three 20-second windows per renderer, serial headful Chrome on Apple M1. Artifacts are in build/nfs3-guest-profile/{glide,d3d}/: per-window handler/block JSON, micro-op census logs, page/worker CPU profiles, screenshots and results. WASM SHA-256: ce8952abbd2f2881096e7133e84f3a6b5ce4b24dca649b151f790684790db2b3.

node tools/nfs-renderer-bench.js --cases=glide,d3d --seconds=20 --samples=3 --guest-profile --profile --out=build/nfs3-guest-profile
node tools/browser-handler-hist.js build/nfs3-guest-profile/glide/sample-1-hist.json --top=15 --blocks=15
node tools/uop-census.js build/nfs3-guest-profile/glide/sample-1-uop.log --hist=build/nfs3-guest-profile/glide/sample-1-hist.json --top=12
node tools/hot-loop-census.js build/nfs3-guest-profile/glide/sample-1-hist.json build/nfs3-guest-profile/glide/sample-2-hist.json build/nfs3-guest-profile/glide/sample-3-hist.json --top=12

The harness enables the existing histogram only on the actual guest-main worker, reads all handler slots and block-table entries between slices, and captures existing micro-op census events from initialization. It does not read counters from the idle page instance or enable other workers' shared histogram storage. Census capture is bounded and fails on truncation. Each census log contains startup-to-window events plus that window's final live-program dump; do not sum logs as independent windows. Aggregate runtime counters are differenced against guestBefore for each window.

Dynamic DLLs were absent from the page module map in this threaded run. Attribution was completed from recorded LoadLibrary addresses and the actual DLL PE preferred bases; the harness now does this automatically. Renderer DLL load base is 0xc30000, preferred base 0x60000000. Executable addresses below are unchanged preferred/runtime VAs.

Repeatable guest-code targets

Percentages below are shares of recorded residual threaded block entries, not all guest instructions and not time spent in each loop.

Target Glide windows 1 / 2 / 3 D3D windows 1 / 2 / 3
0x4c5f28, 0x4c5f34, 0x4c5f3a combined 13.77 / 12.95 / 12.55% 12.19 / 11.97 / 11.04%
0x4dec44 repeated x87 stores 2.91 / 2.85 / 2.71% 2.63 / 2.67 / 2.48%

The first loop scans up to 2,000 eight-byte entries backwards from 0x7d7dd4. While the first local pair word is zero it copies the next pair into stack locals using MOVSD; MOVSD at 0x4c5f3f/0x4c5f40; its other branch links nonzero entries. The census identifies MOVSD as unsupported, with a no-backedge decline at the copy branch. Partial programs at neighboring heads retire as poor; D3D retains one partial program, but still returns to the same threaded copy branch. The useful fix is covering the full loop, not disabling poor-program retirement globally.

The implementation uses the existing guarded load from ESI and store to EDI, then adjusts both pointers using runtime DF without changing arithmetic flags. Each copy completes separately, preserving sequential overlap. Prefixed forms retain their existing threaded paths. Remote tests cover both DF directions, overlap, flags, noncontiguous sparse page crossings, writable-code invalidation, and the actual relocated record-loop shape.

The second loop performs four FST m64 stores of ST0, advances the destination 32 bytes, and repeats. It is rejected as head-unsupported at 0x4dec44. A bounded repeated-store fusion is a separate candidate, provided x87 stack/status behavior and write notifications remain correct. It does not justify a wholesale x87 rewrite.

Tier behavior

Counter, per-window range Glide D3D
Successful tier entries 2.05–2.16 million 2.50–2.85 million
Blocks retired in tier 30.17–31.52 million 23.28–25.14 million
Blocks per entry 14.15–15.18 8.55–9.33
Memory guard failures 0 0
New installs 0–3 0–2

About 64–65% of residual block entries have no compile verdict at that exact head. Their hot-table slots are shared with other addresses, but this does not establish collisions as the cause: fallthrough/call-return entries can also lack hot-head eligibility. Do not claim that resizing the table would recover this share.

D3D repeatedly declines 0x4bf939 (320 cumulative attempts by window 3), sharing a verdict-map slot with other heads. This is a secondary churn lead; the measured CPU profile does not establish compilation as a major cost. One D3D live dump contains an implausible head 0x6918000 with no matching compile record. Per-head lifetime work totals from that dump are excluded; the table above uses independent per-instance runtime-counter deltas.

CPU samples and limits

Across three guest-main worker profiles:

Interval-weighted sample category Glide D3D
WASM leaf frames, including instrumentation 84.15% 71.05%
Explicit histogram helpers, included above 11.02% 8.85%
Idle 3.20% 5.96%
D3DIM draw, inclusive — 15.29%
D3DIM fence, inclusive — 4.06%

Both renderers share $run, $uop_fast, $x87_island_fast, and load/store handlers among their largest guest hotspots. Raw sample counts broadly corroborate these rankings. Glide's host RPC wrapper also occupies 11.49% of weighted intervals, mostly under Glide flush/idle barriers; that includes waiting for the main thread and is not all host computation.

Histogram dispatch counts are not instruction counts: one x87-island or micro-op handler can execute many instructions. The block histogram omits internal tier execution and reports collision counters of 316k–473k for Glide and 73k–100k for D3D; percentages use recorded hits, not an exact census of all blocks. Proximity-based region grouping does not prove a control-flow loop, which is why the priorities above use disassembled individual blocks.

Histogram helpers alone consume roughly 9–11% of sampled intervals and instrumentation affects dispatch paths. CPU-profile intervals can include waiting/descheduling; they are not OS CPU accounting. System load was about 27–52. These runs identify repeated targets; they provide neither clean FPS comparisons nor a predicted speedup. Rebenchmark any implementation with histograms and CPU profiling disabled.

Remote MOVSD implementation results

Reserved Linux box, four vCPUs on AMD Ryzen 9 9950X, Node 24.15.0, Chrome 152. Browser runs used SwiftShader, with no hardware WebGL. Baseline WASM SHA-256 is ce8952abbd2f2881096e7133e84f3a6b5ce4b24dca649b151f790684790db2b3; candidate is dd0ffaeacddf4f05753b4babf97132eface2824238a933c10647a786a2587473. Collected artifacts are under build/nfs-movsd-box/; profile function indices were resolved against its build/baseline.wat and build/combined.wat respectively.

The complete compiler differential suite passed remotely, including forward and backward copies, sequential overlap, flag preservation, unaligned seams, noncontiguous backing pages, warmed store-window invalidation and the NFS record scan. See build/movsd-compiler-test.log within that artifact directory: 288 programs compiled, 97 declined, zero guard failures and 22 invalidations. The initial code-write fixture was corrected to enter its installed loop head; the final pass verifies actual compiled entry before the guarded code write.

Fixed-work CPU measurement

movsd-loop.json records seven alternating AB/BA pairs per shape after two warmups per artifact. Every sample processes 500 scans of 2,000 records. Resetting inputs and checking results occur outside timing. Registers, flags, EIP and the complete 128 KB working buffer match between artifacts on every run. Load average was zero at the start and end of this short measurement.

Record shape Baseline median CPU Candidate median CPU Median paired CPU reduction Median paired wall reduction
All zero 41.229 ms 18.203 ms 56.07% 55.43%
Mixed zero/nonzero 35.965 ms 18.986 ms 47.20% 47.21%

This isolates the periodic guest loop and excludes rendering. Process CPU accounting includes any process background activity; warmup and alternating order reduce that concern. The result is a loop improvement, not an FPS forecast.

Game profile coverage

Two 20-second diagnostic windows per renderer/artifact are saved in build/box-profile-{before,after}/{glide,d3d}/. The targeted addresses are 0x4c5f28, 0x4c5f34 and 0x4c5f3a.

Counter, aggregated per frame Glide before → after D3D before → after
Targeted residual block entries 24,488.37 → 6.93 24,466.45 → 7.03
All residual block entries 176,976 → 146,023 165,191 → 131,900
Recorded handler invocations 1,155,711 → 1,041,310 1,124,771 → 994,202

The targeted blocks fall from 13.66–14.02% to 0.00472–0.00477% of Glide's recorded entries, and from 14.62–15.00% to 0.00532–0.00535% of D3D's. These counters exclude work moved inside the tier; they demonstrate coverage, not equivalent reductions in total guest instructions.

Worker profiles remain dominated by renderer waits: Glide's synchronous RPC wrapper accounts for 65.19% → 66.28% of weighted intervals; D3D's readPixels accounts for 55.03% → 56.72%. Excluding explicit histogram helpers, weighted WASM intervals per frame decrease from 16.46 to 15.45 ms for Glide and 15.25 to 14.25 ms for D3D. These are sampled intervals, not OS CPU measurements. Histogram overhead, scene differences and SwiftShader waits limit attribution. The common $run, x87, micro-op and load/store hotspots remain.

Whole-game measurement limits

Uninstrumented ABBA sessions are in build/box-{before,after}-{1,2}/. Aggregate results do not establish an improvement or a regression:

Renderer Baseline FPS Candidate FPS Change Baseline CPU ms/frame Candidate CPU ms/frame
Glide 16.7066 16.6749 −0.19% 202.048 203.298
D3D 17.8151 17.6136 −1.13% 199.995 201.641

CPU figures sum the browser process tree, including SwiftShader's GPU process; they do not isolate the guest interpreter. Session FPS ranges overlap: Glide baseline 15.528–17.886 versus candidate 15.760–17.590, and D3D 17.632–17.998 versus 16.948–18.280. Paired FPS changes reverse sign (Glide +13.28%, −11.89%; D3D −3.88%, +1.56%). Fourteen of sixteen windows trigger the four-vCPU contention flag. Rendering phase and software-GPU load dominate this comparison. The supported conclusion is improved loop execution and demonstrated tier coverage, with no proven whole-game speedup.

Reproduction on the reserved box, with its existing fixture and Chrome path:

STACK_BENCH_WASM=build/wine-assembly.wasm node test/test-uop-compiler.js
node tools/nfs-movsd-loop-bench.js --baseline=build/baseline.wasm --candidate=build/wine-assembly.wasm --iterations=500 --pairs=7 --warmup=2 --out=movsd-loop.json
node tools/nfs-renderer-bench.js --wasm=build/baseline.wasm --cases=glide,d3d --seconds=20 --samples=2 --swiftshader --no-sandbox --guest-profile --profile --out=build/box-profile-before
node tools/nfs-renderer-bench.js --wasm=build/wine-assembly.wasm --cases=glide,d3d --seconds=20 --samples=2 --swiftshader --no-sandbox --guest-profile --profile --out=build/box-profile-after

Use the same browser commands without --guest-profile --profile for uninstrumented game measurements. A display/Xvfb and CHROME pointing to the remote executable are required. No laptop benchmark was used for this change.