MW3 mixed integer/x87 island experiment

2026-09-30. The isolated prototype speeds up the authentic projection kernel on both ARM64 and x86-64. The real-game ABBA does not establish an FPS improvement, so the change is not enabled in production.

Change and correctness boundary

tools/bench-mw3-mixed-build.js reproduces frozen 837f0a74 byte-for-byte, then transforms only 07b-loop-match.wat in the compiler's virtual inputs. It extends H451 islands across ADD reg,imm32 (H3), INC reg (H64), and DEC reg (H65), retaining the existing integer flag helpers and x87 semantics.

Projection before:
  FLD → FDIV → ADD → ADD → DEC → [remaining x87 island] → JNZ

Projection candidate:
  [FLD → FDIV → ADD → ADD → DEC → remaining x87 operations] → JNZ
   one island entry; FP state retained across integer operations

The matcher still starts/ends on an x87 operation, limits spans to 255 records, and stops at unsupported operations, branch terminators and later EA_TEMP consumers. It does not absorb the loop backedge or change arithmetic order. Integer operations update the actual guest register file and lazy flags.

Existing block-executor eligibility checks reject these integer records inside an x87 span. This conservatively prevents running a mixed island with stale registers cached by a caller. It can also forfeit a block-executor opportunity; that is a performance limitation of this prototype, not something to bypass.

The first candidate tested three handler IDs at every operation. Native V8 code exposed the boolean-combining overhead. The final evaluator uses one unsigned fn < 66 check: the matcher admits only H3/H64/H65 or H188..190. Admission remains an explicit whitelist. No new opcode or export is required.

Final artifacts under build/lazy-games/mw3-mixed3/:

No shared runtime source or canonical build was changed. The earlier mw3-mixed2/ candidate is retained for disassembly comparison and is superseded.

Validation and fixed-work timing

tools/bench-mw3-mixed-parity.js reuses the existing x87 differential fuzzer, injects integer bridges, and checks all GPRs in addition to FP file/shadows, TOP/tags, control/status, scratch memory, lazy flags and exception traces. Final result: 1,000 sequences, 26,138 operations; fast == generic == unfused. Every sequence exercised an island. Random FP states include NaNs, infinities, signed zeros, denormals, empty tags and differing rounding modes.

tools/bench-mw3-mixed-kernel.js copies the actual EXE bytes 0x4fd394..0x4fd3e0, including its original backedge and return. It compares complete output hashes, relevant GPRs/flags and exposed FPU state for ordinary vertices and special float inputs. Both architectures passed.

Timing uses 500,000 vertices per call, twelve alternating warmups, then ten ABBA/BAAB groups (twenty samples per arm). Input setup, output checks and diagnostic census are outside timing. These are Node guest-thread CPU times, not game frame times and not timings from the separate d8/Chrome code capture.

Machine Baseline median Candidate median Reduction Same-artifact control
Laptop ARM64, Node 24.21.0 59.987 ms 49.460 ms 17.55% 0.11% difference
Remote Ryzen 9 9950X, Node 24.18.1 32.992 ms 30.991 ms 6.07% 0.00% difference

Node V8 versions were respectively 13.6.233.17-node.53 and 13.6.233.17-node.50. The x86 control's two medians were 34.001 ms; do not compare absolute times from separate runs as a frequency-independent number. The earlier shorter pilot is not directly comparable to this warmed, longer test and cannot isolate the range-check revision's performance effect.

The post-timing census proves the intended work was absorbed:

Dispatched handler Baseline Candidate
ADD immediate 1,000,000 0
DEC register 500,000 0
Scalar x87 absolute memory 500,000 0
Scalar x87 base memory 500,000 0
x87 island 500,000 500,000

Raw timings, parity state and census live in kernel-arm64.json, kernel-x64-final.json, and their *-control.json files. parity.log records the randomized differential test.

Actual V8 native code on both architectures

The existing tools/wasm-native.js --engine=v8 method was already available: --perf-prof emits code bytes even when V8's text disassembler is compiled out. tools/bench-mw3-mixed-native.js follows that method but retains raw jitdump files and original code addresses for correct PC-relative destinations.

The final added operation classification is:

# ARM64: fn is in w25
cmp  w25, #0x42
b.cc integer_operation

# x86-64: fn is stored in a stack slot
cmpl $0x42, -0x88(%rbp)
jb   integer_operation

ST0 stays in d2 / xmm0. Integer arithmetic and flag maintenance are inlined in the captured evaluator; the integer path does not publish/reload the whole FP stack. INC/DEC still contain work to preserve carry from the lazy flag representation. Existing floating stores still use the common memory helper and its associated spills. This is still an interpreter with per-op decoding.

x87_island_fast native object Baseline Candidate
ARM64 d8 17,984 bytes 17,952 bytes
x86-64 Chrome 19,840 bytes 20,160 bytes

The ARM64 stack reservation stays at 224 bytes; x86 grows from 224 to 232. Object sizes include padding/tables and are not dynamic instruction counts. ARM64 keeps the candidate cursor in w24 on the ordinary loop path, while the baseline loop reloads it from a stack slot; register allocation changed alongside the intended dispatch saving. That observation is engine/version specific, not a measured explanation of the Node speedup.

Baseline native files are in mw3-mixed2/arm64-base/ and mw3-mixed2/native-x64-base/; final files are in mw3-mixed3/arm64-candidate/ and mw3-mixed3/native-x64-candidate3/. Each has capture.json, functions.json, raw dumps, .bin and .asm files. Chrome shutdown left unrelated final records incomplete; the parser skips those tails and requires all four requested functions to be fully captured.

Real gameplay: no established gain

Dedicated box bx_4r5uzdwv, 4 vCPU / 8 GB, headful Chrome 151, SwiftShader-backed WebGL, worker guest, shipping lazy-sync setting. One sequential ABBA, two 15-second stationary-cockpit samples per launch:

Run Artifact FPS windows Triangles/present
A1 Baseline 21.11 / 21.85 2,229 / 2,233
B1 Candidate 22.64 / 22.17 1,848 / 1,843
B2 Candidate 22.85 / 22.12 1,943 / 1,978
A2 Baseline 24.18 / 23.91 1,842 / 1,842

All runs reached the cockpit with zero reported errors; candidate screenshots were inspected. B2's first window recorded 135 renderer fallbacks; all other windows recorded zero. Source/worker hashes and rendering mode are recorded per run in mixed-final-*/result.json.

Scene work varies substantially, and baseline launch-to-launch variation is larger than any apparent candidate effect. B1 is slower than the similarly sized A2 scene. These results cannot support a whole-game improvement claim. The initial aborted baseline lacked restored MFC42/MSVCRT DLLs and is excluded; both were restored before this ABBA. The dedicated box is confirmed archived.

Decision and next useful experiment

Keep this as an isolated, reproducible prototype. The kernel result establishes that retaining FP state across integer updates can help; it does not establish that this loop accounts for enough frame time to improve MW3 overall.

Before promotion, measure mixed-island coverage/cost in the game and the cost of lost block-executor eligibility. One bounded follow-up is routing mixed spans to a separate evaluator, preserving the existing pure-x87 fast path so it does not pay the additional classification check. That trades code size for avoiding overhead on unrelated islands. It still needs game A/B evidence.

Reproduce locally:

node tools/bench-mw3-mixed-build.js 837f0a74 build/lazy-games/mixed-new build/lazy-games/postmerge-profile3/postmerge.wasm
X87_FUZZ_CASES=1000 node tools/bench-mw3-mixed-parity.js build/lazy-games/mixed-new
MIXED_VERTICES=500000 MIXED_ROUNDS=10 node tools/bench-mw3-mixed-kernel.js build/lazy-games/mixed-new/baseline.wasm build/lazy-games/mixed-new/candidate.wasm
node tools/bench-mw3-mixed-native.js capture-d8 build/lazy-games/mixed-new/candidate.wasm build/lazy-games/mixed-new/candidate.named.wasm build/lazy-games/mixed-new/native
node tools/bench-mw3-mixed-native.js decode build/lazy-games/mixed-new/candidate.wasm build/lazy-games/mixed-new/candidate.named.wasm build/lazy-games/mixed-new/native

Use capture-chrome with DISPLAY/CHROME on the x86 box, then decode with GNU objdump (OBJDUMP override). Capture directories must be fresh.