Selective dirty tracking: real-game measurements

Selective page tracking is now implemented in src/03a-page-watch.wat and lib/d3dim-gpu.js. Normal texture-cache hits poll page generations without copying or comparing texture bytes. The historical baseline and preliminary cost experiment below are retained; neither establishes a causal FPS speedup.

Production tracker

Only registered texture, palette and render-target backing pages have live observers. A shared 4 KB root indexes lazily allocated 16 KB tables, each covering 4 MB of backing memory. Each 4 KB page has an atomic reference count and 64-bit generation. Guest aliases and different WASM instances see the same generation. Each cache consumer retains its own last-seen versions; one consumer never clears another consumer's dirty state.

guest store / REP / bulk copy ---- guest page translation ---+
native DDraw / GDI / host copy ---- backing address ---------+--> watched?
                                                               | yes
                                                        atomic generation++
                                                               |
texture cache: identity + page generations <--------------------+
  unchanged -> reuse GPU texture
  changed   -> decode/upload; remember observed generations

The CPU accessors, optimized copy loops, native drawing boundaries, filesystem copies, canonical JS surface writes and GPU readbacks notify the tracker. Acquisition invalidates existing uop windows via the shared epoch; new store windows and the widened 64 KB reguard refuse watched pages. Read windows remain eligible. Surface release, backing retirement, format/palette changes and renderer shutdown are covered. Failed acquisitions roll back their references; overflow or a debug audit miss disables the optimization process-wide and restores byte comparisons. Versions are notifications, not synchronization: existing resource ownership and GPU fences still apply.

Framebuffer uploads use page changes to bound the candidate rows, then trim identical edge rows within that range. This is intentional: Unlock can mark a whole surface, and a DIB swap can retain the same pixels. Neither should force a full upload for a small HUD change. Texture hits do not use this scan.

set_page_watch_audit(1) additionally compares unchanged pages with byte shadows. --audit in the gameplay tool enables it before guest startup. Both NFS III and GTA2 passed real Worker/hardware gameplay audits with zero misses (554,818 and 280,744 total triangles respectively). The regression test deliberately omits a notification and verifies that the audit detects it and that subsequent unnotified changes remain visible through the fallback. The final GTA2 build was audited again after the DIB-swap fix: zero misses through 208,636 triangles, with 562 MB of debug shadow comparisons (gta2_demo-page-watch-audit-final/).

Normal NFS III measurement: 357 flips over two 15-second windows, zero texture byte checks, 663,669 page-generation checks, and 0.024/0.025 ms of framebuffer comparisons per frame. Observed FPS was 11.67/12.13, versus the earlier baseline's 8.09/8.97. These are not controlled A/B timings: system load was 35–37 during the new run. The eliminated texture scans are established by work counters; the FPS gain still needs a quiet paired measurement.

Final GTA2 normal measurement: zero texture byte checks across 705 flips, 147,719 generation checks, and 20.54/26.46 observed FPS. Partial framebuffer uploads averaged 79/160 rows per frame, rather than all 480 rows. Framebuffer comparison time was 1.78/0.97 ms per frame; unlike NFS, GTA2 swaps DIB pointers, so its framebuffer still needs comparison across different backing ranges. Load rose from 13.7 to 27.1; these FPS numbers also are not a controlled A/B.

Artifacts: build/d3dim-gameplay-perf/nfs3_demo-page-watch-normal-1/, nfs3_demo-page-watch-audit-1/, gta2_demo-page-watch-normal-2/, and gta2_demo-page-watch-audit-1/. The first GTA2 normal run revealed full uploads across DIB swaps; that issue was fixed by comparing the target's pixel shadow across compatible layouts. Do not use gta2_demo-page-watch-normal-1/ as the final implementation result.

Validation: full build gates; test-page-watch.js (real VM stores, split pages, aliases, multiple instances, uop windows, native drawing/decoding, readback aliases, DIB recycling, palettes, filesystem copies, JS surface writes, row trimming, release, overflow and audit fallback); test-code-write-granularity.js, test-gdi-surface.js, test-d3dim-gpu-depthless-z.js, test-d3dim-texture-wrap.js, and test-d3dim-gpu-edge-web.js.

node tools/bench-d3dim-gameplay.js --app=nfs3_demo --audit --label=audit --seconds=15 --windows=2
node tools/bench-d3dim-gameplay.js --app=gta2_demo --label=tracked --seconds=15 --windows=2

MechWarrior III menu: remote regression check (2026-09-28)

Compared the complete tracker commit ff6dc0f4 against its parent 76c548ed on reserved box 8, in isolated ~/mw3-watch-ab. Each arm used its own compiled WASM, GPU executor and generated region map; common host files came from the tracker snapshot. Later main-branch optimizations were excluded from both arms. Headful Chrome 151, real guest Workers, 1024x768 viewport, four remote vCPUs (Ryzen 9950X reported), no concurrent benchmark. The renderer reported ANGLE / Intel UHD 620 / Mesa in every measured window. Although --swiftshader was requested, that is not what the runtime reported; do not describe this as a verified SwiftShader run or extrapolate these rates to Safari.

Route: ?debug&threads&d3dim-gpu, launch mw3, wait for 120 DirectDraw presents, settle for 60 seconds, then two 15-second windows. Captures confirm the animated main menu. dx_trace kind 5 counts guest presents, not unique displayed frames. All five runs completed both windows with no page errors or GPU errors.

Run order Tracker Window 1 presents/s Window 2 presents/s
before-a off 23.92 42.72
before-b off 23.98 43.07
after-a on 23.24 32.46
after-b on 24.26 42.06
before-c off 24.20 32.33

The two-window run means average 31.70 without tracking and 30.50 with it (-3.8%), but the baseline's own repeated-run spread is 15.7%. The slow second window occurs in both arms. This does not establish a regression, nor prove zero overhead. Fixed wall windows cover a variable amount of the animated menu: second-window GPU draws range from 340 to 506. A tighter attribution needs a matched animation phase or longer route plus CPU profiling.

There were zero texture uploads and zero texture-byte comparisons in every measured window. GPU framebuffer comparisons cost only 4.08-5.14 ms across each entire second window (zero in the first), so removing texture scans cannot substantially speed up this menu route. Watched-page store barriers remain a possible CPU cost, not a demonstrated cause from this experiment.

Artifacts: build/mw3-watch-ab-results/mw3-{before-a,before-b,after-a,after-b,before-c}/ contains manifests, checksums, counters and start/end captures. The remote directory also retains pinned runtime artifacts and both build inputs. Earlier before-probe / before-1 failed launch, and before-2 overlapped startup; none are included in the table.

Example baseline command, from the isolated remote directory:

DISPLAY=:0 CHROME=/usr/bin/google-chrome node tools/bench-d3dim-gameplay.js \
  --app=mw3 --label=before-a --wasm=build/watch-ab/before.wasm \
  --gpu-source=before/lib/d3dim-gpu.js --region-map=before/lib/region-map.generated.js \
  --swiftshader --seconds=15 --windows=2

Individual present intervals (2026-09-29)

Interpretation corrected by the investigation below: the long gaps are scripted ATTRACT waits, not rendering stalls. The idle route crossed animation phases, and the original benchmark also queried the GPU renderer name on every stats snapshot. Preserve these raw measurements as historical evidence, not as clean shipping-runtime FPS or a precise dirty-tracking cost estimate.

Repeated the pinned comparison with --frame-times --seconds=60 --windows=1 in A/B/B/A order. Each launch again settled for 60 seconds. The worker records performance.now() at each kind-5 present into a preallocated 120,000-entry Float64Array; it is read only after measurement. Screenshots and CPU profiles are outside the timed interval (profiles disabled). No overflow, page errors or GPU errors. These are guest-present intervals, not unique browser scanouts.

Run order Tracker Presents/s Median ms p95 ms p99 ms Longest gap ms
before-interval1 off 22.88 42.92 80.02 83.80 5000.15
after-interval1 on 21.61 46.04 81.16 84.54 5002.52
after-interval2 on 22.36 44.26 80.16 84.81 4999.97
before-interval2 off 23.11 43.07 78.96 83.04 4999.42

Pooled: 23.00 presents/s before, 21.99 after (-4.38%), median 43.00/45.30 ms, p95 79.57/80.54 ms, p99 83.50/84.75 ms. The control repeat spread is about 1%, the candidate spread about 3.4%. Both candidate averages are lower in this small repeated sample, suggesting modest overhead; do not treat -4.38% as a precise universal cost. The final 30 seconds average 13.17/13.08 presents/s.

Every run contains exactly two gaps over 100 ms: one approximately 1.01 seconds and one approximately 5.00 seconds, followed by a brief burst of presents. The 5-second gap accounts for four complete zero-present one-second bins in each capture. Intervals over 50 ms: 1187/2758 before and 1200/2637 after. The menu pacing is therefore not stable. The large pauses occur without tracking too; whether these are intentional game waits or emulator stalls requires tracing and is not established by timestamps alone. Burst present counts also explain why earlier short-window rates overstated sustained pace.

Raw captures: build/mw3-watch-ab-results/mw3-{before,after}-interval{1,2}/ (window-0-frame-times.json, results.json, start/end PNGs, console). Aggregated results: build/mw3-watch-ab-results/frame-time-summary.json. Interval differences were independently checked against the stored timestamps.

Gap and low-FPS diagnosis (2026-09-29)

Two separate causes were found on the pinned ff6dc0f4 snapshot:

The ~1s and ~5s gaps are guest-script waits. reader.zbd at file offset 0x3487a contains the ATTRACT script: PLAYAVI intro.avi, WAIT 1.0, LOADIMAGE mech3splash, WAIT 5.0, then FADEOUT. The floats are stored as 0x3f800000 and 0x40a00000 following their WAIT tokens. The EXE parser at 0x562dd0 compares the command against the WAIT string at 0x5bc2c4 and constructs the object with vtable 0x599cd8. Its update method 0x563c60 compares elapsed time against start time plus duration; it does not draw. Samples inside the gap hit that method (0x563c73), its animation dispatcher 0x5633f0, input polling and the main message/timer loop. No Sleep or blocking wait yields were recorded. The host page was ~98% idle during the gap. The benchmark's unattended warm-up allowed this attract sequence, so calling the two gaps emulator stalls was incorrect.

Slow active animation is a code-cache retirement storm. The RGB565 alpha loop fold publishes 0x528064..0x528111 (173 bytes). Execution also resumes at its internal stores 0x5280f4 and 0x52807b. Publishing the full fold retires those entries; recompiling an interior entry retires the fold. The existing --trace-code-writes instrumentation, captured in a bounded local buffer, recorded 1,000 retirements: 442 fold-to-0x5280f4, 440 reverse, and 118 involving 0x52807b / 0x5280f7. Every record has in_code_write=0. This is overlapping compiled entry ownership, not changing guest code or texture byte comparisons. The micro-op tier can resume at internal instructions; the native whole-loop matcher does not apply the fuse_stop protection used by smaller folds.

Over a 29-second tail sample: 47,595,892 decoded blocks, 46,995,406 retirements, zero directory/index evictions, and one full cache clear. With the benchmark driver-query issue fixed, sampled worker self time in the last 30 seconds was 22.84% page_publish, 12.74% decode_block, 11.82% alpha-loop recognition, and 11.47% uop_fast. Main-thread canvas conversion/upload was small. The function names were verified against a compiler-emitted name section whose non-custom WASM sections exactly match the measured runtime; no current-main function-index guesses were used.

Benchmark correction: instrumentGpu().snapshot() previously repeated getParameter(UNMASKED_RENDERER_WEBGL) on every stats snapshot. It accounted for 7.6% of sampled worker wall time in the first profile. The tool now caches the name once per GL context. That overhead was in the benchmark, not the shipping renderer. Corrected profiled runs still show the retirement storm; their throughput is not an unprofiled A/B replacement for the earlier tables.

Artifacts under build/mw3-watch-ab-results/: mw3-after-gap-debug1 (initial CPU profile), mw3-after-gap-debug2 (cached driver query, periodic guest PCs/cache counters and CPU profiles), mw3-after-cache-trace (bounded retirement records), and mw3-after-gap-callers (targeted 9-second gap sample). In the diagnostic profiles, worker-0 is the guest and worker-1 is the browser page.

Alpha-fold fix and controlled A/B (2026-09-29)

$try_emit_rgb565_alpha_run now checks the existing $fuse_stop predicate while validating its 173-byte candidate. An independently compiled interior entry declines the whole-loop fold; ordinary decoding preserves that suffix. A cold loop still receives the native fold. No guest address special case, thread-mode change, or global optimization disable is involved.

test/test-fused-entry-overlap.js embeds the real loop and alternates its head with both internal store entries (offsets 0x17 and 0x90). It checks pixel output, cold native-fold execution, and zero steady-state retirements or recompiles. The test fails on the old matcher and passes with the guard. The isolated current-main build and code-write granularity regression also pass.

The performance comparison pins ff6dc0f4 against that exact snapshot plus only this guard (1,603,908 vs 1,603,925 WASM bytes), using matching GPU JS and region map. Thus unrelated current-main changes cannot explain the result. Same quiet box 8, Chrome 151, headful real Workers, no CPU profiler; cached GPU-name query in both arms. The renderer reports ANGLE / Intel UHD 620 / Mesa 23.2.1 despite the requested SwiftShader launch flags. These are guest DirectDraw presents, not monitor refreshes or measured combat FPS.

ABBA run Presents/s Median interval p95 interval Block decodes Retirements
Control 1 29.06 27.23 ms 72.31 ms 36,895,507 36,395,126
Fixed 1 76.82 11.04 ms 18.18 ms 4,000,314 3,796,783
Fixed 2 73.55 11.54 ms 18.20 ms 3,974,221 3,771,630
Control 2 31.39 25.84 ms 70.61 ms 36,667,625 36,167,231

Each window lasts 30 seconds after a 60-second warmup. Mean throughput rises from 30.22 to 75.19 presents/s (2.49x); retirements fall 89.6% despite more frames. The control repeat spread is 7.7%, versus a 148.8% improvement. Remaining retirements are not zero and are not explained by this experiment. All four captures still contain the scripted one/five-second waits. Small mouse movements every two seconds did not prevent them; the experimental keepalive option was removed. Do not describe the whole window as stable FPS. Start/end screenshots show the same animated menu, with no rendering errors.

A separate control/fixed pair with --warmup-ms=95000 captures 30 seconds of the later active animation with no scripted gaps in either arm:

Active-animation run Presents/s Median p95 Worst Intervals >50 ms Retirements
Control 14.59 70.49 ms 74.96 ms 80.00 ms 437/437 47,686,768
Fixed 78.74 13.42 ms 18.04 ms 31.53 ms 0/2,362 2,182,344

That is 5.40x throughput and 95.4% fewer retirements in this phase. Per-second counts span 13–19 before and 60–151 after: animation work still varies, so this is not constant FPS. This phase-specific pair is one repeat per arm; the preceding ABBA establishes the repeated improvement across the mixed sequence. Screenshots still show the animated menu, and both runs report no browser/rendering errors. Artifacts: mw3-after-active, mw3-fixed-active.

Artifacts: build/mw3-watch-ab-results/mw3-after-fixcontrol{1,2} and mw3-fixed-fix{1,2} hold screenshots, manifests, raw frame timestamps and start/end cache counters. Use --frame-times --seconds=30 --windows=1 with tools/bench-d3dim-gameplay.js --app=mw3, pinning --wasm, --gpu-source and --region-map for each arm. --trace-yields adds guest PCs/cache samples; --trace-cache adds bounded retirement records. Both require --frame-times. Use unprofiled captures for throughput and --profile for attribution.

Follow-up: color-key fold overlap

The alpha-only retirement trace identifies another conflicting pair: the 19-byte color-key row 0x528268..0x52827b and store entry 0x528271. Of 995 captured retirements, 988 alternate between this pair; all have in_code_write=0. The candidate applies $fuse_stop to interior bytes after the exact-body match, preserving cold native folding and existing entries.

Remote box 8, same pinned runtime/JS/map and 95-second warmup, 30-second unprofiled captures in ABBA order:

Run Presents/s p95 interval Worst interval Retirements Decodes
Alpha-only 1 82.74 17.82 ms 34.70 ms 2,379,486 2,510,921
Both guards 1 74.54 18.77 ms 33.89 ms 0 0
Both guards 2 84.08 17.97 ms 28.09 ms 0 0
Alpha-only 2 83.41 17.88 ms 29.59 ms 1,818,723 1,919,686

The candidate eliminates steady-state cache churn in both captures. FPS is not an established win: arm means are 83.08 versus 79.31 (-4.5%), with a 12.0% candidate repeat spread. The short windows cover different animation phases; the repeated control spread alone is only 0.8%. This does not prove performance neutrality. All four samples have no intervals over 50 ms and no reported browser/rendering errors. Artifacts under build/mw3-watch-ab-results/: mw3-fixed-keycontrol{1,2} and mw3-keyfixed-key{1,2}.

The longer pair (--seconds=120, same 95-second warmup) resolves the initial tradeoff against the guard-only candidate: 82.98 versus 79.42 presents/s (-4.3%). Their scripted gaps occur at nearly identical offsets (~47.6/52.6 and ~114.7/119.7 seconds), so this is not explained by one capture omitting the waits. Retirements fall 7,556,801 → 4, but p95 rises 17.43 → 18.30 ms. Artifacts: mw3-fixed-long, mw3-keyfixed-long. The guard-only version was not accepted as a performance fix.

The revised implementation keeps the native row useful: a pure exact-body matcher establishes a block boundary before it even on the first encounter, and the micro-op compiler treats that row as an exit to threaded/native code. This prevents the predecessor or a micro-op program from compiling its interior store first. The overlap guard still handles genuine interior entries. The expanded regression verifies predecessor entry reaches the native fold and that micro-op compilation does not replace it.

With that revision, the 120-second capture (mw3-keynative-long) measures 81.78 presents/s against the alpha-only control's 82.98 (-1.45%), with four retirements and five decodes versus 7,556,801 and 7,972,493. p95 is 18.42 ms (control 17.43 ms); the four intervals above 50 ms are the scripted one/five second pauses. Captures remain visually coherent and browser/rendering errors are zero. This establishes removal of sustained cache churn, not an FPS improvement or proof of identical performance. This final revision has one long run; the guard-only ABBA must not be presented as repeats of it.

A fresh 30-second active-menu confirmation (mw3-keynative-confirm) measures 85.40 presents/s, p95 17.76 ms, maximum 39.27 ms, with zero retirements and zero decodes. No interval exceeds 50 ms. This confirms the final version also retains the stable cache in a fresh launch; it does not turn the mixed short/long windows into a claimed FPS improvement.

Validation: isolated build, expanded entry-overlap regression, H440 pixel/register/flag/overlap equivalence and the micro-op compiler suite pass. Code-write granularity passed with the overlap guard before adding native-row priority; that revision did not alter the write/invalidation implementation. The additional local menu-to-cockpit acceptance did not finish its first cooperative run within its 300-second timeout; its child needed to be stopped. It produced no cockpit capture, so neither combat correctness nor combat performance is established by this follow-up. No causal attribution of that timeout to the candidate has been made.

Further GDI / DirectDraw consumers (not implemented)

  1. lib/host-imports.js _flushGdiSurfacePresentation: combine the existing dirty rectangle with backing-page generations before rgbaRect and putImageData. Skip unchanged uploads; convert only changed row spans. DirectDraw rebinding currently marks the full surface dirty even when its backing is unchanged. Writes without Lock can be observed through the same CPU/native writer notifications, provided presentation is scheduled.
  2. _refreshGdiSurfacePalette currently rebuilds RGB tuples and clears the nearest-colour cache on every flush. Watch palette storage separately and preserve the decoded palette and GdiSurface.rgbaRect lookup table until its generation changes. Palette changes must invalidate indexed pixels even when no pixel page changed.
  3. Repeated source conversion for blits can use source versions as a cache key. Skipping the actual destination write requires additional proof: destination versions, clipping, ROP, palette, colour key and geometry.

Start with displayed surface and palette pages only. Each consumer must own its observed generations, and rebind/release must handle Flip's backing swaps, layout changes and memory reuse. Preserve existing synchronization and a fallback when tracking is unavailable. The present CPU implementation refuses fast uop store windows for watched pages, so broader registrations need their own A/B before enabling them by default.

Method

The sections below record the pre-implementation measurements. The historical bench-d3dim-watch-cost.js intentionally refuses the production source tree, where synthetic hooks would double-count the write cost.

tools/bench-d3dim-gameplay.js launches the installed NFS III or GTA2 demo in visible Chromium with real Worker threads and the default WebGL executor. It refuses cooperative fallback. It counts actual DirectDraw Flip calls, not browser animation callbacks, and captures the start and end of every measurement window. Both titles reach live gameplay; NFS is at the starting position with an advancing race timer, and GTA2 is in the Wild Demo playfield with active pedestrians and tutorial dialogue. No synthetic drawing workload.

Server-only instrumentation times texture and framebuffer byte comparisons separately. Existing draw, upload, readback and error counters remain intact. The byte-volume counter includes successful comparisons only; early-mismatch reads are excluded. Per-comparison timer calls add some instrumentation cost. The recorded renderer is ANGLE Metal Renderer: Apple M1, not SwiftShader.

WASM and the D3DIM JavaScript source are pinned in memory during each run and copied into its artifact directory with hashes in results.json. Other application assets still come from the shared checkout. These are local Chrome 153 measurements, not Safari measurements or universal FPS claims.

node tools/bench-d3dim-gameplay.js --app=nfs3_demo --seconds=15 --windows=2 --label=baseline
node tools/bench-d3dim-gameplay.js --app=gta2_demo --seconds=15 --windows=2 --label=baseline

Existing comparison cost

Two approximately 15-second windows per game, Apple M1, threads on:

Game Guest Flip FPS Texture checks, ms/frame Framebuffer checks, ms/frame
NFS III 8.09 / 8.97 18.35 / 17.74 1.39 / 1.25
GTA2 20.93 / 16.88 1.78 / 1.87 1.63 / 1.29

NFS checked 39,100 and 46,062 textures in those windows, rereading 0.95 and 1.10 GB of unchanged texture bytes. Only one texture upload occurred across both windows (257 frames). GTA2 performed 39,295 and 35,110 checks, with nine and zero changed-texture uploads respectively. GPU executor errors were zero in both games. Screenshots confirm race/HUD and Wild Demo gameplay.

The comparisons are a meaningful NFS cost, but removing roughly 19 ms from a 111–124 ms frame cannot by itself make the game reach 60 FPS. GTA2 has much less texture-comparison work. These timings are the measured cost of the existing work, not savings already achieved by dirty tracking.

Artifacts: build/d3dim-gameplay-perf/nfs3_demo-baseline-2/ and build/d3dim-gameplay-perf/gta2_demo-baseline-1/. The failed NFS baseline-1 run is excluded: its instrumentation response lacked Worker isolation headers, causing a cooperative fallback. The harness now preserves those headers and fails immediately if Worker startup falls back.

Experimental write-watch cost

tools/bench-d3dim-watch-cost.js compiles two isolated artifacts from one source snapshot, without editing production WAT or the canonical build. Both allocate the same heap-owned, 2 MB backing-page table and register the actual textures and render targets seen during gameplay. Only those pages are watched. A scalar probe checks the page flag and changes the generation only for watched backing pages. The candidate adds:

Registration bumps the existing uop-window epoch. Two control/candidate smoke checks confirm a crossing DWORD preserves its value and marks both watched backing pages, while an ordinary page remains untracked.

This is a cost probe, not a complete cache-invalidation implementation. Its table pointer is per-instance; secondary workers do not share the main worker's registrations. Native and host direct writes, all bypasses, reference counts, retirement/reuse, generation wrap and cache consumers remain missing. No scan is skipped. Those limitations prevent interpreting the result as the cost or correctness of the finished design.

node tools/bench-d3dim-watch-cost.js
node tools/bench-d3dim-gameplay.js --watch-cost-matrix --seconds=12

For each game the matrix uses control, candidate, candidate, control, with two 12-second gameplay windows per launch. All raw counters, build hashes, GPU identity, machine load and screenshots are saved under build/d3dim-gameplay-perf/. The shared machine has substantial variable background load; repeated controls are necessary to interpret small changes.

Attempt on the shared laptop: inconclusive

NFS III arm Window FPS One-minute load, before → after
Control 7.92 / 7.14 6.0 → 35.6
Candidate, first launch 6.36 / 6.64 47.2 → 83.4
Same candidate, second launch 5.02 / 5.21 77.0 → 75.3

The machine reached a one-minute load of 104.8 between these samples. The identical candidate's second launch was itself substantially slower, and NFS also randomizes weather/scene setup between launches. Do not attribute these FPS differences to the watch checks. There is no defensible overhead percentage from this attempt. The matrix was stopped before the final NFS control and GTA2 candidate/control runs; earlier GTA2 results above measure only the existing comparison cost.

All three NFS runs had zero GPU executor errors and approximately 700 watched resource ranges, and their captures show an advancing cockpit race scene. This is functional evidence that the cost probe executes on real watched resources, not proof of complete write tracking or pixel parity across the randomized launches. Artifacts are the three nfs3_demo-watch-* directories.

The matrix now checks load before each launch and stops above twice the logical CPU count (overridable with --max-load). This catches gross contention; it is not a statistical noise guarantee. New matrices use unique output directories, and single runs refuse to overwrite existing results. The remaining performance work is a quiet repeated A/B, followed by a net before/after comparison once write coverage and cache invalidation are complete.