Dual-mode crash sweep, 2026-10-06

Every game has to work under the cooperative scheduler (run.js --no-threads, the default) and with real guest threads (run.js --threads, the headless twin of the browser's Threads mode). tools/crash-sweep.js now runs both and compares them:

node tools/crash-sweep.js --apps=a,b,... --modes=coop,threads --seconds=15 \
  --stuck-after=0 --jobs=1 --jsonl=dual.jsonl
node tools/crash-sweep.js --compare --jsonl=dual.jsonl [--md] [--gate]

Each run records its signature (ok / unimpl / trap / exit / stuck / ...), the final screen (content / blank, from the exit PNG) and the waveOut PCM (sound / silent / none). --compare lists every app whose two modes disagree, crash first, then early stop, then picture, then sound; --gate exits 1 when any do.

Two things to know before reading a result:

Result: curated top 38, build 61686c8e + cedede26

scratch/runs/20261006T1325Z-dual-mode-top38/dual.jsonl (coop vs threads, 15 s each, one run at a time, default memory). Cells are signature / frame / audio.

app coop threads
age_of_wonders_demo ok / content / none ok / content / none
alien_shooter ok / content / none ok / content / none
arcanum_demo ok / content / none ok / content / none
blood2_demo ok / content / none ok / content / none
caesar3_demo ok / content / none ok / content / none
carmageddon2_demo ok / content / none ok / content / none
colin_mcrae_rally_demo ok / content / none ok / content / none
crimsonland ok / content / none ok / content / none
deus_ex_demo ok / content / none ok / content / none
diablo2_demo ok / content / none ok / content / none
diablo_shareware ok / content / none ok / blank / none
die_by_the_sword_demo ok / content / none ok / content / none
halflife_uplink ok / content / sound ok / content / sound
heroes3_demo ok / content / none ok / content / none
hitman_glide_demo ok / blank / none ok / blank / none
hype_glide_demo ok / content / none trap:unreachable / none / none
icewind_dale_demo ok / content / none ok / content / none
jazz2_demo ok / content / none ok / content / none
liquid_war ok / blank / silent ok / blank / silent
moorhuhn ok / content / none ok / content / none
mshearts16 ok / content / none ok / content / none
myth_tfl ok / content / sound ok / content / sound
nfs3_demo ok / content / none ok / content / none
nfs3_glide_demo ok / content / none ok / content / none
pinball ok / content / sound ok / content / sound
quake2_demo ok / content / none ok / content / none
red_alert_95_demo ok / content / none ok / blank / none
simgolf_demo ok / content / none ok / content / none
ski32 ok / content / none ok / content / none
starcraft_shareware ok / content / none ok / content / none
tomb_raider_2_demo ok / content / none ok / content / none
tomb_raider_3_demo ok / content / none ok / content / none
unreal_special_demo ok / content / none ok / content / none
ut348_demo ok / content / none ok / content / none
warcraft3_demo ok / content / none ok / content / none
winamp_mod ok / content / sound ok / content / sound
winboard ok / content / none ok / content / none
zuma_deluxe ok / content / none ok / content / none

Fixed: diablo_shareware threw ThreadManager.runSlice is the cooperative backend under --threads (cedede26): run.js called the inline cooperative wait for a wait nested in a synchronous SendMessage, which the worker backend cannot run. host.js already skipped it.

Open divergences:

  1. diablo_shareware, threads: no trap any more, but the menu is Storm's dialog with plain Win32 buttons and the art is never drawn. Storm waits for its MPQ reader thread inside WM_INITDIALOG. Cooperatively the nested wait pumps that thread inline; under workers the wait answers "pending", the dialog procedure yields, and $wnd_send_message abandons it halfway (thread-manager.js's own comment on waitMultipleCooperative). host.js does the same in the browser's Threads mode, so the page should show the same menu. The fix is a blocking nested wait that services worker RPC until the object is signalled -- a worker-backend design change, not a one-liner.
  2. hype_glide_demo, threads: trap:unreachable. MFC's thread at 0xb98b10 (created suspended, then resumed) exits through EIP 0 still holding a critical section within ~1 s; under coop it stays alive. The main thread later reads a zeroed m_hWnd (ShowWindow(0)) and runs into 0x7525.

Not run: 230 registry apps. A boat fork has no test/binaries (30 GB, gitignored, not in the snapshot) and run.js cannot mount app files from a URL, so a full sweep on a boat needs a feed: a static server here, reverse forwarded into the fork, and a tool that fetches each app's registry files before its run.

Follow-up (THREADS-DIVERGENCES-3): none of the three is a product bug

Checked in the page with Threads on (tools/web-input-probe.js --threads, which serves COOP/COEP, so the guest main thread runs in a Worker):

Diablo and hype fail only under run.js --threads, and the reason is the harness, not the emulator: run.js keeps the guest main thread in-process, while the browser's Threads mode runs it in a Worker. A wait the main thread makes inside a synchronous SendMessage (Storm's MPQ reader inside WM_INITDIALOG) can block in a Worker with Atomics.wait while the page services the other threads' RPCs; in-process it cannot block, because the threads it waits on need that same thread for their host imports, so it answers "pending" and the dialog procedure is abandoned. Hype's loader thread calls exit() (MSVCRT doexit, seen through 073ea677's EIP-0 registers) on the CLI only; its root cause was not pinned, but the page runs it correctly.

So --modes=coop,threads on the CLI compares the cooperative scheduler with the CLI's worker backend, which is not quite the browser's. A divergence it finds is worth one page check (web-input-probe --threads) before it is treated as a game bug. Making run.js --threads run the guest main thread in a Worker would close that gap.

Result: full registry, 271 apps, build 7b6a8cc0

Run on this box rather than a boat (the API key in use can boat exec but not ssh/forward, so the reverse-forwarded feed could not be attached):

node tools/crash-sweep.js --all --modes=coop,threads --seconds=15 --stuck-after=0 \
  --jobs=1 --min-free-mb=1024 --no-build --jsonl=dual.jsonl
node tools/crash-sweep.js --compare --jsonl=dual.jsonl --md

542 runs, 14:25-16:38Z, with --min-free-mb holding the next run whenever other sessions pushed available memory under 1 GB. Evidence: scratch/runs/20261006T1425Z-dual-mode-full/ (dual.jsonl, sweep.log, compare.md, and the rerun below).

signature coop threads
ok 249 248
missing-files (registry points at files not on this box) 17 17
exit (installers / apps that quit by themselves) 4 4
timeout (windows_installer_20) 1 1
trap 0 1 (hype_glide_demo, the known CLI-only case above)

The two schedulers fail in exactly the same places. Apart from Hype, every crash, exit and timeout appears in both modes. 9 of the 271 apps disagree, and 8 of those disagree only on the final picture. Re-running those 8 once (rerun.jsonl) cleared ut2004_demo (noise) and reproduced the other seven:

app coop threads
deus_ex_demo, diablo_shareware, dungeons_of_dredmor_release, icy_tower, nfs2_demo content blank
dungeons_of_dredmor, tworld no capture content

None of them crashes, exits or loses audio in either mode. Diablo is the case already checked in the page above (the Threads page renders it; the blank is the CLI's in-process main thread). The other six have not had that page check yet and are the follow-up list, not bugs on their own. The two "no capture vs content" rows go the other way: under coop the CLI found no surface to write at exit.

Fixed (HYPE-THREADS-CLI-TRAP): the CLI starved its own main thread

The one crash divergence above was the harness after all, but not where the follow-up section guessed. run.js --threads gave every worker slice max(--batch-size × --thread-slices, 20000) blocks against the main thread's --batch-size, so a worker retired about 27 times the work main did per batch. The page never runs that shape: host.js hands its main step count straight to runWorkerSlices. Hype's loader thread then reached its sprite-row table (0x757200) while main was still filling the 64K colour table beside it (0x466b4f), followed a NULL row, and MSVCRT's _XcptFilter turned the fault into ExitProcess. --threads-serial still failed, so it was never a parallel race; --thread-batch-size=1000 passed.

Workers now get max(--batch-size, 1000), the browser's rule. A threads-mode re-run of the 44 apps above (scratch/runs/20261006T1715Z-threads-slice-parity/): 40 unchanged, Hype reaches its main menu, dungeons_of_dredmor_release and ut2004_demo now match their coop picture, liquid_war goes blank to content, and nothing gained a crash, an exit or lost audio.

October 7 reconciliation of later evidence

The 271-app table is the dated startup sweep, not an all-gameplay pass. e4145fe2 fixed the Hype CLI worker-slice mismatch; its 44-app rerun above remains the supporting evidence. Later investigation of the picture-only splits is retained in scratch/runs/20261006T1850Z-threads-picture-splits-w6/result.json:

The old feed and CreatePipe prerequisites must not be reimplemented: the feed and full sweep landed, and WinBoard now has real pipe/child support. Browser gameplay, sound correctness beyond the limited sweep signal and rows with missing fixtures remain separate obligations.

A read-only entrypoint audit on October 7 found 16 of the 17 previously missing executable paths present in the current registry checkout. The one absent entrypoint is test/binaries/candidates/baldurs-gate-interactive-demo/installed-extracted/MinimumData/BGDemo.exe. This checks file existence only, not dependencies or launch success; it does not rewrite the dated sweep outcomes. Exact paths, sizes and registry hash: scratch/sweep-reconciliation-20261007/entrypoint-availability.json.