Run
Browser: Open
index.html(at repo root), select an app, click Launch. Live build deployed at https://wine-assembly.berrry.app viatools/deploy-berrry.js --update.CLI:
node test/run.js --exe=path/to/exe [options]— headless execution with auto-build. A registered EXE gets that app's explicit file manifest; an arbitrary bare--exemounts only the executable. Add repeatable/comma-separated--vfs-include='*.dat,plugins/**/*.dll'patterns (relative to the EXE directory) for ad-hoc companion assets, or prefer--app=ID. Key flags:--verbose,--trace,--trace-api,--trace-gdi,--trace-host=fn1,fn2,--no-close,--break=0xADDR,--break-api=Name,--watch=0xADDR,--dump-gdi=DIR,--max-batches=N,--batch-size=N,--tick-ms-per-batch=N--quiet-apiis close to free speed on any API-heavy app, so pass it unless you are reading the API log. Every Win32 call prints an[API]line by default — Diablo emits 96,787 of them by its main menu and 724,015 by the time it reaches gameplay — and that write is blocking I/O on the thread the guest runs on. Measured back to back on the same 1000-batch Diablo command line: 3:53 wall / 25.1s user CPU by default against 1:17 wall / 25.1s user CPU with the flag. Identical CPU, three times the wall clock — the whole difference is the process waiting on stdout, and on a loaded box it is the difference between a run finishing and a run being SIGKILLed at its timeout.--trace-apiand friends are unaffected; this only suppresses the unconditional one-liner.--tick-ms-per-batch=N(default 200) sets how much guest time one batch is worth on the headless batch-driven clock. Reach for it whenever a game's engine steps on aWM_TIMERand the capture shows the clock already expired: at the default 200ms/batch Chip's Challenge burns its entire 100-second level timer in 500 batches and puts up "Ooops! Out of time!" before any--inputlands.--real-ticksis not the fix — a 16-bit app runs 5000 batches in a third of a second of wall clock, so its 110ms timer fires about three times in a whole run and nothing ever moves.--tick-ms-per-batch=5reaches real gameplay.The guest calendar is pinned by default to
1999-06-01T12:00:00Z(it then advances with guest time), so two runs of one command are the same run even when the guest seeds from the time of day — Deus Ex diverged per wall-clock second before this.--wall-clock-ms=Nor an app'swallClockinlib/apps.jspicks another date;--real-calendargives today's date;--real-tickskeeps the real calendar. The time zone is pinned too, to UTC, soGetLocalTimereads 12:00 on a PDT laptop and on a UTC box alike, and agrees withGetTimeZoneInformation's zero bias.--tz=IANA/Namepicks a zone;--real-timezone,--real-calendarand--real-tickskeep the host's. The browser always uses the real date and zone.--batch-sizeis a budget of BLOCKS, not steps, and a block is not a fixed amount of work.run(N)spends one$block_budgetper basic block, and the$stepsquantum is 256 block transfers per visit to the run loop, counted at block ends rather than in every dispatch — so one block can be a 3-instruction loop head or an entire folded sprite row. Measured on Diablo (--batch-stats+--handler-hist): 6.9 ops per block in its Smacker intro against 282 ops per block in its menu, a 41x spread inside one app. Ops per second is flat at ~7-9M across both, so the emulator is not "slower" in the menu at all — batches are simply a meaningless unit of work, and any per-batch number (batches/s, ms/batch) silently changes meaning when the guest's code shape changes. Quote ops, or quote wall time for fixed work.This is also why time-paced content stalls. The headless clock is
batch * TICK_MS_PER_BATCH, so guest time advances per batch while work per batch varies 41x — and it varies the wrong way: an app sitting in a tight polling loop retires tiny blocks, so it is granted the fewest ops per guest-second exactly when it is waiting for time to pass. At the defaults that intro is being run on a machine doing ~35,000 ops per guest-second against ~100M for a real Pentium. Anything pacing itself offtimeGetTime/GetTickCount— an intro video, an animated menu, a timed fade — then renders a fraction of a frame per guest second and looks stalled, which reads convincingly as a broken decoder. Diablo's Blizzard North logo is byte-identical for 23,000 batches for exactly this reason (--trace-schedhas the main thread inside smackw32 working the whole time, and an API census counts 10,068timeGetTimecalls against 62 surface Lock/Unlock pairs) and plays straight through at--batch-size=200000. When a timed animation appears frozen, raise--batch-sizebefore you go looking for a decoder bug.CLI, by app id:
node test/run.js --app=sol— takes the exe, its DLLs, its data files and its command line fromlib/apps.js, the same registry the desktop icons read, so the CLI mounts exactly what the browser mounts.--app=with an unknown id prints the full id list. An explicit--exe/--argsoverrides the registry.PNG render:
node test/run.js --exe=path/to/exe --png=output.pngRemote/headless control: add
--control-stdinfor newline-delimited JSON commands on stdin. Add--frozento park before the first batch and advance only through{"action":"step","n":N}commands; snapshots and PNG captures remain available while parked. Use--max-seconds=Nas the in-process wall clock guard for controlled runs, so no externaltimeout -s KILLwrapper is needed.Real threads (CLI): guest threads use the cooperative scheduler by default;
--no-threadsselects that mode explicitly for reproducible A/B commands.--threadsinstead gives every guest thread its own nodeworker_threadover the one shared memory — the headless twin of the browser's Worker backend. The guest's main thread stays in-process, so this is not identical to the browser, but it is the only automated coverage of the shared-memory locks, the per-thread RPC blocks and the wait-completion path. The two mode flags are mutually exclusive. Add--threads-serialto run one guest thread at a time (splits "race" from "wrong per-thread state" in one run) and--thread-batch-size=Nto set the steps per worker slice (default--batch-size, min 1000 — the browser's own rule: host.js hands its main step count to every worker, so main and workers retire work at parity. The old default,--batch-size×--thread-sliceswith a 20000 floor, let workers outrun main ~27× and manufactured a startup race in Hype that no real machine and no page run has).--trace-threadalso reports why a thread got no slice this batch. Runs print acritical sections:line, aheld critical sections at exitlist (each section's owner,LockCountandRecursionCount) and a per-threadcsPark/steal,csBadLeave,waitingOnCS=… heldBy=…column whenever a guest thread blocked on aCRITICAL_SECTION. Read the held list first on any threads hang: a climbingcsParksays somebody is waiting, but only the owner/counter state says why — a section whose counters sit below the valuesInitializeCriticalSectionwrites (lock=-1 recursion=0) can never be released by the guest again, and a nonzerocsBadLeavenames the thread that put it there. A section owned by a thread that has already exited is released automatically at exit. Sections are never taken by force by default, because doing so rewritesLockCount/RecursionCountunder the owner and the guest CRT reads those —--cs-steal-after=Nturns it back on for the one measurement it is good for, telling "waiting forever" apart from "took it". Seedocs/design-real-threads.mdfor what worker mode does and does not get right yet.Child processes (CLI): a guest CreateProcess with redirected std handles always runs as a real child emulator (pipes over the vlan wire,
docs/design-anonymous-pipes.md). Any other CreateProcess only does with--spawn-processes(orspawnProcesses: trueon the app inlib/apps.js); without it the call reports success and no child runs, which installer-extraction tests rely on. A spawned child gets a snapshot of the guest's C:\ and registry, finds its exe on that C:\ (current dir, exe dir, SYSTEM, WINDOWS), and what it leaves is merged back before the parent's wait on it returns — soinstmsi -> msiinst -> msiexecinstalls Windows Installer in one run.--pipe-child-args="..."passes flags to every child down the tree (--input=B:dlg-cmd:IDanswers a grandchild's dialog). Each child is another 512 MB guest: a three-level chain is boat-sized. The browser does the same with visible in-page instances (host.jsprocess_spawn,VirtualFS.mergeChildFrom).Prefer the CLI's own
--max-seconds=Nguard. Controlled test harnesses should launchrun.jsdirectly and ask it to stop or quit over stdio; do not wrap normal runs intimeout -s KILL. The internal guard is checked between batches, preserves shutdown diagnostics and lets the process clean up. An external SIGKILL deadline is only a last-resort harness guard for an untrusted regression that might never return from one WASM batch.
Superseded shared-worktree defaults (preserved baseline)
The paragraph below preserves the pre-reconciliation shared text verbatim. The current defaults and instructions are in the Run section above.
- Real threads (CLI): guest threads use the cooperative scheduler by default;
--no-threadsselects that mode explicitly for reproducible A/B commands.--threadsinstead gives every guest thread its own nodeworker_threadover the one shared memory — the headless twin of the browser's Worker backend. The guest's main thread stays in-process, so this is not identical to the browser, but it is the only automated coverage of the shared-memory locks, the per-thread RPC blocks and the wait-completion path. The two mode flags are mutually exclusive. Add--threads-serialto run one guest thread at a time (splits "race" from "wrong per-thread state" in one run) and--thread-batch-size=Nto set the steps per worker slice (default--batch-size×--thread-slices, min 20000 — not--batch-size, which is tuned for the opposite constraint).--trace-threadalso reports why a thread got no slice this batch. Runs print acritical sections:line, aheld critical sections at exitlist (each section's owner,LockCountandRecursionCount) and a per-threadcsPark/steal,csBadLeave,waitingOnCS=… heldBy=…column whenever a guest thread blocked on aCRITICAL_SECTION. Read the held list first on any threads hang: a climbingcsParksays somebody is waiting, but only the owner/counter state says why — a section whose counters sit below the valuesInitializeCriticalSectionwrites (lock=-1 recursion=0) can never be released by the guest again, and a nonzerocsBadLeavenames the thread that put it there. A section owned by a thread that has already exited is released automatically at exit. Sections are never taken by force by default, because doing so rewritesLockCount/RecursionCountunder the owner and the guest CRT reads those —--cs-steal-after=Nturns it back on for the one measurement it is good for, telling "waiting forever" apart from "took it". Seedocs/design-real-threads.mdfor what worker mode does and does not get right yet.