Tracing (reach for this BEFORE editing source to add console.log)
Ad-hoc console.log / DBG_* env vars rot. Use the built-in flags first; extend them when they fall short.
| Flag | What it prints |
|---|---|
--trace-api (=Name1,Name2) |
Every Win32 API call with args + return; with =NAMES filter, only those APIs. Args/returns are typed via args:[{name,type[,out:true]}] / ret fields in src/api_table.json (LPCSTR, HWND, LPMSG, flags:WS, etc.) — untyped entries fall back to an nargs-sized hex dump (or 6 dwords if nargs is unknown). Args flagged out:true are decoded after the handler runs (e.g. LoadStringA buf=, GetMessageA msg=) on a separate out: line. |
--trace-api-dedup |
Collapse N consecutive identical API trace lines into a (xN) summary. |
--trace-from=N / --trace-to=N |
Confine every trace category to a window of batches. Without it a bug that only appears at batch 100,000 cannot be traced at all: the flags themselves are affordable, but writing their output for the 100,000 batches of boot before the question is blocking I/O on the guest's own thread — the same effect --quiet-api exists for, and on a long route it is the difference between a run finishing and a run hitting its deadline. Setup and the exit summary are never windowed; only per-batch output is. --trace-dx additionally skips its formatters outside the window rather than only dropping their output, because those scan 64KB of DIB per Lock/Unlock/Present. |
--trace-stack[=DEPTH|=Name1,Name2|=Name:DEPTH,...] |
Walk EBP frame chain on each matched API call (default depth 12). =N overrides default depth for all; =Name1,Name2 limits to those APIs; =Name:N sets per-API depth. Add --trace-stack-scan for code built without frame pointers, where the EBP walk prints garbage like frames=[0x0003ffff]: it lists the stack dwords that point just after a call instead (+off:0xVA*R direct, *i indirect). A stale return address in a dead slot matches too, so trust the innermost few and confirm with --count — that is how NFS II's T2 game step was found. |
--trace-gdi |
Every wrapped GDI primitive: CreateBitmap, BitBlt, StretchBlt, FillRect, DrawEdge, DrawText, TextOut, Rectangle, Ellipse, Polygon, MoveTo/LineTo, Arc, SetPixel, SetTextColor, SetBkColor, SetBkMode, SelectObject, DeleteObject, DeleteDC, GetClipBox, LoadBitmap, CreateSolidBrush, GetObject, PatBlt |
--trace-dc |
Every _getDrawTarget resolution: hdc → resolved hwnd, top-level hwnd, canvas ox/oy, canvas size. Logs NO_CANVAS when resolution fails. Use when a draw call fires but nothing appears — shows which surface each DC lands on. |
--trace-ctrl |
Every WAT-native control paint: [ctrl] paint hwnd=0x… Button at 254,422 75x24 vis=1, in dispatch order, with the screen rect and the effective-visibility bit. Reach for this first on "these pixels should not be there". GDI rasterizes inside WAT now, so --trace-gdi sees only surface binds and cannot say who drew what; this can. Two entries for one hwnd at different origins = a control repainted after a move and nothing erased the old rect (children own no surface). vis=0 = a paint that leaked through while the control or an ancestor was hidden — those pixels are a ghost nothing will clean up. |
--trace-input |
Which routing branch in lib/renderer-input.js consumed each mouse event: the candidate window list for a down, the child it resolved to, and every early return that dropped one (DROPPED: outside modal …). Reach for this first on "the click does nothing". A swallowed click makes no API call at all, so --trace-api shows a healthy message pump and nothing else, and the down and up paths gate differently — a control can accept a press and never see the release. This names the branch that ate it. |
--trace-reg |
Every registry op (open/query/create/set/enum/close) with key path, value name, and result ("found"/"not found"/actual data). Use to discover which keys an app probes when storage returns empty. |
--trace-fs |
Every VFS op: CreateFile (with decoded access/creation + handle/FAIL), GetFileAttributes, FindFirstFile, FindNextFile — each with path and hit/miss result. Use to see which files an app looks for but can't find in the VFS. |
--dx-surfaces |
At exit, one line per live DirectDraw surface: slot, size, bpp, pitch, caps flags, DIB address, the palette actually bound to that surface, and a sampled colour count. Reach for this when a DX app "renders nothing": it tells a primary that was never written (nonZero=0) apart from an offscreen texture that has content, and it names the slot the --png= capture picked. |
--gl-census |
At exit, every GL entry point this run actually issued, with a call count, plus whether the WAT fixed-function mirror (src/09a8f-gl-matrix.wat) stayed trusted. This is how "is app X covered" stops being a guess. No GL family latches UNTRUSTED any more — gluPerspective/gluLookAt/gluOrtho2D, then glPushAttrib/glPopAttrib, all grew mirrors, so the list of unmirrored families prints none by construction and the latch is the reading that can still move. It does still move: the attribute stack is capped at GL's required minimum of 16 and a deeper push latches, because a dropped push leaves every later pop restoring the wrong nesting level. A run whose latch is clear licenses a consumer to lower draws from the WAT state for that app, and nothing more. Two independent readings are printed together — what JS saw called, and what the WAT observer noticed — because a disagreement between them is itself the finding. The counters are load-immune, so this is readable on a box far too busy for any timing number. Needs --headless-gl on the CLI. |
--trace-gl[=Name1,Name2] |
One line per GL command as the backend receives it: [gl] b25453 glBlendFunc(0x1, 0x1), each argument as hex with a float reading beside any word that looks like one (GLfloat and GLenum share stack slots), and for a packed draw its vertex count, first vertex and first vertex's RGBA. Honours --trace-from/--trace-to. Reach for it on "the GL app draws nothing" before reading renderer code: it is what showed ptct never calls glViewport (GL's default is the drawable) and that its geometry arrives with alpha 0 under additive GL_ONE, GL_ONE blending. Works on both --gl-renderer=software and --headless-gl. |
--trace-net |
Every vln/1 frame on the virtual LAN wire, decoded: -> SYN 10.0.0.2:49152 -> 10.0.0.1:8035. The room is 10.0.0.0/24 and addresses are seats — the host is always 10.0.0.1, which inet_addr also accepts written 10.1. Pair with --vlan-ip=A.B.C.D (this process's room address) and --vlan-wire (join the segment offered by the parent process over child IPC). |
--trace-host=fn1,fn2 |
Generic wrap of any host import by name — logs raw args + return. Use when no category fits yet. Example: --trace-host=gdi_draw_edge,wnd_set_state_ptr |
--rpc-census |
With --threads: per-thread histogram of the host imports each guest thread went out to the main thread for. This is the throughput question in worker mode — a blocking import costs that thread a postMessage plus an Atomics.wait, so it stops until the main thread takes a turn — and --host-census cannot answer it, because it wraps the main thread's table and never sees a worker's calls. It is what found Winamp's decode thread spending 9,909 round trips on math_pow and 10,378 on the API-name log. |
--host-census[=N] |
Counts every host import and prints a top-12 histogram (plus the top single-int argument values, so log_i32(0xca00f10f)=9599404 names the exact marker) straight to stdout every N calls, default 1M. Reach for this first when a run hangs or the harness dies with a JavaScript OOM. Everything else we log is buffered and drained only between batches, so a batch that never returns prints nothing at all no matter how much it is doing — this is the one flag that can see inside one. It wraps the final import table, after run.js overrides lib/host-imports.js's versions with its own logging ones — unlike --profile-host, which needs the names up front, wraps before those overrides, and reports at exit. |
--trace |
Every decoded block's EIP |
--trace-eip-stream |
With --trace-eip-range, writes matching EIP lines immediately instead of buffering until the current batch returns. Pair with --trace-eip-detail to include register state, synchronous-message depth and remaining block budget when diagnosing a long nested guest call. |
--trace-sched[=N] |
One compact line each time the set of thread states changes, plus a heartbeat every N batches (default 5000). Shows main + every worker as M:run@0xEIP T1:sleep T2:wait(0xHANDLE). Reach for this first on any "it hangs" or "it's slow" report with threads involved: a stalled system prints the same line repeatedly, a healthy one churns. Doubles as a cheap sampling profiler — --trace-sched=50 then histogram the M:...@0x... addresses. |
--time-scale=N |
Run the guest clock N× the wall clock. Separates "the app is waiting for time to pass" from "the app is doing work": if a slow boot doesn't get faster at --time-scale=10, it is not timing-bound. |
--max-seconds=N |
Stop the batch loop after N seconds of wall clock (the timer starts at loop entry, so app load is not counted), whatever --max-batches says — pass a huge --max-batches with it. The exit line then reads N batches in Ns (M batches/s), and that batch count is the throughput number. This is the axis to benchmark on, because cost per batch is not constant within a run: Caesar is ~0.1ms/batch through its boot and several times that once a city simulates, so a batch count picked to land near a target duration is per-app guesswork that goes stale as soon as the app gets further in the same budget. Fix the duration, compare how far each build got. |
--batch-stats[=FROM_BATCH] |
How many blocks each batch actually retired (p50/p90/p99, share that spent the whole budget) and a histogram of why each batch stopped: budget spent, EIP zero, yield_flag, blocking wait, debug facility. Reach for this before concluding that a region of a run is "slow per batch" — that phrasing hides two opposite causes. A low p50 with budget spent rare means batches keep bailing early and the host pays its per-batch cost for blocks the guest never ran; a p50 at the budget means the batches are full and the blocks themselves are expensive. On Diablo both regions came back full (mean 991 intro, 853 menu), which is what proved the 6x "slowdown" was a unit artifact and not a cliff. Pair with --handler-hist-thread=0 to divide ops by blocks. |
--decode-stats[=FROM_BATCH] |
Per-batch distribution of block decodes, plus the guest slice's wall time beside it. Reach for this on any change to how decoded code is stored or invalidated: it is the one series that is both deterministic (identical across runs of one build, so it is safe to diff between builds) and pointed at the mechanism. --frame-stats's load-immune interval batches cannot see decode cost at all — a batch is a budget of blocks and decoding retires none, so a batch that re-decodes a thousand blocks and one that decodes none look identical there. Read the p50, the decode-free share and the storm line, not the mean: a cache that evicts live blocks pays a steady drizzle every batch, while one-time page compilation concentrates ~95% of its work into ~5% of batches. |
--slice-split=B1[,B2,...] |
Cut one run into phases at those batch numbers and report each phase's total guest-slice wall time, mean ms/batch and block decodes. Reach for this when a run is several workloads and only one is the question — "how long does the save take" is not answerable by timing two runs with different --max-batches, because this box sits at load 10-40 and the two runs are then measured against different machines with both runs' noise in the subtraction. Splitting one run's own per-batch series shares the load across every phase, so the phases are comparable to each other. Read the shares, not the seconds. Measured on StarCraft's save route: boot+gameplay 90.9%, menu 4.3%, the save itself 4.8% — while the same save is 10.9% of route ops and 20.0% of route block entries, a spread that is itself the finding (the save runs at 5.85 ops/block against the rest of the route's 11.98, so it is twice as block-transfer-bound and eats a large share of transfers for a small share of the clock). It turns on --decode-stats collection from batch 0 by itself, since a phase split over a truncated series is meaningless. |
--trace-loopmatch[=0xEIP] |
At decode time, dump every self-loop block the decoder emits (or just the one entered at 0xEIP): entry, op count, and each (handler index, operand) — exactly the input the loop-idiom matcher in src/07b-loop-match.wat sees. Pipe the log through node tools/loopmatch-decode.js <log> [--eip=] [--uniq] to get handler names. Reach for this when asking "why did this loop not get lowered": the answer is a role the matcher does not recognize, and this shows which op it is. Pairs with --loopmatch-stats (self-loop/match/run/byte counts at exit). LUT_RUN is on by default; --no-lut-superops is its narrow A/B partner. COPY_RUN remains off because its broad corpus safety has not been re-established; opt in with --copy-superops. The legacy --loop-superops / --no-loop-superops switches control both families. All flags are applied separately to every guest-thread WASM instance. See §§14–16 of docs/loop-idiom-superops-design.md. |
--trace-seh |
SEH chain operations |
--trace-fpu |
Every x87 exception flag as it is raised ([fpu] raise ZE at 0x…), and every FCLEX/FNINIT that takes them down again. The flags are sticky — nothing but those two instructions clears them — so a program that reads the status word sees whatever the last few thousand instructions left there, and a "Division by zero" message can be reported an arbitrary distance from the divide that set ZE. Reach for this when an app blames the FPU: zero [fpu] lines means the complaint is a software check, not an x87 status read. |
--trace-code-writes |
Why decoded blocks are being thrown away. Every guest write that retired compiled code (address, width, writer EIP), and every block retired because a newer overlapping block was published, as retired <- publisher pairs; the exit summary ranks both. Reach for it when --decode-stats shows storms in steady gameplay. Zero writes plus a few pairs recurring thousands of times is two entries re-decoding each other, an emulator bug, not guest SMC — that was Heroes III's 143K retirements (a load-run fold swallowing an instruction that was already a compiled entry; $fuse_stop in 07-decoder.wat). For real SMC, tools/code-drift.js <log> --pe=PATH compares --input=B:dump-mem: dumps of .text with the PE on disk and classifies each rewritten instruction as operand-only or opcode. |
--fault-null[=stop|=raise] |
Report every guest access no mapping covers — [fault] unmapped guest access 0xADDR from eip=0xEIP — instead of letting $g2w absorb it into NULL_SENTINEL (reads 0, writes go nowhere). =stop traps on the first one so the crash dump names the instruction. =raise does what the hardware does: raises EXCEPTION_ACCESS_VIOLATION at the faulting instruction, so the guest's own __except gets it and an unhandled one exits instead of running on. Reach for =raise when the sentinel is not merely hiding the cause but changing the behaviour — a circular-list walk that never reaches its start node reads 0 forever and loops without bound, which reads as a heap or a hang bug and is neither (B&W2's 0x9e17d0). The faulting instruction still completes against the sentinel before control reaches the handler, so one register may hold garbage that real hardware would have left alone; the block it was in is abandoned, not resumed. Reach for it when a symptom appears far from its cause: a pointer that got zeroed reads as plausible data for thousands of instructions before anything notices. Expect a nonzero baseline — notepad's startup alone probes ~266 unmapped addresses legitimately, so read the addresses and EIPs, not the count. The check lives in the $g2w miss path, after every translation attempt already failed, so an off-run pays nothing. Propagated to worker instances. |
--break=0xADDR[,...] / --break-api=Name[,...] |
Pause emulator at address / API call |
--break-once |
Don't re-arm WASM bp after first hit. Plus prints bp_first_caller (sticky dbg_prev_eip snapshot from the very first time $eip == $bp_addr) — recovers the true caller when the bp lands inside a tight self-loop that would otherwise overwrite dbg_prev_eip with the bp address itself. |
--trace-at=0xADDR (--trace-at-dump=0xADDR:LEN[,...]) |
Log regs + optional hexdump of given regions each time EIP hits addr (no stop). Add --trace-at-watch to diff each hexdump vs previous hit (bytes marked *). Multi-addr --trace-at=A,B,C works but forces BATCH_SIZE=1 (only useful for early-execution probes); for late-code multi-addr fan-out use --count instead. |
--count=0xADDR[,...] (max 16) |
Native WASM hit-counter per address. Reports Hit counts: summary at run end. Full speed (no BATCH_SIZE penalty). Address must be a basic-block entry (call-return landing, branch target, fn entry). Use this for "of N addrs, which fire and how often?" probes. |
module+0xVA syntax |
--trace-at, --count, --break accept module+0xORIG_VA (e.g. d3drm+0x647c3905) — auto-resolved to runtime VA after DLLs load using each DLL's PE-header-declared origBase. Use module name without .dll/.exe; exe is also valid. Eliminates manual + delta arithmetic for cross-DLL probes. |
--watch=0xADDR / --watch-byte=ADDR / --watch-word=ADDR (--watch-value=0xVAL, --watch-log) |
Break when memory at ADDR changes. Size: dword/byte/word. --watch-log logs every change without stopping into debug prompt (essential for non-interactive runs). --watch-value filters to a specific target value. |
--show-cstring=0xADDR[,...] |
On every --trace-at hit and debug prompt, decode 1-byte-refcount + ASCII-at-+1 CString layout. Prints [CString@ADDR] rc=N len=M "text" — great for MFC/Borland apps where strings are packed this way. |
--skip=0xADDR[,...] |
Simulate ret when EIP hits — step past a fn |
--dump=0xADDR:LEN, --dump-seh, --dump-backcanvas |
Post-run memory hexdump / SEH dump / per-window back-canvas PNGs |
Extending: the tracing infrastructure lives in lib/host-imports.js under if (trace.has('gdi')) — it uses a wrap(name, fn, formatter) helper. To add a category, duplicate that block for your category and add a matching if (TRACE_X) traceCategories.add('x') in test/run.js. The generic --trace-host= should cover most one-off investigations without needing a new category.