Tools catalog
Commands and code paths are relative to the repository root.
tools/gen_dispatch.js— Generates09b2-dispatch-table.generated.wat(br_table + calls +$init_dx_com_thunks) fromapi_table.json. COM vtable start IDs are auto-computed from interface prefixes (e.g.IDirectDraw_*), so adding a new API never requires manual ID fixups.tools/gen_api_table.js— Generates the API hash table (01b-api-hashes.generated.wat)lib/pe.js— the PE header/section reader (readPE(fileOrBuffer)→{buf, imageBase, sections, va2off, va2offInfo, off2va, sectionForVa, isCodeVa}). Use it instead of re-derivingreadUInt32LE(0x3c)+ section walk in a new tool. Two rules live here so every tool inherits them:section.isCodeis true for Borland sections named CodeSeg/DataSeg even when flagged as data, andva2offInfo().hasRawis false for BSS addresses that have no bytes on disk (va2offreturns -1 for those).tools/xrefs.js— also importable:require('./xrefs').scanXrefs(file, targetVa, {near, codeOnly})returns the classified hits, which is whatcaller_census.jsuses instead of parsing printed output.tools/disasm.js— x86 disassembler for debugging (importable module)tools/disasm_fn.js— disassemble at one or more VAs:node tools/disasm_fn.js <exe> 0xADDR[,0xADDR,...] [count]. Warns when the start looks like a mid-instruction desync.tools/xrefs.js— find all references to a data/code VA:node tools/xrefs.js <exe> 0xADDR [--near=0xN] [--code]. Classifies each ref asload/store/branch/other; handles Borland-style code-in-data sections (sections namedCodeSeg/DataSegeven when flagged data). Use--nearto catch branches into any byte of a trampoline region.tools/find-refs.js— find every 4-byte pointer literal to a VA (vtable slots, fn-pointer tables, dispatch tables):node tools/find-refs.js <pe> 0xVA [--code-only|--data-only] [--context=N]. Complementsxrefs.js(which finds branches) — use this when xrefs returns 0 but you suspect the fn is reached via an indirect call through a stored pointer. Example: 0 data refs ⇒ fn is not in any vtable; investigate fall-through or computed-address paths.tools/find_fn.js— given an interior VA, locate the enclosing function's entry:node tools/find_fn.js <exe> 0xADDR[,0xADDR,...]. Walks back to the nearest55 8B ECprologue,CC/90padding boundary, orC3/C2ret. Use when a trace hit lands mid-function and you need the entry for--break=or a cleandisasm_fnstart.tools/find_field.js— find all accesses to a struct field[reg+OFFSET]by scanning ModRM displacements:node tools/find_field.js <exe> 0xOFF [--reg=esi,edi] [--op=write,read,lea,cmp,imm,indirect] [--context=N] [--fn]. Use when REing C++ class layouts to locate setters/getters of a specific member offset.tools/caller_census.js— count runtime hits per static caller of a callee:node tools/caller_census.js --exe=PATH --module=NAME --callee=0xORIG_VA [run.js args]. Uses--count(WASM-native, full speed) + module-relative addr resolve. Default probe = callsite+5 (post-call landing ofe8rel32). Output: per-callsite hit count. Max 16 callers (HIT_COUNT slot limit). When you need to know which of N call sites of a hot fn actually fire and how often, this is the one-shot tool.tools/find_vtable_calls.js— locatecall dword [reg+disp](FF /2) sites in a PE/DLL by vtable slot or raw displacement:node tools/find_vtable_calls.js <pe> <slot>(or--disp=0xNN, or--slotsfor a per-slot histogram). Filter base reg with--reg=ecx,edx. Use to enumerate COM call sites for a specific interface method (e.g. slot 32 =IDirect3DRMFrame::AddVisualat disp 0x80). Complementsfind_field.js(data accesses) andxrefs.js(data-VA refs), neither of which filter call-indirects by displacement.tools/find_string.js— find every VA where a string literal occurs in a PE:node tools/find_string.js <exe> "<literal>" [--utf16] [--all]. PrintsVA [section] raw=0xOFF "literal". Use as the first step of a string-driven xref hunt — feed the printed VA intotools/xrefs.js.tools/find_bytes.js— locate every occurrence of a byte pattern in a PE:node tools/find_bytes.js <pe> <hex>or--push=0xIMM(push imm32 callsites:68 ll ll ll ll) or--imm32=0xVAL(any 4-byte LE literal). Filter by--section=.text, optional--context=N. Use to enumerate all call sites pushing a specific msg id (e.g.--push=0x3e8finds everypush 0x3e8site), or scan for a signature instruction sequence. Reach for this BEFORE writing inlinepython3 -cbyte-search scripts.tools/bench-loops.js— time a synthetic guest loop instead of a whole app:node tools/bench-loops.js [--list] [--shapes=lut,store_stream] [--bytes=16m] [--reps=N] [--toggle=case_chain|rle_run|rect_run] [--top=N] [--json]. Injects hand-encoded x86 into a live wasm instance (thetest/test-x86-ops.jspattern) and runs both A/B arms in one process, alternating every rep with the order rotated. Measured noise floor ±1% against the 24-42% that left every whole-app A/B in interpreter-dispatch-perf.md unresolvable — so reach for this before opening another timing worktree. Readblocks/iter, notops/iter: on its calibration run CASE_CHAIN came out +57% faster while printing 7.7% more handler ops (folds 420-424 re-record the ops they replace, and the real variable — block entries — is invisible to the handler histogram). Thenop_chain/jmp_chainshapes price the two primitives with dispatch count held equal by construction: a dispatch is ~8ns and a block transfer adds ~9ns on top of it. Never quote a microbench % as an app % — multiply by the profile share of the machinery it exercises (that is how the +57% here and the ≤2% on Caesar turn out to agree). It understates dispatch cost by construction (a periodic loop is perfectly BTB-predicted) and says nothing about whether a shape occurs in real code, so pair it withfind-loops.js/match-loops.js.blk_mix512is the shape that does not understate it — 512 distinct blocks cycled, so no indirect site sees a learnable target sequence — and the two answers are not close: the block executor's one-block leaf is worth 0% on everyblk{k}shape and +13% onblk_mix512, and the executor as a whole goes from +42% onblk2to -11% onblk_mix512with that leaf disabled. Quoteblk{k}for a hot inner loop andblk_mix512for anything about an app's block working set; a shape re-entering one site forever cannot price a BTB at all. Itsblk_mix512_u{2..12}siblings fix the block size so a breakeven has an answer. A multi-page shape must declarecodeStridelarger than its own code, or rep N's first block lands on rep N-1's tail, the two arms silently share decoded code, and the off arm runs the on arm's folds (oneReprefuses it now). See docs/loop-microbench-harness.md.tools/find-loops.js— census of self-contained inner loops in a PE:node tools/find-loops.js <pe> [--min-body=3] [--max-body=20] [--family=lut,copy] [--skeletons=30] [--json]. Finds short backward branches landing on an instruction boundary, classifies each body (copy/fill/lut/scan/blend/reduce/load2-store/call/rep) and prints a normalized skeleton so the same loop over different registers collapses to one string. Use it to answer "which loop shapes recur across apps" before writing a superinstruction for one of them — see docs/loop-idiom-superops-design.md. Linear sweep, so data-in-code misdecodes: a skeleton seen once is a lead, one seen in several binaries is signal. Says nothing about how hot a loop is; pair with--handler-hist. Its family guess is a loose regex and over-counts badly (stack-counter loops read aslut) — it is a candidate finder; usematch-loops.jsfor the real answer. Importable asrequire('./find-loops').findLoops(file, {minBody, maxBody}).tools/match-loops.js— applies the loop-idiom matcher from docs/loop-idiom-superops-design.md to those bodies and reports what would actually be lowered:node tools/match-loops.js <pe> [<pe>...] [--list=LUT_RUN] [--why] [--json]. Builds the design's real summary (roles → induction variables → memory streams → side effects) and applies the COPY_RUN/FILL_RUN/LUT_RUN/SCAN_RUN predicates, so a match means the predicate held, not that a regex fired.--whyprints the decline histogram — that histogram is the work list for making the matcher more general. Measured across 10 apps: 2.0% of static self-loops match, andcall+multi-branchare 58% of all declines (those are Design B's territory, never A's). Match rate is the wrong metric on its own — check whether the matched VAs are the hot ones from--handler-hist.tools/find-diamonds.js— which loops can$loop_match_blocknot even see?node tools/find-diamonds.js <pe> [<pe>...] [--max-span=N] [--max-blocks=N] [--detail] [--skeletons=N] [--keep-nomemory] [--json]. The matcher only runs on a block that branches to itself, and this classifies every short backward branch into the three populations that precondition creates:SELF(one block, no exit — visible),SELFEXIT(one block + a conditional exit — invisible, because the decoder splits the block at that branch so the head stops branching to itself), andDIAMOND(an internal branch splits the body). Conflating the first two is the error this tool exists to stop: a loop whose exit targets an address past the back edge has zero internal edges and reads asSELF— "already covered" — when it is nothing of the kind (COMI'sexe+0x40340e). Measured 2026-09-19 over 1271 PEs: SELF 47,419 / SELFEXIT 22,917 / DIAMOND 44,633 / COMPLEX 8,253, so 59% of candidate loops are invisible, and the top skeleton is the CRT sentinelstrcpyin 378 binaries — a shared idiom, unlikeRLE_RUN(1 of 287 PEs).--detaillists every cycle with its class and every declined head with the reason, so "the hot loop is absent" stops looking like "the hot loop was declined forrep". Its load/store/advance heuristic is right for a static blit census and wrong for a runtime one —--keep-nomemorykeeps the structural class instead of declining. Static reach is not hotness; pair with the next tool.tools/loop-class-share.js— which loop class does an app actually spend its block entries in?node tools/loop-class-share.js <hist.json> [...] --exe=PATH [--pe=NAME=PATH ...] [--max-span=N] [--top=N] [--json]. Feed ittest/run.js --handler-hist --handler-hist-thread=0,0,0 --handler-hist-start=A --handler-hist-stop=B --hist-json=F(or the browser page probe); it places every hot block inside the cycle containing it and reports the share by class. Read theSELFEXIT/DIAMONDrows (what the matcher cannot reach) and thenonerow — a largenonemeans the hot code is not a short backward-branch loop at all, which is the common case: four of six measured hot regions across five profiled apps, including StarCraft's #1 at 25.8%, which is straight-line branchy code with calls. The share is a property of the app and swings three orders of magnitude (Caesar III ~59-88%, Diablo 0.06%), and one app's own windows disagree about which loops — so quote a VA with its window, and confirm across several withtools/hot-loop-census.jsbefore building anything.~on a class means the cycle failed the blit heuristic: structurally that class, without the memory shape a fold would want.tools/find-rle-nests.js— census of RLE sprite-blit loop nests:node tools/find-rle-nests.js <pe> [<pe>...] [--min-cases=4] [--detail] [--json]. Finds acmp r8,imm8 / jzladder whose targets are branch-free load/store bodies that all jump back to one head — the run-length blit Caesar III writes at0x40f71c, where a token byte selects a fixed-width literal copy and one token is a transparent skip.find-loops.js/match-loops.jscannot see this shape at all: they classify a single self-loop block, and this is a nest of ~20 blocks reached through a jump ladder. Measured 2026-08-24 over all 287 PEs intest/binaries(83 more are 16-bit NE and are rejected): c3.exe is the only hit, so a fold for it is a one-app fold, not a reusable primitive. Its body decoder is a deliberate instruction subset — an unknown opcode rejects the body rather than guessing, so it under-reports; the known blind spot is a decoder that dispatches through a jump table instead of a compare ladder.tools/find-ck-lut-nests.js— census of colour-keyed LUT blit loops:node tools/find-ck-lut-nests.js <pe> [<pe>...] [--detail] [--json]. Finds a short backward conditional jump, decodes linearly from its target, and keeps the nest only when the decode lands exactly on the jump it started from — so it sees the multi-block "diamond" shapefind-loops.js/match-loops.jsare blind to by construction ($loop_match_blockonly ever runs on blocks that branch to themselves, so acmp byte [esi],0xff / jnb skipsentinel test makes a loop invisible to every recognizer). Classifies each hit asCK_LUT{8,16}_{SRC,DST}— keyed, table element width, and whether the table index comes from the source cursor or from the destination word (dst = shadow_tbl[dst]) — orLUT{8,16}_NOKEYwhen there is no key test. Measured 2026-09-12 overtest/binaries(1494 files, ~290 distinct PEs): 613 keyed matches, 297 of them in SimGolf'sjgl.dll, every other binary 1-3. So a keyed-LUT fold is a one-app fold likeRLE_RUNwas; see §19 of docs/loop-idiom-superops-design.md. Match ≠ hot: Pocket Tanks has a textbook keyed LUT16 at0x409a60and its actually-hot blitter is an indirect-call-per-pixel pipeline instead — pair with--handler-histortools/browser-handler-hist.jsbefore building anything.tools/loopmatch-decode.js— turn a--trace-loopmatchlog back into structure:node tools/loopmatch-decode.js <run.js log> [--eip=0xVA] [--uniq]. One entry per self-loop block the decoder emitted, each op named from the(elem ...)list insrc/02-thread-table.wat(so it cannot drift from a renumber). This is the runtime companion tomatch-loops.js, which works statically off the PE: use this one when you need the ops the decoder actually emitted, fusions and all, rather than the ones a linear disassembly predicts.tools/disasm-dump.js— disassemble guest memory captured in a run.js log:node tools/disasm-dump.js <log> [--addr=0xVA] [--nth=N] [--count=N] [--list]. Pair with--input=B:dump-mem:0xVA:LEN. Reach for it for packed, decrypted or runtime-generated code, wheredisasm_fn.jsreads the PE on disk and shows only ciphertext.tools/file2va.js— convert PE file offsets ↔ VAs:node tools/file2va.js <exe> 0xOFFSET[,...]or--va=0xVA[,...]. Use afterstrings -t x/ hex-editor finds, or to translate a VA back to a file offset for patching/inspection.tools/dump_va.js— peek static PE/DLL bytes at one or more VAs:node tools/dump_va.js <exe> 0xVA[,0xVA,...] [len=32]. Marks BSS ranges (no raw data) explicitly so a zeroed sentinel doesn't masquerade as initialized data. Use this instead of--trace-at-dumpwhen you only need to inspect static.rdata/.data.tools/vtable_dump.js— dump function pointers from a vtable in a PE/DLL:node tools/vtable_dump.js <exe> 0xVTABLE_VA [n_slots=16]. Per slot, prints slot index, slot address, target VA, and the first instruction at the target — fast way to enumerate COM/C++ vtables and verify each slot points at a real prologue rather than NULL/garbage.tools/data_offsets.js— print the address of every NUL-terminated string in a WAT(data ...)segment:node tools/data_offsets.js src/01-header.wat 0x11300 [--check=0xADDR,...]. Ordinal-import tables in08b-dll-loader.wataddress these strings by absolute offset, so inserting or renaming one silently shifts every later entry. Use this to confirm an offset still names what its comment claims.tools/hlp-dir.js— list a .hlp's internal files and decode its|SYSTEMtagged records (title, contents, config macros, window definitions):node tools/hlp-dir.js <file.hlp> [--dump=|NAME]. Works on files the WAT parser refuses, so it answers "what is actually in this file".tools/hlp-wat-check.js— what the WAT parser makes of a .hlp:node tools/hlp-wat-check.js <file.hlp> [...] [--topics]. Prints load result (with the named error code and offset), the topic/context/keyword/phrase inventory, and per-topic decode+layout results including the numbered layout and Hall-decompression failure reasons. Reach for this first on any "this help file does not render" report; pair withhlp-dir.jswhen the file will not even load.tools/png-diff.js— compare two PNGs:node tools/png-diff.js a.png b.png [--tolerance=N] [--region=X,Y,W,H] [--out=diff.png]. Prints changed-pixel count/share, worst channel delta and the changed bounding box; exits 1 when they differ, so it chains in a shell. Importable asrequire('./png-diff').diffPng. Use it for "did this refactor change what the screen shows" instead of copying another privatepixelDiff()into a test.tools/record-probe.js— record a real app through the browser recorder and read back what the encoder actually produced:node tools/record-probe.js [--app=ID] [--seconds=N] [--target=screen|window] [--width=W] [--height=H] [--out=FILE]. Prints resolution, real bitrate, bits/pixel, profile, the I/P/B census and mean keyframe spacing. Reach for this on any change tolib/recorder.js, because every quality knob there is a request that MediaRecorder may ignore.It drives headless Chrome, and what it finds holds for real Chrome but not for Safari. Measured 2026-08-26 across all three:
knob headless Chrome real Chrome Safari avc1.640028High profileprofile=Highprofile=Highignored — Constrained Baseline start(5000)keyframe cadenceevery 3.50s every 3.37s every ~1.7s (own ~2s GOP) videoBitsPerSecondhonoured honoured (spent 53% of the offer) honoured (spent 78%) audioBitsPerSecondhonoured honoured honoured B-frames none none none So the tool is trustworthy for Chrome — headless and headful agreed on profile and GOP — and Safari is the outlier: VideoToolbox pins the weakest toolset and runs its own GOP no matter what is asked, leaving bitrate as the only lever that reaches a Safari recording. There is no WebDriver path in this tool, so a Safari answer means recording by hand and running
ffprobeon the file in~/Downloads.Where the ceiling is differs by engine, and it decides whether raising
TARGET_BPPhelps at all: Safari spends most of what it is offered, so more budget reaches the picture; Chrome spent barely half of an 8.7 Mbps offer, meaning its own rate control is the limiter and a bigger cap changes nothing. Check the spend, not the request, before turning that knob.Three traps: a bitrate far under the request usually means a static screen, not a broken setting — check the capture actually moves (
ffmpegtwo frames +png-diff.js) before reading anything into it; headless fps is not a real number, so judge motion only from a headful run; and the audio bitrate is the cheapest tell for which build ofrecorder.jsa file came from (browser default ~186 kbps vs our explicit 128 kbps), which beats guessing whether a cached script was in play. Needsffprobeon PATH for the analysis half; without it the .mp4 is still written.tools/twitter-clip.js— cut a recording down to what an upload will accept:node tools/twitter-clip.js [in.mp4] [--out=FILE] [--crop=W:H:X:Y] [--crf=N]. With no argument it takes the newestwine-assembly-*.mp4in~/Downloads. Measures the content box withcropdetectat five points across the clip and takes the union, so a window that only appears late is not sliced off; fits the result inside 1920x1200, warns past 140s, and re-encodes to H.264 High + AAC with+faststartso the platform re-encodes once instead of resizing first.Since b22dca52 the recorder no longer captures the letterbox bars at all, so the crop half is for clips recorded before that — the resolution and duration fit still apply to every upload.
It refuses to cut a strip whose peak luma says it is picture (measured: encoded black sits around Y=17, so the threshold is 32) and tells you which edge, rather than quietly trimming the dark left side of a scene. Two things it exists to stop you re-learning:
cropdetectandsignalstatswrite to stderr and ffmpeg exits 0 either way, so readingexecFileSync's return value silently yields "no crop detected" while the measurements scroll past — that shipped a still-pillarboxed 43MB file once. And a bar that is not black is content: check before you cut.tools/web-input-probe.js— drive one real browser app with phone-sized touch input and capture the page:node tools/web-input-probe.js --app=sol --query= --viewport=375x667 --touch --steps='wait:4000;key:Enter;key:F2;shot:/tmp/sol.png'.--query=uses the shipping shell without the debug toolbar;--viewport=667x375checks landscape. This is the underlying one-app probe for the sweep below.tools/mobile-app-sweep.js— photograph onlyDESKTOP_APPSon an iPhone SE-sized page in both orientations, using each app's startup recipe so the final picture shows content rather than an empty board or splash:node tools/mobile-app-sweep.js --out=/tmp/mobile-release --jobs=2. It writesportrait/<id>.png,landscape/<id>.png, corresponding-restand-fillshots, andmeasurements.jsonwith fit/fill placement. Use--apps=a,b,--orientations=portrait,landscape,--settle=MS,--launch=MS, and--timeout=SECONDSfor bounded reruns. Review the images; a successful launch or nonblank pixel count alone does not prove gameplay.tools/app-contact-sheet.js— tile a pile of app screenshots into one labelled sheet:node tools/app-contact-sheet.js --dir=/tmp/mobile-release/portrait --out=/tmp/portrait.png --cols=6 --cell=180x320and repeat withlandscapeand a wide cell. Recursively finds every PNG, resolves each name back to an app id inlib/apps.js(tolerating the-a/-b,d_,sol2decorations sweeps produce), picks one capture per app and letterboxes them into a grid with the id under each tile.--pick=largestis the default because a failed capture is a flat desktop-teal PNG of ~2KB while a real frame is 50-400KB, so file size ranks content. Reach for this after any registry-wide sweep: 156 apps in one image is the only practical way to eyeball "does everything in the dropdown still draw". Pure JS (pngjs + a built-in 5x7 font) — no ImageMagick.tools/startup-modal-sweep.js— what is behind the message box an app greets you with?node tools/startup-modal-sweep.js --apps=a,b [--all] [--seconds=N] [--answer=ID] [--shots=DIR] [--json=out.json] [--no-build]. Two passes per app: run it and collect every[MessageBox]line, then re-run it pressing that box's own default button (IDOK for a notice, IDYES for a question) via--input=N:dlg-cmd:every couple thousand batches and photograph what is left. The verdict — clean / content / modal-only / stuck / crash — is about the second picture, so "shows a box" and "shows a box and nothing else" stop looking alike. Reach for this whenever a launch check scores an app as a pass and the contact-sheet tile is a grey box on teal: About boxes (Four Stones, Funtris, Peaks), a welcome screen (Klotski), a question (HyperTerminal's "You need to install a modem"), or a real complaint naming an emulator gap (Imaging's "The Image Admin control cannot be found").tools/crash-sweep.js— does every registry app survive launch, in both scheduler modes?node tools/crash-sweep.js --apps=a,b|--all --modes=coop,threads --seconds=15 --stuck-after=0 --jsonl=dual.jsonl [--jobs=N], then--compare --jsonl=dual.jsonl [--md] [--gate]. Each run is classified into a crash signature (unimpl:,trap:,exit:,stuck,ok, ...) plus the final screen (content/blankfrom the exit PNG) and the waveOut PCM (sound/silent/none);--comparelists every app whose--no-threadsand--threadsruns disagree, worst first, and--gateexits 1 when any do -- the check to run after a large merge. Pass--stuck-after=0: a loader whose main thread waits on worker threads trips run.js's default 10-batch STUCK detector in 0.2 s, so both modes "stick" identically and nothing is compared. Worker scheduling is not deterministic, so re-run a frame-only divergence before acting on it (Red Alert went blank once and matched on the repeat). A boat fork has notest/binaries: give it--feed=(next entry). See docs/crash-sweep-dual-mode-20261006.md.tools/app-files-feed.js— game files for a boat fork, which carries the repo but not the 30 GB of gitignoredtest/binaries. On this box:node tools/app-files-feed.js serve --port=8711(binds 127.0.0.1, serves onlytest/binaries, read-only) andboat forward <id> --reverse --local 8711; neverhostit, these are commercial demos. In the fork:node tools/app-files-feed.js fetch --base=http://127.0.0.1:8711 --apps=a,b|--all [--dry-run] [--dest=DIR]pulls exactly whatrun.js --app=<id>mounts (exe, path dlls,files,localFileManifestand its list, a cdAudio cue and its images, plustest/binaries/dlls), skipping files already present.tools/crash-sweep.js --feed=http://127.0.0.1:8711does the fetch per app just before that app's first run. A fork also comes with an emptynode_modules: runnpm cithere first, or every PNG capture silently falls back to nothing.tools/block-exec-sweep.js— does--block-execstill show the same screen, for every app in the registry?node tools/block-exec-sweep.js --all --no-build [--control] [--budgets=300,600] [--batch-size=20000] [--jobs=4] [--json=] [--md=] [--sheet=]. Four runs per app — {executor off, on} × {two budgets} — and it keeps the off arm's own budget-to-budget diff as the null band, because one budget cannot tell a rendering difference from a frame caught at a different point of a clock-paced animation and without that control there is no scale to judge the A/B against. Classes: IDENTICAL, PACING (differ by less than the null band and converging), DIFFERENT, CRASH (which arm, the crash line, the block-exec counters), TIMEOUT, NOPIC, NONDET. Every DIFFERENT/CRASH row carries the exact repro command line. Pass--controlon anything you are going to quote: it adds a fifth run (the off arm at the larger budget, twice), and on the first registry sweep it reclassified 13 of 21 DIFFERENT rows as the app disagreeing with itself. Measured 2026-09-15 over 201 ids: 157 IDENTICAL, 4 PACING, 0 executor-introduced crashes (all 14 CRASH rows crash with the executor off too), no on-vs-off difference larger than that app's own null band. See docs/block-executor-design/sweep-2026-09-15.md. Known blind spot: on a title whose animation has already settled by the first budget the null band is ~0, so a genuine fraction-of-a-frame phase difference has nothing to be measured against and reads as DIFFERENT.tools/browser-handler-hist.js— the handler/hot-block histogram for an app that only runs in a browser:node tools/browser-handler-hist.js <hist.json> [--top=N] [--blocks=N] [--dump=FILE].--handler-histlives intest/run.js, and anything on the OpenGL path can never reach it —lib/gl-compat.jsneeds adocument, sowglCreateContextreturns 0 headless and the guest exits during startup. The counters themselves are plain wasm exports, so a page can read them; what a page cannot do is name handler 64 or say which DLL block0x00a98462is in. Feed it the JSON fromtools/page-probes/arm-handler-hist.js+read-handler-hist.js(pass those toprofile-web-frames.jsas--after-launch/--report-eval, cooperative mode only — with--threadsthe page instance's counters stay at zero) and it names every handler fromsrc/02-thread-table.watand attributes each block tomodule+0xVA. Counts are load-immune, which is what makes them usable on a box too busy for timing. Read ops/block first: SimGolf comes back at 3.88 against Diablo's 282, which says the cost is block transfer around a per-pixel loop, not the ops inside it.tools/hot-loop-census.js— which loops are hot across several profiling windows:node tools/hot-loop-census.js <hist1.json> [hist2.json ...] [--gap=256] [--top=N] [--blocks] [--json]. Takes the same page-probe JSONbrowser-handler-hist.jsreads, groups adjacent hot blocks into loop regions, and prints each region's share in every window plus the spread between them. Read the spread column, not the mean. One window is not evidence about where an app spends its time, and this is the tool that exists because that cost a session: SimGolf'sjgl+0x10017b6fwas 9.65% of block entries in one sample and absent from the top forty in the next, and a superinstruction (H455) got built for it that moved the frame rate not at all. A region at 53% in one window and 0% in another is a scene, not a fold target; the closing "hot in EVERY window" list is the only line worth building against. Grouping is by proximity, not control flow, so--blocksprints every member VA and a bad grouping is visible rather than silent. Counts and present counts are both load-immune, soblk/framesurvives a busy box. Shares attribution withbrowser-handler-hist.jsviatools/hist-blocks.js.tools/memory-series.js— is a browser run leaking, or is it just big?node tools/profile-web-frames.js --app=ID --seconds=N --headful --before-load="$(cat tools/page-probes/arm-memory-series.js)" --report-eval="$(cat tools/page-probes/read-memory-series.js)" --relay='memseries', thennode tools/memory-series.js <json>(or--from-log=<run.log>). The probe samples wasm bytes, JS heap, canvas count and app count every 2s from before page load, so load-time allocation is in the series instead of being invisible to a sampler armed afterwards. Two conclusions it exists to prevent: "the run got OOM-killed, so we leak" — one tab is a fixed 512 MB of guest heap plus a JS heap that sawtooths to ~490 MB under GC, so it can touch ~1 GB at a peak with nothing wrong; and "first sample to last sample went up, so we leak" — every healthy run allocates during load, so the verdict is a least-squares slope over the second half only, and a series too short to separate a slope from GC noise reports UNKNOWN rather than a number. The wasm and JS series are printed separately because guest-arena size must never be charged to a JS leak.--from-logrebuilds everything from the relayed lines when the page died before--report-evalcould return — which is exactly the run a memory question gets asked about. Measured on Warcraft III 2026-09-13: 951 samples over 31.6 min, wasm flat at its 512 MB floor (0.00 MB/min), JS heap slope 2.6 MB/min — no leak.tools/cache-slots.js— is the block cache too small, or is its index aliasing?node tools/cache-slots.js <hot-block-dump> [--mask=0x3fff] [--hash=fold12|fold8|mul] [--top=N]. Feed ittest/run.js --handler-hist --handler-hist-thread=N --hot-block-dump=FILE, which writes the whole executed-block working set for the window (the printedtop blockslist is only the top 20). It applies the real$cache_slotindex —(ga ^ ga>>12) & CACHE_MASK, withCACHE_MASKread out ofsrc/01-header.watso it cannot drift — and prints occupancy, contested slots, and the cost of each collision as hits into a slot's loser.--maskand--hashmodel a resize or a new index before any WAT is written. Reach for this before growing the cache: on Caesar III the working set is 1861 blocks in 4096 slots and every alternative index scored the same or worse, so its 129610 decodes are not a cache-size problem.tools/mpq-dir.js/tools/mpq-extract.js— Blizzard MPQ archives (Diablo'sspawn.mpq).mpq-dir.jsreads the tables: header, block table,--pos=0xOFF(which block a traced read came from),--table=IDX(sector offsets),--name='dir\file.ext'(hash-table lookup, no listfile needed).mpq-extract.jsdecodes an entry all the way to bytes — decryption plus per-sector PKWARE DCL explode — with--out=,--png=(8bpp RLE PCX → PNG using the trailing 769-byte palette),--frame-height=N(slice a tall multi-frame sprite sheet into one PNG per frame),--palette(index histogram, which names the colour key) and--verify(decode every block, check each against itsfsize). Shared code is intools/mpq.js. Reach for this when a guest asset renders wrong: it is the host-side ground truth for what the bytes are supposed to be, with no emulator in the loop.tools/dump2png.js— render a--dumpregion as a picture:node tools/dump2png.js <log|bin> --width=320 [--bpp=1|4|8] [--flip] [--mode=gray|mask|index] [--skip=N] [--height=N] [--scale=N] [--grid=N] [--addr=0xADDR] [--nth=N] [--binary]. Parses run.js's own hexdump output, so any buffer you can--dumpyou can look at. Reach for it the moment a question becomes "is this run of bytes the right picture" — a hexdump cannot answer that, and thepng-*.jstools all start from a PNG, which is exactly what you do not have while the surface is still inside the emulator.--mode=mask(zero black, non-zero white) shows what guest code reading "not the paper index = ink" will make of an indexed surface, which is often the finding. Get--bppand--flipright before believing anything: a 1bpp 320px sheet is a 40-byte stride, and read as 8bpp it looks like 8x-wide glyphs on every eighth scanline over an eighth of the buffer — convincingly like a rasterizer that gave up early. A positivebiHeightis bottom-up, so--flipis the common case, not the exception.--nthpicks among repeated dumps of one address.tools/hexdump.js— Memory hexdump utilitytools/parse-rsrc.js— PE resource section parsertools/pe-imports.js— PE import table dumper (--alllists all functions,--dll=NAMEfilters by DLL)tools/gfx-app-census.js— which registry apps are the 3D apps:node tools/gfx-app-census.js [--apps=a,b] [--desktop] [--family=gl,d3d9] [--list] [--json]. Walkslib/apps.jsand reports, per app, which of ddraw/d3drm/d3dim/d3d8/d3d9/gl its own exe and DLLs can reach, and whether the evidence is a linked DLL name (dll) or an entry-point literalGetProcAddressis handed (name). Reach for it before any "get all the 3D apps working" claim, because that set was previously recalled from memory rather than measured. Measured 2026-09-22: 87 of 225 apps reach a 3D API, but 85 of those are DirectDraw — the true OpenGL/Direct3D set is 27, and seven of them are thescr_*D3DRM screensavers, which move as one cluster. Data files are deliberately not scanned: a .pak can contain any bytes, so a hit there says nothing. Static reach is not use — Quake II shipsref_soft.dllbesideref_gl.dlland Warcraft III names d3d8, gl and ddraw all three; confirm which renderer a route took with--gl-censusor a traced run. Since 2026-10-06 it also reads every PE module an app mounts (files,localFileManifest) and matches the IDirect3D/2/3/7 IID bytes, the only static trace Direct3D 1-7 reached throughIDirectDraw::QueryInterfaceleaves: the set grew from 38 to 69 apps (Blood II'sd3d.ren, MechWarrior 3, Tomb Raider II/III, NFS III). Aniid-only row is a candidate, not a finding: any DirectDraw game linkingdxguid.libcarries those GUIDs too (Jazz2, Moorhuhn).tools/gl-name-census.js— which binaries can reach a GL entry point the WAT mirror doesn't cover:node tools/gl-name-census.js [path ...] [--families=a,b] [--list] [--json]. The import table cannot answer this: every GL engine in the corpus resolves GL throughGetProcAddress, sope-imports.js ref_gl.dlllists KERNEL32/USER32/GDI32 and no OpenGL at all. The names are still in the file as literals, so this searches for those. Measured 2026-09-22: 35 GL-using binaries, 18 nameglPushAttrib/glPopAttrib— which is what justified building the attribute stack that now mirrors them, so no family latches UNTRUSTED any more. Read--list, not the percentage: a GL driver (3dfxgl.dll,pvrgl.dll) and SDL's loader table name the whole API by construction and are not evidence about any app, whileref_gl.dll's hand-written QGL table, GoldSrc'shw.dlland the Unrealopengldrv.dlls are. Static reach is not hotness and the two already disagree:ref_gl.dllnamesglPushAttriband a 40,000-batch Quake II menu census counted zero calls to it — pair withtest/run.js --gl-census.tools/ne-dump.js— 16-bit NE (New Executable) structure dumper:node tools/ne-dump.js <file.exe> [--segments] [--imports] [--entries] [--relocs=N] [--all]. Every other tool here assumes a 32-bit image, so this is the only way to readtest/binaries/win98-16bit/. Prints the segment table with file positions and flags, the module-reference/imported-name tables, the entry table, and per-segment fixups already resolved toUSER.#113/seg 1:0x0form. Start here for anything about the Win16 games or Explorer's QT_Thunk path.tools/pe-sections.js— PE section header dumpertools/pe-version.js— dumpVS_VERSION_INFOfrom a PE:node tools/pe-version.js <pe> [<pe>...] [--json]. Prints the fixed file/product version plus every StringFileInfo pair (FileVersion, ProductName, OriginalFilename…). Use it to answer "which release is this DLL from" — the corpus mixes DirectX/OLE builds and only this resource says which.parse-rsrc.jswalks menus/dialogs/strings/icons and skips RT_VERSION, and macOSstrings(1)has no-eflag, so the UTF-16 block is invisible to a grep.tools/wasm-native.js— see the machine code a wasm JIT makes of one of our functions:node tools/wasm-native.js --func='$next' [--index=N] [--tier=ion|baseline] [--limit=N] [--top=N]. Every "the JIT probably does X" argument ends here instead. Node's V8 has--print-wasm-code/--print-codecompiled out and macOS refuses a debugger attach, so this goes through SpiderMonkey'swasmExtractCode, which hands back the native code plus a segment table naming each function index — no privileges needed. It is SpiderMonkey Ion, not V8 TurboFan: read it for structure (how many loads a handler really does, whether a bounds check survived, whether the loop stayed in registers, how big the function is), never for a cycle count attributed to Chrome.--topis a size census over the whole module. Needsnpx jsvu@latest --engines=spidermonkeyandbrew install binutils(GNU objdump; Apple's has no-b binary);$SM/$OBJDUMPoverride the paths. Measured 2026-08-24:$nextis 193 instructions and opens every dispatch with a frame setup, a stack-limit check and an interrupt check.--engine=v8reads V8 TurboFan instead (jsvuv8,$D8): d8's disassembler is compiled out too, but--perf-profstill writes a jitdump carrying every code object's bytes, and that is what this parses. Check both before generalizing a register-allocation finding: on toyvm's 785-arm E1 loop (2026-09-30), Ion kept$pcon the stack in every arm once the function had two host-call sites, and TurboFan kept it inw0throughout.tools/indirect-census.js— how many data-dependent indirect branches one of our wasm functions really contains:node tools/indirect-census.js --preset=block-exec [--funcs='$next,$g2w'] [--sites]. It runstools/wasm-native.js's SpiderMonkey extraction once and classifies every register-indirect jump in the disassembly, which is the only way to answer "does this rewrite cost more dispatches than it saves" without guessing at the JIT. On arm64 the three shapes look nothing alike and only one of them is a branch predictor's problem: abr_tableiscmp+ldr x16,[base,idx,lsl #3]+br x16(a jump table, counted with its arm count), areturn_callto a known function is anadr x17trampoline and is not data-dependent, and areturn_call_indirectisldr x8,[table]; add; ldr; br x8(counted as a tail-call indirect).--sitesprints each one's offset and arm count, which is what tells a per-block table apart from a per-micro-op one. Read the site count for BTB pressure and the per-op path for cost — confusing the two is what made the block executor look like it paid "four to six indirect jumps per micro-op against the threaded path's one plus two" when the two paths are in fact level (docs/block-executor-design.md section 25.1). Needs the samejsvuSpiderMonkey and GNU objdump aswasm-native.js, and reports Ion, not TurboFan.tools/inline-verdicts.js— which of our wasm functions V8 refuses to inline, and what that refusal costs:node --trace-wasm-inlining test/run.js ... > /tmp/inl.log 2>&1thennode tools/inline-verdicts.js /tmp/inl.log [--top=N] [--func='$g2w'] [--index=N,N]. The raw trace is tens of thousands of per-site lines; this collapses them to one row per callee, weighted by the call count V8 reports at each site, which is the only ranking that matters (denied at one cold site is noise; denied at sites carrying three million calls is the lever). V8 prices a callee's whole graph against a growth budget, so a hot leaf whose fast path sits behind a cold tier is refused everywhere it is actually needed — this names those leaves. Measured 2026-08-31 (Heroes II, 40000 batches): after the accessor split in accessor-fastpath-split.md,$nextis denied at sites carrying 3.15M calls and$set_regat 1.32M with zero inlined, while$g2whas flipped to 2.42M inlined — so the ~8-13%--wasm-inlining-min-budget=600is still worth is a dispatch/register problem, not a memory-translation one. Indices are per-build; name them withnode tools/func-index.js N.tools/render-png.js— Headless PNG renderertools/check-parens.js— WAT parenthesis balance checker (auto-diffs vs git HEAD)tools/build.sh— Build script (gates + concat +lib/compile-wat.js)tools/check-wat-manifest.js— assertssrc/main.watx's include list ==src/*.watas a set and as an order (missing/duplicate includes are compiler-enforced; this catches the file nobody includes)lib/wat-manifest.js— parsessrc/main.watx's include list; the one place Node reads the source order fromtools/concat-wat.js— writesbuild/combined.watfrom that include list (not a shell glob)tools/check-apps-registry.js— asserts everylib/apps.jsentry points at files that exist (exe, path-form DLLs, data files). Both hosts read that registry now, so a typo'd path breaksrun.js --app=<id>as well as the desktop icon.tools/check-api-table.js— assertsapi_table.jsonids are array positions and the array is append-only vsHEADtools/check-guest-span.js— census of the "translate once with$g2w, then read or write a whole struct through that pointer" pattern:node tools/check-guest-span.js [--min=64] [--file=09a4] [--caller-only] [--json] [--record] [--check]. A$g2wresult is only good for one guest page — two adjacent sparse guest pages need not be adjacent in WASM memory — so a plaini32.store offset=1024through it can land in unrelated memory, which is how SimCity 2000's power-plant picker lost half a BITMAPINFO. It ranks every such local by the furthest byte the function touches through it and marks each CALLER (the guest address came in as one of this function's parameters) or internal. Read the CALLER rows on$handle_*front doors, largest first — that is where the buffer is the guest's own and can straddle. The other rows are mostly false positives by construction: an address the emulator allocated itself lives in an arena we laid out and is linear, and a helper's parameter is often exactly that, so a 22KB CALLER row on a helper means "follow who calls it", not "bug". The fix is$guest_span_in+$guest_span_writeback(or$guest_span_releasefor a read-only span), or per-field$gl32/$gs32on the guest address.--record/--checkratchet the count againsttools/guest-span-baseline.json; the gate is not wired intotools/build.sh, because the pattern is legitimate wherever the pointer is emulator-owned.tools/check-data-strings.js— asserts every(i32.const 0xADDR) ;; Namestring-address annotation still names the string at that addresstools/gen_dispatch.js --check— fails if the generated dispatch table is stale rather than regenerating ittools/deploy-berrry.js— Deploy to berrry.app.--updateupdates an existing app and by default fetches the server's sha256 manifest, then uploads only files whose hash differs (so a no-op redeploy ships zero files).--fullforces a complete reupload.--files=a,b,cuploads an explicit comma-separated list of repo-relative paths and skips diffing. Note: by default--updatewill push uncommitted working-tree changes, since the diff is against the live server, not git.node test/test-diablo-shareware-browser-web.js— drives the shipping desktop shell through Diablo's intro video, red post-skip title, main menu, character selection, loading bar, and gameplay. Saves all six stages tobuild/diablo-shareware-browser/and checks title color, both menu selection markers, portrait, loading bar, and HUD orbs.BASE_URL=http://127.0.0.1:8142targets an existing release server; otherwise it starts its own. Use this browser test for launch-art regressions: CLI captures do not exercise HTTP asset mounts or browser composition.