User mandate (2026-08-02): "do tests like this from now on" -- after the #110 grind bench, where the field composition (peer fire destroying an arm over the wire) replaced the constructed proxy that had let #86 be called fixed while players kept hitting it. The topic carries two halves on purpose: DOCTRINE -- what counts as VERIFIED: 1. reproduce the REPORTED scenario, not a convenient adjacent one 2. scalpel hooks (BT_SELF_DAMAGE_ZONE / BT_KILL_SUBSYS / BT_FORCE_*) locate defects; they support a "fixed" claim only with proven path-identity to the field composition -- and the field composition still gets run 3. MP symptom -> two-node proof (master-side correctness says nothing about what a peer sees) 4. visual symptom -> pixel proof (gotcha 23) 5. coverage claims need the axis enumerated and measured (all gates, all chassis), because per-chassis behaviour lives in authored data 6. an unexplained extra effect in a passing run means the run has not passed MACHINERY -- the bench_common.sh contract (summarized, file = source of truth), single-node and two-node skeletons (relay, ports, affinity, fire cadence, GOTO_STOP standoff), process hygiene (stale-node taskkill first, never double-background, teardown kill order, stale-exe tells), and log-reading rules (capped diagnostics are not evidence of absence; alarm lines are not trends; name the actor at every refusal; field logs have no gates set -- spawn-time summaries ungated, per-frame traces gated). Routed: Quick Lookup row, CLAUDE.md reasoning step 4, build-and-run parity section, reconstruction-method Key Relationships. checkctx CLEAN. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
7.9 KiB
id, title, status, source_sections, related_topics, key_terms, open_questions
| id | title | status | source_sections | related_topics | key_terms | open_questions | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| test-harness | The Test Harness — how benches run, and what counts as verified | living | scratchpad/night6/bench_common.sh (the contract); build-and-run.md §bench-parity; nights 6-9 bench scripts; the #86/#110 verification failure |
|
|
The Test Harness — how benches run, and what counts as verified
Two subjects, deliberately in one file: the MACHINERY (how to launch and read a bench) and the DOCTRINE (what a bench must show before a fix may be called fixed). The doctrine exists because the machinery was used wrong more than once.
THE DOCTRINE: verify the FIELD COMPOSITION, not a constructed proxy
The archetype failure (#86 → #110). "Weapons on damaged arms keep firing"
was fixed and verified by destroying the WEAPON directly (BT_KILL_SUBSYS).
The field scenario was the ARM ZONE dying while the weapon stayed healthy — a
path the bench never took. The claim "fixed" stood for two playtests while the
players kept hitting the bug; the actual defect (a stubbed cascade walk,
gotcha #26) sat exactly in the gap between the proxy and the field. A fix
verified only against a reproduction you constructed is not verified.
Rules, in order of application:
- Reproduce the reported scenario first — the player's composition, not a
convenient adjacent one. If the report is "peer's arm shot off in MP", the
proof is a two-node run where PEER FIRE kills the arm (
mp_armgrind.sh), not a solo self-damage run. - Scalpels are for isolation, not for verdicts. Deterministic hooks
(
BT_SELF_DAMAGE_ZONE,BT_KILL_SUBSYS,BT_FORCE_SEEK,BT_ARMOR_FORCE) are the right way to LOCATE a defect. Before a scalpel result supports a "fixed" claim, prove PATH IDENTITY to the field composition (e.g. the self-damage harness dispatches a realEntity::TakeDamageMessage— the same message + handler a peer round takes — so the scalpel shares every line from the handler down; verified 2026-08-02, and then the field composition was STILL run). - If the symptom is multiplayer, the proof is two nodes. Master-side correctness says nothing about what a peer sees; replication is its own layer (records, edge-detects, instance gates). #94, #84 and #110 each had a peer-side half invisible to any solo bench.
- If the symptom is visual, the proof is pixels —
BT_SHOT_EVERYframes diffed against a control region, not a log line saying the value changed (gotcha #23: a colour change that logs perfectly and renders nothing). - Coverage claims need a sweep, not an instance. "Works for every weapon
type / chassis" = enumerate the axis and measure each point (all 3 fire
gates through a zone death; all 8 chassis' authored cascade flags —
zonesweep.sh). Per-chassis behaviour usually lives in AUTHORED DATA, so fixed code on one chassis proves nothing about the others. - Chase every anomaly in a passing run to ground. The Owens rack cascade also killing the sensor suite looked like a walk bug; it was authored skeleton data (the mast is a child of the rack). A pass with an unexplained extra effect is not yet a pass.
THE MACHINERY
The contract: scratchpad/night6/bench_common.sh
The single source of truth for launching a node THE WAY A PLAYER'S SHORTCUT DOES — read the file, its comments carry the incident history. Summary:
bt_player_env—BT_PLATFORM=glass,BT_START_INSIDE=1(cockpit, not chase cam),BT_DEV_GAUGES=1. Copied verbatim from the shipped .bats;bt_assert_player_envwarns on drift.bt_launch <log> <egg> <affinity> [args]— one node, env-scoped, PID tracked in/tmp/bt_bench_pids.$$.bt_expert_egg— shipped eggs are ALLexperience=expert; a novice egg silences heat/crits/jams and is only legal for the gimp bench (which must say so out loud). See experience-levels.bt_kill_ours— teardown kills ONLY this run's PIDs (a blankettaskkill /IMonce shot down two live sessions).
Single-node bench skeleton
. /c/git/bt411/scratchpad/night6/bench_common.sh
cd /c/git/bt411/content || exit 1
taskkill //F //IM btl4.exe > /dev/null 2>&1; sleep 2 # stale-node clear
sed "s/^map=.*/map=grass/; s/^time=.*/time=day/; 0,/^vehicle=.*/s//vehicle=$V/" MP.EGG > X.EGG
export BT_<gates>=... # diagnostics + scalpels
bt_launch x.log X.EGG 0x03
sleep <duration>; taskkill //F //IM btl4.exe; sleep 2 # then grep x.log
The vehicle sed must replace the WHOLE line (s/^vehicle=.*/), or you mint
vehicle=thr1bhk1 and the spawn FATALs. Chassis codes: ava1 bhk1 lok1/2 mad1/2
own1 snd1 thr1 vul1 (own1 has NO arm zones — rack zones instead).
Two-node bench skeleton (the MP pattern)
bt_assert_player_env
bt_expert_egg MP.EGG X.EGG
BT_MP_LOG=1 <gates> bt_launch x_b.log X.EGG 0x0C -net 1601 # node B first
sleep 2
BT_MP_LOG=1 <gates> bt_launch x_a.log X.EGG 0x03 -net 1501 # node A
sleep 5
python ../tools/btconsole.py X.EGG 127.0.0.1:1501 127.0.0.1:1601 & # the relay STARTS the mission
sleep <duration>; kill $relay; bt_kill_ours
- A
-netnode renders NOTHING untilbtconsole.py(the headless relay) starts the mission — a black window is pre-mission, not a hang. - Affinity: TWO logical processors per node, disjoint sets (
0x03/0x0C); one LP starves the gauge executive (multiplayer peer-shakiness fix). - Give the SHOOTER dense fire (
BT_AF_PERIOD=3+BT_AF_MISSILE=1) and the victim sparse (=9) when one side must lose a grind. Unthrottled autofire trips the FailureHeat all-weapons brick and combat dies. - Combat drive:
BT_GOTO=enemy BT_GOTO_STOP=<units>— 100 ends in a visual ram scrum; 180 gives a standing exchange. - Read each claim on the node that can see it: damage on the VICTIM's log,
replication on the OBSERVER's (
[zone-repl],[mlrec], mirroredDETs).
Process hygiene (every one of these cost a session)
- Stale-node taskkill FIRST. A leftover node joins the next lobby as a third instance and poisons the run.
- Never double-background.
run_in_background+ an inner&orphans the script mid-startup: nodes launch, the script's own cleanup never runs. - Teardown kill order closes windows one at a time — a "crashed" window at the end of a timed run is usually the bench's own taskkill.
- Stale exe:
LNK1104/ a bench ignoring new diagnostics = the previous process still holdsbtl4.exe, or the build silently didn't run — check thebtl4.vcxproj ->line printed, then rerun. - Bench artifacts (
*.EGGcopies,*_NNN.png, bench logs) must be deleted fromcontent/before any dist cut; field logs are never committed.
Reading logs without fooling yourself
- A capped/throttled diagnostic that is silent is NOT evidence of absence
— the #99
[seqrun]shared counter "proved" a sequence never ran; it was throttle starvation. Per-object throttles for per-object questions. - Alarm-line counters are not trend lines —
free=0printed AT the failure is true by definition; the 30s census carries the trend (#32, three separate times). - Log the actor's NAME at every refusal/decision (
'PPC' fire REFUSED); anonymous counters (FIRED #20) cannot support attribution claims. - Field logs have NO diagnostic gates set — a fix whose acceptance evidence sits behind an env var is unverifiable in the field. Spawn-time summaries (one line, ungated) answer questions retroactively; per-frame traces stay gated.
Key Relationships
- Uses: build-and-run (parity, env gates, BT_SHOT capture) · experience-levels (expert vs novice gating)
- Informs: reconstruction-method (step 4 "verify honestly" — this file is the how)
- Incident sources: reconstruction-gotchas §23 (pixels), §26 (silent stubs); multiplayer (replication layers)