Files
BT411/context/test-harness.md
T
Joe DiPrimaandClaude Fable 5 152249cb76 #95: the SCOREBOARD banked one missile per salvo -- credit the DELIVERED amount
Two consumers of one impact, only one was ever verified: the victim gets
TakeDamage{amount=per-missile, burstCount=cluster roll} and applies it
burstCount times (armor was always right); the score post on the next line
sent the bare per-missile amount. An LRM10 salvo dealing 7-35 armor banked
3.5 points -- Rajel's "~3 points to score", to the digit.

Ground truth: the binary's score is the victim handler's tally (amount once
per applied burst @0x4a04da + crit bonuses) reported to the INFLICTING
player (the id-0x16 tail, deferred #45). The shooter-side stand-in now
posts amount x burstCount -- the identical figure handed to the victim, at
the identical one-post-per-TakeDamage granularity. Same pass:

* splash never credited score at all -- the binary tallies every
  TakeDamage; now posted per splash victim (amount x falloff bursts);
* the bridge credited the LOCAL player for ANY registered hit -- AI-master
  fire on the player, a dying mech's death-blast splash; now refused
  unless the shooter IS the local vehicle (MP unaffected: only local fire
  carries live damage on a node);
* direct-fire unchanged -- beams author burstCount=1 (emitter.cpp:355).

Verified per the harness doctrine: single-node field composition (madcat,
real fire, real enemy) 22/22 impact credits paired at damage x burst, zero
bare 3.33 posts; two-node replicant-victim run BOTH directions 31/31
paired across three missile authorings (3.33/2.0/5.0 per-missile) + 25-pt
ballistics, leftovers all burst-1 beam amounts. test-harness.md gains the
map=grass note (MP.EGG authors cavern/night; GOTO mechs shoot rock).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-08-02 14:11:28 -05:00

152 lines
8.6 KiB
Markdown

---
id: test-harness
title: "The Test Harness — how benches run, and what counts as verified"
status: living
source_sections: "scratchpad/night6/bench_common.sh (the contract); build-and-run.md §bench-parity; nights 6-9 bench scripts; the #86/#110 verification failure"
related_topics: [build-and-run, reconstruction-method, multiplayer, reconstruction-gotchas, experience-levels]
key_terms: [bench, harness, field-composition, scalpel]
open_questions: []
---
# The Test Harness — how benches run, and what counts as verified
Two subjects, deliberately in one file: the MACHINERY (how to launch and read a
bench) and the DOCTRINE (what a bench must show before a fix may be called
fixed). The doctrine exists because the machinery was used wrong more than once.
## THE DOCTRINE: verify the FIELD COMPOSITION, not a constructed proxy
**The archetype failure (#86 → #110).** "Weapons on damaged arms keep firing"
was fixed and verified by destroying the WEAPON directly (`BT_KILL_SUBSYS`).
The field scenario was the ARM ZONE dying while the weapon stayed healthy — a
path the bench never took. The claim "fixed" stood for two playtests while the
players kept hitting the bug; the actual defect (a stubbed cascade walk,
gotcha #26) sat exactly in the gap between the proxy and the field. **A fix
verified only against a reproduction you constructed is not verified.**
Rules, in order of application:
1. **Reproduce the reported scenario first** — the player's composition, not a
convenient adjacent one. If the report is "peer's arm shot off in MP", the
proof is a two-node run where PEER FIRE kills the arm (`mp_armgrind.sh`),
not a solo self-damage run.
2. **Scalpels are for isolation, not for verdicts.** Deterministic hooks
(`BT_SELF_DAMAGE_ZONE`, `BT_KILL_SUBSYS`, `BT_FORCE_SEEK`, `BT_ARMOR_FORCE`)
are the right way to LOCATE a defect. Before a scalpel result supports a
"fixed" claim, prove PATH IDENTITY to the field composition (e.g. the
self-damage harness dispatches a real `Entity::TakeDamageMessage` — the
same message + handler a peer round takes — so the scalpel shares every
line from the handler down; verified 2026-08-02, and then the field
composition was STILL run).
3. **If the symptom is multiplayer, the proof is two nodes.** Master-side
correctness says nothing about what a peer sees; replication is its own
layer (records, edge-detects, instance gates). #94, #84 and #110 each had a
peer-side half invisible to any solo bench.
4. **If the symptom is visual, the proof is pixels**`BT_SHOT_EVERY` frames
diffed against a control region, not a log line saying the value changed
(gotcha #23: a colour change that logs perfectly and renders nothing).
5. **Coverage claims need a sweep, not an instance.** "Works for every weapon
type / chassis" = enumerate the axis and measure each point (all 3 fire
gates through a zone death; all 8 chassis' authored cascade flags —
`zonesweep.sh`). Per-chassis behaviour usually lives in AUTHORED DATA, so
fixed code on one chassis proves nothing about the others.
6. **Chase every anomaly in a passing run to ground.** The Owens rack cascade
also killing the sensor suite looked like a walk bug; it was authored
skeleton data (the mast is a child of the rack). A pass with an
unexplained extra effect is not yet a pass.
## THE MACHINERY
### The contract: `scratchpad/night6/bench_common.sh`
The single source of truth for launching a node THE WAY A PLAYER'S SHORTCUT
DOES — read the file, its comments carry the incident history. Summary:
- `bt_player_env``BT_PLATFORM=glass`, `BT_START_INSIDE=1` (cockpit, not
chase cam), `BT_DEV_GAUGES=1`. Copied verbatim from the shipped .bats;
`bt_assert_player_env` warns on drift.
- `bt_launch <log> <egg> <affinity> [args]` — one node, env-scoped, PID
tracked in `/tmp/bt_bench_pids.$$`.
- `bt_expert_egg` — shipped eggs are ALL `experience=expert`; a novice egg
silences heat/crits/jams and is only legal for the gimp bench (which must
say so out loud). See [[experience-levels]].
- `bt_kill_ours` — teardown kills ONLY this run's PIDs (a blanket
`taskkill /IM` once shot down two live sessions).
### Single-node bench skeleton
```bash
. /c/git/bt411/scratchpad/night6/bench_common.sh
cd /c/git/bt411/content || exit 1
taskkill //F //IM btl4.exe > /dev/null 2>&1; sleep 2 # stale-node clear
sed "s/^map=.*/map=grass/; s/^time=.*/time=day/; 0,/^vehicle=.*/s//vehicle=$V/" MP.EGG > X.EGG
export BT_<gates>=... # diagnostics + scalpels
bt_launch x.log X.EGG 0x03
sleep <duration>; taskkill //F //IM btl4.exe; sleep 2 # then grep x.log
```
The vehicle sed must replace the WHOLE line (`s/^vehicle=.*/`), or you mint
`vehicle=thr1bhk1` and the spawn FATALs. Chassis codes: ava1 bhk1 lok1/2 mad1/2
own1 snd1 thr1 vul1 (own1 has NO arm zones — rack zones instead).
**Weapon/combat benches need `map=grass time=day`.** `MP.EGG` authors
`map=cavern time=night`; a bench that copies it without the map sed (e.g. via
`bt_expert_egg` alone — mp_double.sh has this gap) parks the GOTO-driven mechs
against cavern rock and they shoot terrain for the whole window. Always sed the
map/time on the COPY, single-node and two-node alike.
### Two-node bench skeleton (the MP pattern)
```bash
bt_assert_player_env
bt_expert_egg MP.EGG X.EGG
BT_MP_LOG=1 <gates> bt_launch x_b.log X.EGG 0x0C -net 1601 # node B first
sleep 2
BT_MP_LOG=1 <gates> bt_launch x_a.log X.EGG 0x03 -net 1501 # node A
sleep 5
python ../tools/btconsole.py X.EGG 127.0.0.1:1501 127.0.0.1:1601 & # the relay STARTS the mission
sleep <duration>; kill $relay; bt_kill_ours
```
- A `-net` node renders NOTHING until `btconsole.py` (the headless relay)
starts the mission — a black window is pre-mission, not a hang.
- Affinity: TWO logical processors per node, disjoint sets (`0x03`/`0x0C`);
one LP starves the gauge executive ([[multiplayer]] peer-shakiness fix).
- Give the SHOOTER dense fire (`BT_AF_PERIOD=3` + `BT_AF_MISSILE=1`) and the
victim sparse (`=9`) when one side must lose a grind. Unthrottled autofire
trips the FailureHeat all-weapons brick and combat dies.
- Combat drive: `BT_GOTO=enemy BT_GOTO_STOP=<units>` — 100 ends in a visual
ram scrum; 180 gives a standing exchange.
- Read each claim on the node that can see it: damage on the VICTIM's log,
replication on the OBSERVER's (`[zone-repl]`, `[mlrec]`, mirrored `DET`s).
### Process hygiene (every one of these cost a session)
- **Stale-node taskkill FIRST.** A leftover node joins the next lobby as a
third instance and poisons the run.
- **Never double-background.** `run_in_background` + an inner `&` orphans the
script mid-startup: nodes launch, the script's own cleanup never runs.
- **Teardown kill order closes windows one at a time** — a "crashed" window at
the end of a timed run is usually the bench's own taskkill.
- **Stale exe:** `LNK1104` / a bench ignoring new diagnostics = the previous
process still holds `btl4.exe`, or the build silently didn't run — check the
`btl4.vcxproj ->` line printed, then rerun.
- **Editing source through bash-heredoc python collapses `\` one level** even
single-quoted -- a `"\n"` you meant as backslash-n arrives as a real newline,
the "fix" silently no-ops (`good == bad`), and the C2001 hunt repeats. Build
backslashes as `bytes([92])`/`chr(92)`, and verify a replace by LENGTH DELTA
or a re-read, never by the script printing "fixed".
- Bench artifacts (`*.EGG` copies, `*_NNN.png`, bench logs) must be deleted
from `content/` before any dist cut; field logs are never committed.
### Reading logs without fooling yourself
- **A capped/throttled diagnostic that is silent is NOT evidence of absence**
— the #99 `[seqrun]` shared counter "proved" a sequence never ran; it was
throttle starvation. Per-object throttles for per-object questions.
- **Alarm-line counters are not trend lines** — `free=0` printed AT the
failure is true by definition; the 30s census carries the trend (#32, three
separate times).
- **Log the actor's NAME at every refusal/decision** (`'PPC' fire REFUSED`);
anonymous counters (`FIRED #20`) cannot support attribution claims.
- Field logs have NO diagnostic gates set — a fix whose acceptance evidence
sits behind an env var is unverifiable in the field. Spawn-time summaries
(one line, ungated) answer questions retroactively; per-frame traces stay
gated.
## Key Relationships
- Uses: [[build-and-run]] (parity, env gates, BT_SHOT capture) · [[experience-levels]] (expert vs novice gating)
- Informs: [[reconstruction-method]] (step 4 "verify honestly" — this file is the how)
- Incident sources: [[reconstruction-gotchas]] §23 (pixels), §26 (silent stubs); [[multiplayer]] (replication layers)