Commit Graph
13 Commits
Author SHA1 Message Date
CydandClaude Fable 5 1233a1d5f6 BT410 5.3.66: RIO-fault eliminations, and TM mode's own NULL-deref surfaced
Three clean eliminations on the live-RIO fault, all on the ack-fixed emulator:
not the ack heuristic (reproduces after the fix), not the debug logging
(pod_render_quiet faults identically with no BT_* env), and not the failed
test-mode handshake (the quiet run's init PASSED and it faulted anyway).

The Thrustmaster isolator -- L4CONTROLS=THRUSTMASTER so the game never opens
the RIO while vRIO keeps streaming into serial1 -- was blocked by its own
discovery: TM mode dies at boot in ControlsInstanceDirectOf<int>::Update+0x6,
mov eax,[eax+0x1c] with EAX=0, at guest 00485E1A.  That address resolves
cleanly in our map and the loader names BTL4REC.EXE CODE 0x75E1A directly: a
real defect in our controls wiring for the non-RIO path, and its own
reconstruction brick.

Standing facts for the hunt: our exe + live vRIO = fault in ~30-40s; shipped
exe + the same stream on the same emulator = fine (two baselines).  The ISR
and LBE4ControlsManager are archive code; ours in the per-event path are the
mapper (MECHMPPR) and the event/receiver plumbing.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 10:24:22 -05:00
CydandClaude Fable 5 e846717283 BT410 5.3.65: THE MECH STANDS -- wire matrix convention proven from engine source
pose_chase.png: the MadCat assembled from its 19 parts, standing in the arena
at its live articulated position.  Feet, legs, torso, weapon pods, canopy
plate.  Not a collapsed stack at the origin, which is what every prior frame
actually showed.

The convention is no longer inferred -- it is read from the engine.
RootRenderable's ctor writes an entity pose into a DCS with
*(Matrix4x4*)dpl_GetDCSMatrix(myDCS) = localToWorld, so
Matrix4x4::operator=(const AffineMatrix&) IS the wire layout: row-major,
rotation in rows 0-2, zeros in column 3, translation in ROW 3 (entries
12/13/14).

The old 3/7/11 choice rested on two errors, both corrected in the code
comments: AffineMatrix does keep translation at 3/7/11, but because it is
COLUMN-major -- citing it as precedent for a row-major layout read the storage
and ignored the indexer (AFFNMTRX.HPP:99).  And 'the last row made the maths
blow up' came from the contaminated bisect; those crashes were the RIO fault.

Rotation now composes through the engine's own Matrix4x4::operator=(const
EulerAngles&) from the page's pitch/yaw/roll, and the node write uses the
RootRenderable idiom (dpl_GetDCSMatrix + assign).

Also banked: the chase-frame fifobridge variant -- camera framed on the
articulated root instead of the arena bounds -- as the offline proof harness
for any future pose question.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 10:02:06 -05:00
CydandClaude Fable 5 d143f833ab BT410 5.3.64: THE COCKPIT EYE IS LIVE -- root+eye composition, and the emulator
deadlock it exposed

The camera gap root-caused and fixed through three layers:

LAYER 1 (ours): "chain the engine after the skeleton" was a non-fix.  The
engine's MakeEntityRenderables only accepts Object/Rubble resources, so
chaining it with a Skeleton printed "wrong video resource type" and built
NOTHING -- the fifodump proved it: zero vr_flush_dcs_artic records.  The real
fix mirrors the engine's Mover composition (L4VIDEO.CPP:4795-4860) inside our
mech case: a Dynamic RootRenderable (whose ctor seeds its DCS from
localToWorld and whose Execute re-flushes on change -- the ONLY source of
per-frame wire articulation), the skeleton hung UNDER its DCS via ReadSKLFile's
new parent_dcs parameter, and for the inside view a DPLEyeRenderable on that
root with the published EyepointRotation.  Also explains the mystery monolith
in the first frame: the skeleton was parked at the world origin with the
camera inside its shins.

LAYER 2 (1995 library, read from our own linked symbols): dpl_DrawSceneComplete
= velocirender_frameack(0), and frameack's unsolicited-reply path prints
"dpl error - unsolicited input during frame ack", adds the 1995 authors' own
puzzled "flush artic??" when the stray action is 0x1f, and calls exit(9).
Documented before it ever fires.

LAYER 3 (emulator, the actual deadlock): with per-frame articulation flowing
for the first time, the game stopped drawing at frame 232 -- one 0x1f burst
per frame forever, no receives, no error.  The VPX device feeds the frame ack
only after 6 CONSECUTIVE empty polls and reset that counter on EVERY
outputData write; per-frame articulation writes made the count unreachable, so
dpl_DrawSceneComplete never went true.  Fix: the reset is gated on
!frame_outstanding -- once a draw is outstanding the only receive the game
will do next is the frame ack.  The iserver-drain concern the reset guarded is
boot-time only, when no frame is outstanding.

Verified: two agreeing runs of ours on the fixed emulator (611 draws / 222
artic batches in run 2 -- the 232 wall is gone), camera travelling with the
mech (cam -329.7,0,60.5 -> -362.3,0,54.5 across 8s), anim_abs 0 -> 1 on the
bridge, view now INSIDE the cockpit cage.  Shipped binary re-run on the fixed
emulator: launches and runs, no regression.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 09:46:23 -05:00
CydandClaude Fable 5 af010c18c2 BT410 5.3.63: the mech gets both skeleton and engine renderables
MakeEntityRenderables no longer stops after the skeleton.  It builds ours AND
chains DPLRenderer, which is what creates the RootRenderable and the
DPLEyeRenderable the camera follows.  Verified over three runs:

  [mer] 208 class=3001 view=1 res=1
     [vid] type=1 file='mad.skl'
  [skl] video\mad.skl -> 26 nodes, 19 objects
     [chain] into DPLRenderer, skeleton=yes, stack at 0x12dde7

No fault, mission launches, sim runs.  Worth noting separately: the engine's
eyepoint construction -- prime suspect for the page fault through most of a
session -- runs perfectly here, with the stack at 0x12dde7, right beside the
main loop's frame.  Both are consistent with the corrected finding that the
fault is a live-RIO artifact at a fixed DPMI host address.

The camera still does not move: two frames six seconds apart differ by zero
pixels while [sim] shows the mech travelling from (487,20,288) to (493,20,301),
and the bridge reads cam (0.0, 10.0, 0.0) -- world origin plus eye height.  So
the eye renderable was necessary but not sufficient.

Eliminated by reading our own code rather than guessing: the viewpoint entity
IS the mech (BTL4APP.CPP:354), the video renderer IS linked to it
(APP.CPP:1337), and SetupCull would Fail() loudly if the mech lacked
"siteeyepoint" or EyepointRotation -- neither Fail appears.  That leaves the
path from SetupCull's worldToEyeMatrix to what the bridge reads off the wire,
which is bridge territory rather than reconstruction territory.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 08:36:51 -05:00
CydandClaude Fable 5 3c42766f48 BT410 5.3.62: correction -- the fault address was never in our code
Every conclusion that named EulerAngles::operator=(const LinearMatrix&) as the
fault site was wrong, including the disassembly built on top of it.  Three
independent proofs:

1. ROTATION.CPP is ours, so the function was instrumented directly -- print
   this, &matrix and a local's address on the first twelve calls plus any wild
   pointer.  The run faulted with ZERO [euler] lines: never called.

2. After the probe was added, 0x66D9 disassembles to a two-byte conditional
   jump (jnl) that touches no memory.  The dump reports ErrCode 0004, a data
   READ.

3. EIP 0x66D9 and cr2 0x7000FA64 are identical to the byte across every build
   this session, including builds where the code at that offset changed
   completely.  A fault in our CODE segment moves when the code moves.  This
   one does not -- it lives in the DPMI host, which loads low and never gets
   relinked.

The map lookup was a coincidence: our _TEXT is section-relative from 0, so its
offsets overlap the host's low addresses, and the map will name a function for
any small number.  Before resolving a fault address against the map, confirm it
belongs to our segment -- cheapest check is whether it moves on relink.

What is actually true, by same-binary A/B:

    vRIO DOWN, serial1 still a namedpipe -> mission LAUNCHES
    vRIO UP,   same conf, same binary    -> FAULT

The trigger is a LIVE RIO stream, not an unplugged cable.  The emulator reports
continuous serial1 RX overruns, the fault lands wherever the app happened to be
(terrain one run, the mech's inside view another), and the shipped binary
survives the identical conditions.  A fault at a fixed host address, at
arbitrary points in the app, requiring an interrupt-driven stream, is an
interrupt re-entrancy signature rather than a wild pointer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 08:22:15 -05:00
CydandClaude Fable 5 d5ae9c4851 BT410 5.3.61: the camera never moves -- our mech case skips the engine's eyepoint
Two frames captured seconds apart are pixel-identical while [sim] shows the
mech travelling from (180.4, 10, -358.7) to (-604.8, 0, -352.4), with the
bridge title reading cam (0.0, 10.0, 0.0) throughout.  The scene renders, the
eyepoint is never driven.

Our own MakeEntityRenderables explains it.  When ReadSKLFile finds a Skeleton
entry it sets handled = True and the chain to DPLRenderer::MakeEntityRenderables
is skipped -- but that is the call which builds the RootRenderable and the
DPLEyeRenderable the camera follows (L4VIDEO.CPP:4548-4566).  We hand the board
a skeleton with no root and no eye.

That also closes the loop on the fault: the eyepoint construction we skip is
the same code that faults when it runs.  The mech case must end up doing BOTH,
so the EulerAngles::operator=(const LinearMatrix&) fault has to be understood
before the mech can be seen from its own cockpit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 23:34:10 -05:00
CydandClaude Fable 5 9c959e6891 BT410 5.3.60: first rendered frame from a running reconstructed mission
emulator/render-bridge/first_mission_frame.png -- arena city, textured
buildings, horizon, 28fps, captured from the GL bridge while our build ran a
live mission with the mech walking.  The skeleton went in with it:
[skl] video\mad.skl -> 26 nodes, 19 objects.

Corrects yesterday's RIO framing.  I called serial1 an unplugged cable.  Wrong
twice: VRio.App is running on this rig (the dev-rig default since 7/17), and
the emulator log recorded real byte counts on that port -- RX overruns of 471,
503, 528 -- which an unconnected pipe cannot produce.  The port was LIVE and
streaming during every crashing run.

So the defect is not that we mishandle a disconnected port, it is that we
mishandle a live RIO stream, which is worse: every production pod has a real
RIO.  The shipped binary on the same rig logs 'lost RIO analog request', keeps
sending characters, and runs the mission; ours logs 'RIO never came back from
test mode!' and dies partway through the load.  The overrun counts say the
guest is not draining the port fast enough.  L4RIO.CPP's handshake is authentic
archive code, but MechRIOMapper (BT/MECHMPPR.CPP) is ours -- that is the thread.

pod_render_norio.conf stays the way to get a running mission today, but it is a
workaround: it disables the cockpit's primary input device.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 23:33:08 -05:00
CydandClaude Fable 5 475ae3106e BT410 5.3.59: A MISSION RUNS -- the RIO serial port was blocking every load
pod_render_norio.conf is pod_render_rec.conf with one line changed,
serial1=disabled instead of the vrio named pipe.  Same binary, same egg.  With
it, our build launches and the mech walks:

  [launch] state=2 minPriorityEmpty=1 ticks=85316
  [queues] p0=- p1=- p2=BUSY p3=- p4=-
  BTL4Application::RunMissionMessageHandler
  Turning Plasma Score Display On
  [sim] pos=(180.466,10,-358.703) yaw=-0.275605 spd=14.4
  [sim] pos=(251.959,10,-250.498) yaw=-0.89103  spd=14.4

No fault, 290 log lines and counting.  This is the first time the
reconstruction has reached a running mission on the pod.

The hang was never starvation, a deadlock, or a refill loop -- it was VOLUME,
the first candidate I listed and then talked myself out of.  530 renderer
events queue at priority 0 during load; BackgroundTasks::Execute runs exactly
one task per call round-robin over seven tasks, and the 1ms frame budget is
always blown here so no extra background passes happen.  That is ~9 events per
real second, so a load legitimately takes ~60s.  The fault arrived at ~40s,
before the backlog could clear.  Every 'the queue never drains' reading was
really 'the process dies before it can'.

The disproof was already in hand: the post tally climbed to 530 and went FLAT,
which means nothing was refilling.  p0=BUSY with nextReady=1 is equally
consistent with a large finite backlog still draining, and I read it as refill.

Why the RIO port: the pipe has no server attached, which should behave as an
unplugged cable.  The emulator log shows steady serial1 RX overruns, our build
prints 'RIO never came back from test mode!', and the SHIPPED binary prints
'lost RIO analog request' and launches anyway on this identical rig.  So an
unplugged RIO is survivable and our handling of it is not -- a real defect, not
a rig artifact, since a pod with an unplugged RIO cable should still boot.

The wild reference at EulerAngles::operator=(const LinearMatrix&)+0x19 is still
unexplained, but it is now reproducible on demand by enabling serial1 rather
than being a coin flip.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 23:23:29 -05:00
CydandClaude Fable 5 d3b332a62a BT410 5.3.58: the load gate is blocked by a REFILL loop on priority 0, not starvation
Per-priority occupancy, measured every second until the fault:

  [queues] p0=BUSY p1=- p2=BUSY p3=- p4=- nextReady=1

That single line kills the two leading theories.  p3/p4 empty means nothing
above priority 0 is competing, so the pump is free to serve it -- no
starvation.  nextReady=1 means PeekAtNextEvent always has a READY event, so
priority 0 is not holding a timed event whose alarm never comes due -- no
deadlock.  What remains is refill: priority 0 is replenished as fast as the
pump drains it, ~143 events per simulated second.

It is also fully deterministic.  Two runs of the same binary reported
pump=352/1499 at frame 1001 and 495/2643 at frame 2002, identical to the byte.
So 'crashes about half the time' was never a race -- it was comparing runs that
differed in binary or conf.

Only InterestManager::PostRendererEvent posts at priority 0 (it maps every
renderer event there while the app is not RunningMission), so naming the
message names the flooder.  Application::Post now tallies priority-0 posts by
message ID and reports the busiest three.  Notably, all three producers of
NotifyOfNewInterestingEntity are entity-CREATION paths, so if that ID dominates
then something is creating entities in a loop -- which would explain the
unbounded growth behind the fault as well as the hang.

pod_render_noskl.conf now carries the same probes, so one instrumented binary
can be run with the skeleton walk on and off.  Walk-off runs are the only
configuration of ours known to reach 'Turning Plasma Score Display On'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 22:20:11 -05:00
CydandClaude Fable 5 4cf09917ce BT410 5.3.57: the crash is deterministic, and the real blocker is the load gate
The previous commit's attribution was wrong and is corrected in the roadmap.

WRONG: 'the crash is the skeleton walk's' rested on 0 crashes in 2 runs with
the walk disabled -- a 25% coin flip presented as evidence.  The oldest
preserved dump has no [skl] line and ends on the old 'couldn't figure out how
to MakeEntityRenderables' fallback, so it predates the walk and faults
identically.

WRONG: 'it lands at different points each time'.  Every dump carries the same
numbers to the byte (cr2 7000FA64, EIP 66D9).  What varies is how far the log
gets, not where the fault is.

Resolving the address through btl4opt.map names it exactly:
EulerAngles::operator=(const LinearMatrix&)+0x19, a read through the matrix
reference.  A binary scan finds all three call sites of that operator pass a
stack local, so the 1.79GB pointer still has no static explanation -- that is
now its own open item rather than a guess.

Two theories killed cleanly by the new BT_STACK_LOG probe and a binary scan:
ESP drift (drift=0 over 2002 frames; exactly one callee-cleans function exists
in the whole binary) and an undersized stack (our PE and the shipped one have
identical 1MB/8K geometry).

What the probe found matters more: the application is parked in state=2, which
is LoadingMission, not WaitingForLaunch.  Priority 0 is where the interest
manager queues renderer events during load, the gate needs that priority empty,
and it never empties -- so the mission never launches and the renderer holds a
blank screen by design.  The fault arrives ~30s into that wait.  Runs that DO
launch never print a single [launch] line.  So 'crashes half the time' and
'hangs during load' are one event seen twice.

A 15x disagreement between two clocks looked like a reconstruction slip --
BTL4.CPP passes GetTicksPerSecond() where ApplicationManager wants a frame
rate.  Checked against the shipped binary before touching it: same instruction
sequence, same kind of static float pushed.  Authentic.  Documented so nobody
'fixes' it.

Also swept every subsystem DefaultData against its real C++ base.  Fifteen
chain past their immediate parent, but fourteen skip only classes that add no
handlers and no attributes, and no class's attribute-ID base disagrees with its
index chain -- so there are no gap slots there.  The one real defect: Generator
is a HeatSink but chained to Subsystem::MessageHandlers, so it ignored every
ToggleCooling message.  Fixed (compiles next build).

Tooling: podrun.sh stages the build over BTL4REC.EXE, exports the host-side VPX
board env -- without which the run dies at the iserver handshake rather than
merely rendering nothing -- and archives every run's log, marking it -CRASH
when it faulted.  Before this the only preserved dump was an accident, on a rig
where each run costs four minutes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 21:41:38 -05:00
CydandClaude Fable 5 f54e1a3c19 BT410 5.3.56: the intermittent pod crash is the skeleton walk's -- attributed by same-binary A/B
Added BT_NO_SKL, which skips the skeleton build so ONE executable can be run
with and without it.  That is the only clean way to test an intermittent
fault, and it settles attribution:

  walk enabled   ~50% of runs die (page fault at 66D9) at DIFFERENT points
  walk disabled  0 crashes in 2 runs
  shipped exe    0 crashes in 2 runs, same rig and conf

So it is our new code, not the rig, and the varying fault point points at
memory corruption surfacing later rather than a bad instruction in the walk.

Hypotheses killed by reading the private library headers: the matrix array is
the right size (dpl_MATRIX is float32[4][4]); and neither s_dplobject nor the
common dpl_node header retains an object name, weakening the dangling-string
theory (a private name cache inside the library is still possible and is the
best surviving suspect).

Also checked something that would have been much worse than a crash:
~NotationFile REWRITES its file when dirty, and these are the game's shipped
.SKL data files.  MAD.SKL is untouched, and ReadSKLFile now carries the
engine's own Verify(!IsDirty()) guard.

Next tests are listed cheapest-first in the roadmap: keep the NotationFile
alive, copy the Object= names, count board allocations, and re-check whether
MakeEntityRenderables runs twice per mech.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 20:45:28 -05:00
CydandClaude Fable 5 15d509c8a8 BT410 5.3.52: frame rate measured properly -- the missing mech skeleton is the real gap
The pod updates every few seconds.  Measured by sampling a head window every
2s (emulator/render-bridge/headrate.py): OURS changed in 5 of 12 samples,
SHIPPED in 1 of 12.  So the reconstruction is not slower than the shipped
binary -- the emulated board is just expensive.  Caveat recorded with it:
ours was being driven by the throttle hooks while the shipped exe ignores
them and sat parked, so treat those as same-order, not a win.

What did cost us 3x was mine: BT_MECH_LOG/BT_LAUNCH_LOG write per-frame lines
and DEBUG_STREAM=cout is redirected to COM3, so every one goes through an
emulated serial port.  Turning them off took the wire from 480 to ~1480
bytes/sec.  pod_render_quiet.conf is the conf to use for timing work.

And a correction to my own earlier framing: wire bytes/sec is NOT a frame
rate.  Shipped pushes ~8900 B/s against our ~1480 and the difference is
CONTENT, not speed -- shipped submits the mech and we do not:

    OURS:    L4VIDEO.cpp wrong video resource type for object mad.skl
    SHIPPED: (no such line)

The mech's model resource is a SKELETON.  The engine's default
MakeEntityRenderables accepts only Object/Rubble and rejects anything else
with exactly that message, because skeletons are the GAME renderer's job.  So
our arena renders but the MECH IS ABSENT, and with it the cockpit interior --
one unimplemented path explaining both the missing model and the lower wire
volume.

Next brick is now precisely scoped: answer MechClassID by reading the .SKL
notation pages into a dpl_DCS hierarchy and hanging the per-node .BGF
geometry off it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 14:04:35 -05:00
CydandClaude Fable 5 09192367c0 BT410 5.3.51: THE 3-D WORLD RENDERS -- reconstruction to emulated board to a picture
emulator/render-bridge/first-3d-frame.png: the arena from the pod -- sky,
horizon, ground, the arena structures along the skyline -- at ~31fps on the
VelociRender bridge.

The whole chain now works in our build: MakeVideoRenderer ->
BTL4VideoRenderer over DPLRenderer -> board boot -> Renderer::LoadMission ->
DPLReadEnvironment for the art paths -> 40/40 arena objects loaded -> mission
launch -> per-frame submission over the wire.

The mistake worth remembering: the run was never stalled after
InitializePlayerLink.  I called that from a 150-second sample and it was just
too short.  CheckLoadMessageHandler reposts every second and will not advance
until the min-priority event queue drains, and with the video renderer in the
mission that takes minutes.  A BT_LAUNCH_LOG trace showed the queue draining
and then both RunMissionMessageHandler calls, the plasma display, and the
first sensor tick -- matching the shipped binary line for line.  The renderer
deliberately holds a blank screen until RunningMission (L4VIDEO.CPP:5100), so
'black' was the app still loading, not a rendering bug.

Zero unbuildable entities, zero geometry failures.  The frame is static only
because the mech is parked with no input, and the camera in the bridge title
is the bridge's own viewer, not the game's eyepoint.

Also banked the launch-gate trace (BT_LAUNCH_LOG) behind an env var.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 10:30:20 -05:00