Commit Graph
2 Commits
Author SHA1 Message Date
CydandClaude Fable 5 3c42766f48 BT410 5.3.62: correction -- the fault address was never in our code
Every conclusion that named EulerAngles::operator=(const LinearMatrix&) as the
fault site was wrong, including the disassembly built on top of it.  Three
independent proofs:

1. ROTATION.CPP is ours, so the function was instrumented directly -- print
   this, &matrix and a local's address on the first twelve calls plus any wild
   pointer.  The run faulted with ZERO [euler] lines: never called.

2. After the probe was added, 0x66D9 disassembles to a two-byte conditional
   jump (jnl) that touches no memory.  The dump reports ErrCode 0004, a data
   READ.

3. EIP 0x66D9 and cr2 0x7000FA64 are identical to the byte across every build
   this session, including builds where the code at that offset changed
   completely.  A fault in our CODE segment moves when the code moves.  This
   one does not -- it lives in the DPMI host, which loads low and never gets
   relinked.

The map lookup was a coincidence: our _TEXT is section-relative from 0, so its
offsets overlap the host's low addresses, and the map will name a function for
any small number.  Before resolving a fault address against the map, confirm it
belongs to our segment -- cheapest check is whether it moves on relink.

What is actually true, by same-binary A/B:

    vRIO DOWN, serial1 still a namedpipe -> mission LAUNCHES
    vRIO UP,   same conf, same binary    -> FAULT

The trigger is a LIVE RIO stream, not an unplugged cable.  The emulator reports
continuous serial1 RX overruns, the fault lands wherever the app happened to be
(terrain one run, the mech's inside view another), and the shipped binary
survives the identical conditions.  A fault at a fixed host address, at
arbitrary points in the app, requiring an interrupt-driven stream, is an
interrupt re-entrancy signature rather than a wild pointer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-29 08:22:15 -05:00
CydandClaude Fable 5 d3b332a62a BT410 5.3.58: the load gate is blocked by a REFILL loop on priority 0, not starvation
Per-priority occupancy, measured every second until the fault:

  [queues] p0=BUSY p1=- p2=BUSY p3=- p4=- nextReady=1

That single line kills the two leading theories.  p3/p4 empty means nothing
above priority 0 is competing, so the pump is free to serve it -- no
starvation.  nextReady=1 means PeekAtNextEvent always has a READY event, so
priority 0 is not holding a timed event whose alarm never comes due -- no
deadlock.  What remains is refill: priority 0 is replenished as fast as the
pump drains it, ~143 events per simulated second.

It is also fully deterministic.  Two runs of the same binary reported
pump=352/1499 at frame 1001 and 495/2643 at frame 2002, identical to the byte.
So 'crashes about half the time' was never a race -- it was comparing runs that
differed in binary or conf.

Only InterestManager::PostRendererEvent posts at priority 0 (it maps every
renderer event there while the app is not RunningMission), so naming the
message names the flooder.  Application::Post now tallies priority-0 posts by
message ID and reports the busiest three.  Notably, all three producers of
NotifyOfNewInterestingEntity are entity-CREATION paths, so if that ID dominates
then something is creating entities in a loop -- which would explain the
unbounded growth behind the fault as well as the hang.

pod_render_noskl.conf now carries the same probes, so one instrumented binary
can be run with the skeleton walk on and off.  Walk-off runs are the only
configuration of ours known to reach 'Turning Plasma Score Display On'.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-28 22:20:11 -05:00