c578a9b7904e1bd1707f2e5ebac1a181775722da
424
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
c578a9b790 |
BT410 5.3.95: the stride becomes the speed -- and the hull now breathes with the gait
Half of IntegrateMotion's job, taken now because two findings made it safe:
THE GAIT SLEW RATE IS THE AUTHORED ACCELERATION. The binary's ctor sets
forwardCycleRate, gimpCycleRate and groundCycleRate all from the model's
maxAcceleration (rec+0x44 -- 30 for the Mad Cat) and airborneCycleRate from
superStopAcceleration. That closes both of yesterday's unsourced-rate
worries at once: the ctor default of 1.0 (which would have taken ~27 seconds
to reach speed) is retired, and replacing the explicit acceleration model
with the stride rate preserves drive feel EXACTLY -- same slew rate, same
demand, different (authentic) path for the value to travel.
SO: Simulate now takes currentBodySpeed = bodyStride / dt, the binary's own
relationship (IntegrateMotion @004ab1c8: local velocity z = -distance/dt).
The old acceleration model survives only as a [T3] fallback for a model with
no gait clips at all, so unverified fleet mechs stay drivable.
LIVE CURVE, arena run, no fault:
spd = 7.32 stepping off (stand->walk clip)
spd = 26.67 22.16 27.31 30.54 ... the run cycle around the demanded 26.9
The old model pinned spd at exactly 26.9022 every frame. Now the mech steps
off slow and its hull speed oscillates within each stride -- heel-strike vs
mid-swing. Cadence in the motion is not noise; it is the point. The mech is
carried by its animation, which is what "the gait IS the locomotion" meant in
the 1995 design.
Also queued a correction from the raw decomp: IntegrateMotion's replicant
branch writes +0x598 with a VECTOR op, so motionEventName is likely a Point3D
rather than the CString the BT411 member map suggested -- flagged in the
sidecar for the dead-reckon increment rather than churned now.
Still staged from IntegrateMotion: the airborne pick, the dead-reckon fold,
the orientation integration, the turn-in-place dispatcher.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
9a74f2aa09 |
BT410 5.3.94: THE LEGS WALK -- AdvanceLeg/BodyAnimation live, 16 of 22 joints animating in paired strides
The last link is in. AdvanceLegAnimation (@004a5028) and AdvanceBodyAnimation
(@004a5678) reconstructed and wired into Mech::Simulate, and the first live
run put a walking gait on the wire:
before 22 handles, 1 animating (the vehicle root)
after 22 handles, 16 animating, ZERO 2-float records
root 812 poses; six PAIRS at 650/649, 599/596, 563/563,
542/540, 531/529, 434/421; three slow joints at 25
Six left/right pose-count pairs is six joints per leg cycling in alternating
strides -- the symmetry is itself evidence that the handed clip alternation
(Right, Left, Right, each clip one stride) is running correctly. Chain, end
to end, every stage previously verified in isolation and now live together:
mapper demand -> AdvanceLegAnimation (state machine) -> SelectSequence
-> SequenceController::Advance (keyframes -> Joint::SetHinge/SetRotation)
-> BTL4HingeRenderable (5.3.88 matrix transport) -> 12-float flush
-> render bridge applies. Mission drove itself clean, no fault.
RECONSTRUCTED FROM THE RAW DECOMP, NOT THE DONOR -- and the sidecar says why:
BT411's versions carry port-era replicant accommodations (its mapper cell
does not replicate; the binary's does) and a turn-in-place dispatcher it
relocated INTO the leg machine from mech4's master performance. None of that
is 1995 code. Here the leg version reads the mapper unconditionally, exactly
as decompiled, and state 4's ARMING stays where the binary has it -- in
mech4, not yet reconstructed.
THE TWO CHANNELS DIFFER MORE THAN THEIR ClipFinished TWINS DO, all
binary-verified: the leg version has the wind-down block, the turn-in-place
case and the "Standing Not Supported" guard; the body version has none of
those, its case 4 sits in the plain-advance group, and move_joints reaches
every body Advance AND its reset's Reset -- the caller decides whether the
body channel poses joints or only measures stride. Wired accordingly: leg
poses, body measures (move_joints 0), body distance dropped at a seam marked
STAGED -- consuming it as the forward step is IntegrateMotion's job (mech4).
Also faithful: case 0 FALLS THROUGH so a freshly armed clip advances the same
frame it was selected; the plain group's Standing guard is unreachable via
that fall-through and catches direct entry only; each cycle plays its clip at
cycle/stride of the authored rate (a slow walk IS the walk clip played slow);
and the reverse cycle's caps are all negative with the advance ratio folded
positive at the end.
UNSOURCED, named in the sidecar rather than invented: idleStrideScale
(+0x5ac, defaults 1) and runSpeedMax (+0x7a0, the run cycle's upward cap --
LoadLocomotionClips does not set it; defaulted huge so it never binds until
its real writer is found). ForceUpdate(8) is a STAGED no-op pending the
replication emitter.
Still deferred: the airborne Advance* flavours (@004a5bf8/@004a71f4, jump
jets) and the Gimp*ClipFinished limp machines.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
a87c49c158 |
BT410 5.3.93: the clip loader runs LIVE -- the Mad Cat measures itself, and a 1995 copy-paste bug ships again on purpose
LoadLocomotionClips + ResolveAnimationClip + MeasureClipStride + LoadClipSlot
reconstructed and WIRED into the ctor's GameModel block. animationClips[] is
no longer an empty array: every gait slot resolves by the model's animation
prefix, and the locomotion constants are now MEASURED from the authored clips
instead of asserted as bring-up defaults. First live run, arena mission:
[mech] clips 'mad': standSpeed=5.23 walkStride=18.51 revStride=56.05
revSpeedMax=26.26 gimpSpeedMax=-4.23 gimpStride=-20.26
limpSet=1
That line carries three verifications at once: the prefix printing as text
proves the Mech__ModelResource layout is right at +0x40; the reverse figures
come out NEGATIVE exactly as the transition machines expect; and the mission
ran clean to live driving afterwards (703 log lines, no fault).
DRIVING FEEL CHANGED, deliberately: speedDemand at 0.6 throttle went 14.4 ->
26.9, because the placeholder top speed (30) gave way to the measured 56.05.
The Mad Cat is simply faster than the bring-up guess. Authenticity arriving,
not a regression.
TWO BINARY BEHAVIOURS REPRODUCED ON PURPOSE, both documented at the function:
The speed caps read keyframeData[keyframeCount] -- one entry PAST the last
frame. Fencepost is the binary's own (0x690 + 8 + [0x670]*0xc); whether the
authored table has count+1 entries is unestablished, but the clips were
authored against this read and the measured values are sane.
The reverse-cycle stride is computed from STALE locals. The decomp is
unambiguous: bbr and bbl are both measured into local_8/local_c, and the
divide's second terms come from local_10/local_14 -- still holding the
RUN-LEFT figures. gimpStrideLength = -((bbl + rrl_stale)/(...)). A 1995
copy-paste bug, shipped in every pod for thirty years, reproduced here with
a comment pointing at the wwr/wwl block that shows the intended pattern.
(And an earlier scare resolved: the negation IS in the binary -- the very
next instruction is 0x350 = -0x350. My first decomp window cut one line
short and briefly indicted the donor's minus sign.)
ONE DELIBERATE DIVERGENCE, tagged [T3]: the binary dereferences every resolve
result unguarded -- a model missing a mandatory clip crashes on load. Here a
miss stores NullResourceID (SelectSequence resolves it to an inert controller)
and the dependent measurement is skipped. Keeps the boot alive on unverified
clip sets; revisit when the fleet's models are known-good. Measurement binds
pass a NULL finished-callback (measurement parses, never plays -- the binary's
live pointers can never fire there).
Ctor additionally zero-initializes the whole gait channel -- globalTimeScale
defaulting to 1 specifically, because zero would silence every clip advance --
and fills the clip array with NullResourceID before the loader runs, so the
uninitialized-member class of bug (see 5.3.83) is closed here BEFORE the
consumers arrive.
MECH.HPP: the optional limp set carved out (hasGimpClips + 4 measured limp
figures + gyroRumbleTimer); reservedState 150 -> 140.
Next: the four Advance* entry points -- the last link before the legs move.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
64f99ffcc8 |
BT410 5.3.92a: reconcile two stale claims in MECH2.NOTES.md
The slot-map section added in 5.3.92 corrected two things the doc still asserted higher up: that gimpStrideLength is authored negative (the sign is applied at measurement) and that animationClips is sized AnimationCount (it is AnimationSlotCount, 0x21). Both now agree with the correction rather than leaving a reader to hit the stale version first. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3629755d90 |
BT410 5.3.92: the gait slot map -- the clip array is bigger than the name table, and the "gimp" members are the REVERSE figures
Went to source animationClips[] and found two things: a bug in what 5.3.91 committed, and a naming trap that has already cost BT411 a shipped defect. THE BUG, MINE. I sized animationClips[AnimationCount] -- 0x1d, from the enum. Wrong. The name table stops at 0x1d but the ARRAY does not: slot 0x20 (mech+0x64c) is the bump/crash clip the mech binds on a hard wall impact. So the array is [AnimationSlotCount] = 0x21 and Set*Animation's Verify bounds against that instead. Sizing a real array off a name table that stops earlier reads fine and corrupts whatever sits next door; caught by reading the clip loader, not by the compiler. THE SLOT MAP, recovered from LoadLocomotionClips and now written down in full (suffix, meaning, and which measured constant each clip yields): 5 swr stand -> walk standSpeed 6/7 wwr/wwl forward walk CYCLE walkStrideLength = (s6+s7)/(d6+d7) 8/9 wsr/wsl walk -> stand 10/11 wrr/wrl walk -> run reverseSpeedMax 12/13 rrr/rrl run CYCLE reverseStrideLength 14/15 rwr/rwl run -> walk 16/17 sbr/sbl stand -> back gimpSpeedMax 18/19 bbr/bbl reverse CYCLE gimpStrideLength (NEGATED here) 20/21 bsr/bsl back -> stand 22/23 wgl/wgr walk -> limp gimpLeft/RightSpeedMax 24/25 ggr/ggl limp CYCLE gimpLeft/RightStrideLength 26/27 gsl/gsr limp -> stand 0x20 bmp bump / crash stagger THE TRAP: THE "gimp*" MEMBERS ARE THE REVERSE FIGURES, NOT THE LIMP ONES. gimpSpeedMax and gimpStrideLength are measured from sbr and bbr/bbl -- the reverse gait. The real limp has its own gimpLeft*/gimpRight* pair. This is the same bad naming that produced the states-16-19 misreading recorded in 5.3.91, and it has now caused the same error twice from two directions. Also settled: gimpStrideLength's negative sign is applied AT MEASUREMENT, not authored into the data -- which is where the fold in the transition machines comes from. And the limp clips are OPTIONAL: the loader probes for wgl and leaves hasGimpClips 0 with slots 22-27 unfilled if the model lacks it, so the deferred gimp branch must check that before routing into the limp machine. THE ENUM IS NOT THE SLOT MAP, and MECH.HPP now says so at the enum itself. The names are verbatim from the binary and authoritative AS NAMES, but slot 0x0e takes the run-to-walk clip while the table calls it RightReverseAnimation, and the forward walk alternates 6/7 rather than the pair the WalkForward names suggest. Read the slot map for "what does this play"; read the enum for "what did the original call this index". ATTRIBUTION NOTE: the four clip helpers (ResolveAnimationClip @004a7f50, MeasureClipStride @004a8054, LoadLocomotionClips @004a80d4, LoadLocomotionClipsExt @004a86c8) are exactly the four addresses the manifest lists under mech2.cpp that BT411 files under mech3. The manifest's attribution comes from the binary's own file tagging, so they belong here -- which also accounts for all 12 of mech2's functions. BT 51/51. Still nothing calls the gait. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d8aea8d871 |
BT410 5.3.91: the gait transition machine -- two channels, one table, and a reverse cycle that has been misread twice
mech2.cpp exists. Six of its twelve functions are reconstructed: the two
Set*Animation binders, the two *Transition tails, and both *ClipFinished jump
tables (@0x4a69aa leg / @0x4a6e0a body). Compile-verified, BT 51/51, links
clean.
TWO CHANNELS, DELIBERATELY NEAR-DUPLICATE. A mech runs two clip channels over
the same states and the same clips; only the speed they consult differs:
LEG reads the LIVE mapper GetSpeedDemand() -- responds to the stick at once
BODY reads bodyTargetSpeed, a snapshot -- which is what lets a dead-reckoned
or networked mech walk with no controls mapper of its own
So LegClipFinished and BodyClipFinished are twins rather than one shared
routine, exactly as the binary has them. Kept that way on purpose: where the
two jump tables agree, a divergence in this file is a bug, and that mutual
check is worth more than the duplication costs.
EVERY CLIP IS ONE STRIDE, which is why every state is handed -- a walk is
Right, Left, Right, and each entry to and exit from a cycle has its own handed
pair so the mech always leaves on the correct foot. The 29-state enum is
VERBATIM from the 0x3c-stride name table at .data:0050cfe8, the table the
"Unsupported mech animation" assert indexes, so the names and their order are
the original's rather than inferred from behaviour.
THE COMMIT TEST. Both exits that leave a walk cycle test the demand AND the
current cycle speed slewed by one carryover:
if (demand < standSpeed && (cycle - rate*carryover) < standSpeed) -> stop
if (demand > walkStride && (cycle + rate*carryover) > walkStride) -> run
Requiring both is what stops a momentary flick of the stick yanking the mech
out of a stride it has already committed to. Drop either conjunct and you get
a mech that stutters between gaits on noisy input.
TWO THINGS THAT READ WRONG AND ARE NOT:
gimpStrideLength is authored NEGATIVE. The cycle time from it comes out
negative and is folded positive before being spent (@0x4a6c6e / @0x4a6d3d).
That fold is not defensive coding -- remove it and the limp plays backwards.
States 16-19 on the body channel are the REVERSE gait, not a limp, despite
sharing the gimp caps. BT411 records misreading these as "gimp, fall back to
standing" TWICE; that makes the body loop stand -> reverse-entry forever, a
slow reverse with a wrong-footed exit. Read here by structural symmetry with
the leg table, where every previously-decoded body case mirrors its leg twin.
DEFERRED, and named in the sidecar: the four Advance* per-frame entry points
and the two Gimp*ClipFinished limp machines -- with them the gimp-level branch
at the top of both *ClipFinished, so a limping mech currently runs the normal
machine. That branch needs a TU-safe read of the graphic alarm level (BT411
routes it through a mechdmg bridge to dodge an AlarmIndicator ODR split), worth
reproducing carefully rather than reaching for the alarm directly.
NOTHING CALLS ANY OF THIS YET. The Advance* functions are the entry points and
they are the deferred half, so no gait state is ever selected and a run behaves
exactly as before. Same caveat as 5.3.90: a blocker removed, not a behaviour
delivered.
HEADER: MECH.HPP gains the enum, six declarations, and the channel state --
legStateAlarm/bodyStateAlarm (read via GetLevel; the binary's +0x3b0/+0x728 are
mirrors of the alarm level, so no separate int is kept), legCycleSpeed,
bodyCycleSpeed, forwardCycleRate, gimpCycleRate, standSpeed, gimpSpeedMax,
gimpStrideLength, globalTimeScale, animationClips[0x1d]. 41 ints carved from
reservedState, 191 -> 150. walkStrideLength/reverseStrideLength/
reverseSpeedMax/bodyTargetSpeed already existed from the Phase 5.3 locomotion
work and are reused.
STILL UNSOURCED: animationClips[] is declared but nothing fills it. The clip
handles come from the mech's model resource and must be resolved before the
Advance* increment, or SetLegAnimation hands SelectSequence a garbage ID. It
fails soft (SelectSequence tolerates a missing resource with an inert
controller) but it is a hard prerequisite for the gait doing anything.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
c2b0a371ba |
BT410 5.3.90: the gait player is real -- SequenceController::Advance was the Fail() stub under the legs
Asked what was next for the rebuild, went to measure it, and found I had told
the operator the wrong thing one turn earlier.
THE CORRECTION. I said mech2.cpp (the gait state machine) was the next
package. It sits on top of SequenceController's playback methods, and all
three of those -- SelectSequence, Advance, Reset -- were Fail() stubs. Advance
is the function that actually walks a clip's keyframes and writes each
animated joint. mech2 without it would have called straight into a Fail. So
the dependency runs:
Mech::AdvanceLegAnimation mech2.cpp -- the STATE MACHINE, still absent
-> SequenceController::Advance THIS -- keyframes -> joint writes
-> Joint::SetHinge / SetRotation / SetTranslation
-> the joint DCS flush fixed in 5.3.88
5.3.88 fixed the bottom link, this fixes the middle, the top is still missing.
WHAT THIS DOES NOT DO: make the legs move. Nothing calls any of it yet -- no
gait state is ever selected because mech2 is unreconstructed -- so a run looks
identical to yesterday's. This removes a blocker. Saying otherwise would be
easy and wrong.
RECONSTRUCTED (compile-verified, BT 50/50, links clean):
SelectSequence @004277a8 -- find + lock the clip, parse its layout. The
fetch is FindResourceDescription and NOT SearchList: the ID arriving here is
already resolved, and SearchList would treat it as a resource LIST and walk
the clip bytes as IDs.
Advance @0042790c -- snap through every keyframe the new time has passed
(writing authored poses), then interpolate the partial frame. Returns the
forward distance covered, which is what the gait feeds into mech motion.
Reset @004283b8 -- return every animated joint to neutral, so an abandoned
gait doesn't leave the skeleton frozen on a stale frame.
THE PART THAT CANNOT BE SEEKED. The pose block is packed BY JOINT TYPE -- 8
bytes hinge, 12 ball, 24 ball+translation -- so the root-translation table
behind it is not reachable by arithmetic on any stored count. SelectSequence
must walk the entire skeleton summing per-joint sizes to find it. Which also
means the parse depends on the mech's own skeleton agreeing with the clip: an
unresolvable slot returns NULL and contributes 0, keeping the walk honest
rather than drifting keyframeData onto garbage.
TWO CONTRACTS worth recording before mech2 is written against them:
move_joints == 0 is NOT "do nothing" -- it advances the clock and
accumulates distance while leaving the skeleton alone. That is how the body
channel measures a stride without fighting the leg channel over the same
joints.
The finished callback is RE-ENTRANT BY DESIGN. At end of clip it picks the
next state, re-arms this controller through SelectSequence (rewinding it to
frame 0), and advances the carryover itself -- so its return value is the
distance that carryover covered, folded straight into Advance's own return.
mech2's BodyClipFinished has to honour that or the gait double-counts.
NOT CARRIED OVER from the donor: its BT_HIP_LOG diagnostic and audio
footstep-broadcast path, neither of which is 1995 code. footStepThreshold IS
parsed -- the field is real and authored -- but nothing reads it yet.
AND A CAVEAT ON THE 91% I QUOTED: seqctl.cpp is not in the 50-TU BT census at
all (its code sits below the BT address range the census was built from), and
that figure counts a TU as done if the FILE EXISTS -- 18 files still carry 26
Fail() stubs between them. The census understates what "playable" needs.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
33fd921cea |
BT410 5.3.89: the hit-location cylinder MEASURED -- 18 tables, 8 distinct, and two chassis that never twist
The operator recalled the damage model as "a pie wedged cylinder" and asked how it maps across the mechs. It is exactly that, and the shipped data is now extracted rather than described. dmgscan.py brute-forces every offset in BTL4.RES and accepts a candidate only if the ENTIRE nested type-29 structure parses -- thresholds strictly ascending and terminating at exactly 1.0, zone indices in range, names NUL-terminated. A wrong format guess cannot survive that, so finding exactly 18 tables -- the count DAMAGE-MODEL.md already claimed from an independent reversal -- is a confirmation of the format, not a coincidence. MEASURED: 18 tables, every one 7 bands x 8 wedges = 56 cells. Only EIGHT are distinct by content; the other ten are duplicates. 22 zones 4 twisting x3 Avatar / Mad Cat class -- table A 22 zones 4 twisting x2 Avatar / Mad Cat class -- table B 21 zones 4 twisting x3 Loki 22 zones 4 twisting x2 Thor 21 zones 4 twisting x2 SND2 24 zones 4 twisting x2 Battlemaster / Vulture 20 zones 0 twisting x2 Black Hawk 17 zones 0 twisting x2 Owens BLACK HAWK AND OWENS ROTATE NO BAND WITH THE TORSO. Every other chassis rotates its upper four. That is a real behavioural difference in the shipped data, not an absence of it. The zone COUNTS match the per-chassis .SKL dz_ sets exactly, which is what lets a table be fingerprinted back to a chassis. It is not always unique -- Avatar and Mad Cat share a zone set but have two DIFFERENT tables, and no chassis name sits near the stream, so they are recorded A/B rather than guessed. Stated as undetermined in both the notes and the visual. THE GEOMETRY, now named: 18 wedge names in six anatomical rings (Foot, Leg, Hip, Waist, Chest, Top). Slot 0 starts at angle 0 spanning 45 degrees, so under atan2(z,x) the mech's +X is right and +Z is front. Each named face covers TWO adjacent wedges (Right = 7,0 / Front = 1,2 / Left = 3,4 / Rear = 5,6) -- so dead ahead is the SEAM between two Front cells, never the centre of one. Bands 0-2 are chassis-fixed; the live torso twist is added to the impact angle before the wedge pick on the rest. And the scatter is generous in a way worth knowing at the controls: a clean foot-wedge hit is only 50% that foot, 30% the lower leg, and 20% of the time the OTHER foot entirely. ALSO: an interactive plate of all of it -- every cell of all 8 tables, plan and elevation, and a twist slider that rotates the upper bands live -- published for the playtesters. Its dataset regenerates from dmgscan.py. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f0f0e3d70e |
BT410 5.3.88: hinges flush as MATRICES -- the axis was never on the wire, and a DCS remembers how it was last set
The mech rendered with a live root and frozen limbs. Cause, and it is not in
the renderer:
L4VIDRND.CPP:1028 sets SINGLE_AXIS_HINGE True, so HingeRenderable pushes a
joint with dpl_SetDCSXAxis / YAxis / ZAxis -- three separate libDPL entry
points, each handed only (sine, cosine). The AXIS is carried by WHICH
FUNCTION WAS CALLED, and libDPL never puts it on the wire.
Confirmed against a live capture rather than argued: every articulation
record in joints_check.fifodump is 12 bytes, [handle][sin][cos], 26133 of
them, no axis field anywhere. And the renderer cannot recover it from
context either -- that capture contains ZERO action-0x22 name records, so a
DCS handle cannot be matched back to a .SKL node. vrboard was parsing those
records correctly and dropping them on purpose, with a comment saying so.
THE FIX is the archive's own alternative. The #else half of that same #if
(L4VIDRND.CPP:1153) builds a Quaternion from the Hinge and assigns it over
the DCS matrix -- the axis ends up IN the matrix, and it flushes as a
12-float pose the renderer already applies. That is precisely what
BallJointRenderable::Execute does unconditionally, with no #if at all, which
is why ball joints were never affected. So this is the branch the original
authors wrote and did not take, applied to hinges so they behave like the
ball joints beside them.
A DCS REMEMBERS HOW IT WAS LAST SET -- the part that cost a build.
The first attempt subclassed HingeRenderable and overrode only Execute. It
compiled, ran, and did not work: the wire still carried 12-byte records. What
it DID change was their contents -- the pair went from (sin, cos) to
(0.9990, 0.0000), which is m00 and m01 of the matrix being written. That is
the whole diagnosis in one number. The base ctor calls dpl_SetDCS?Axis before
the subclass gets control, which puts the DCS in single-axis mode, and the
flush then serialises two floats out of it no matter what goes in through
dpl_GetDCSMatrix.
So the axis setter must never touch this DCS. BTL4HingeRenderable therefore
derives from ChildOffsetRenderable, whose ctor builds the offset DCS and calls
no axis setter -- making it the exact hinge analogue of BallJointRenderable:
same base, holds its own attribute pointer plus a previous-value copy, writes
a full matrix in Execute. CODE/ untouched; Component::Execute is virtual
(CMPNNT.HPP:40) so the override dispatches normally, and shadowing 1900 lines
of L4VIDRND.CPP to flip one #define was not needed.
VERIFIED on the rig, torso-sweep conf, clean run:
before 26133 records, ALL 2-float, 1 handle, renderer applied NOTHING
after 7626 records, ALL 12-float, ZERO 2-float messages remaining,
22 handles emitted, renderer applied all 22 (anim_abs)
and the swept joint moves: handle 0x684, m00 across 1532 distinct values
from 0.1582 to 1.0000 -- cos of a 0..80 degree sweep, which is exactly what
BT_FORCE_TORSO drives.
LEG GAIT: NOT FIXED, and the earlier note claiming this would fix it was
wrong. The transport was necessary but not sufficient. On a walking run
(new pod_render_joints.conf = the arena mission with BT_JOINTS) all 22 joints
flush as matrices and the renderer applies them, but only handle 0x672 -- the
vehicle ROOT -- animates. The limbs emit their initial pose and never change.
Because nothing drives them. Joint::SetHinge / SetRotation exist, and the
ONLY caller anywhere in the reconstruction is TorsoSimulation
(BT/TORSO.CPP:382, horizontalJointNode->SetRotation). That is why the
torso is the one thing that moves. MAD.SKL declares JointCount=25 --
jointhip, jointlthigh, jointrthigh, jointtorso, jointshakey, jointeye and
the rest -- and the gait that should walk them is simply not reconstructed
yet. Separate piece of work, now unblocked rather than done.
HARNESS CAVEAT: pose_probe frames on "the last-articulated root" in
anim_abs. With 21 joint DCSs now landing there, that heuristic picks a joint
instead of the vehicle root and frames empty arena. The change invalidated
the assumption; the render is not evidence either way until it takes the root
handle explicitly. The counts above are the evidence.
FAULTS SEEN: two crashes during this work, both cr2=7000FA64 at host 66D9 --
the known parked load-window fault, address unmoved across a relink (which is
the standing test for "not ours"). They died at DIFFERENT points ([mer] 217
and [mer] 116), where a defect in new construction code would die
consistently. Third and fourth runs clean.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
df8c2552d0 |
BT410 5.3.87: OPERATOR-CONFIRMED -- both artifacts gone on a live mission
Full arena1 mission on the pod rig (podrun.sh pod_render_norio, 5.3.83-86 staged), run clean to steady state with no fault. Operator: "the artifact appears to be gone and the bleed from the control mode seems to be gone as well." Both reported symptoms -- the yellow-orange bar on the colour head and the bleed into the MFDs -- from the one uninitialised read, closed by the one zero-fill. WEIGHT OF THE EVIDENCE, stated plainly rather than banked as a win: This is ONE run of an INTERMITTENT, heap-content-dependent artifact. A clean run is exactly what the UNFIXED build also produced on the gauge A/B rig, three boots running, which is the whole reason the fix could not be verified there. So the observation is not logically airtight on its own. What makes it convincing anyway is the shape of the fix rather than the count of runs: after it, no index in 64..255 can return anything but zero, whatever the heap holds. The behaviour is deterministic by construction, where before it was a lottery. Combined with the operator having seen the artifact repeatedly across BOTH this reconstruction and BT411, one clean run on the rig that was showing it is meaningful. Further clean runs strengthen it. A recurrence would mean a SECOND source -- most likely whatever blits the *CRIT/*HEAT silhouettes, still unidentified, still carrying 1500-2500px of index 231 apiece -- and NOT a failed fix. That distinction is worth keeping straight if it ever comes back. Still open and unchanged by this: the 354px runtime-generated cluster, and which widget blits the 231-bearing avatar art. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
902ccf1332 |
BT410 5.3.86: colour-mapper escalation KILLED -- they read palettes, they never blit
5.3.85 raised a hypothesis and explicitly flagged it as unverified: the 1345
colour-mapper entries (656 cmCrit, 662 cmHeat, 27 cmArmor) all sit on port 0,
their mode is ModeSecondaryCritical, and they name adpal.pcc / adpal2.pcc /
heatpal.pcc -- three files the art sweep flags, with ADPAL and ADPAL2 being 2x5
and ENTIRELY index 99. If those pixels were blitted as colour indices, 1345
widgets would each be writing out-of-range on the six-bit head, which would
dwarf the 676px bar and would fire in-mission, exactly where the operator sees
the artifact.
Read the implementation. It does not happen.
ColorMapper::ColorMapper (BTL4GAUG.CPP:389) takes the two names into the
palette8Bin -- not a pixelmap bin, not a bitmap bin -- and the execute path
(:526) is
Palette8 *palette = warehouse->palette8Bin.GetIfAlreadyExists(...);
PaletteTriplet &entry = palette->Color[currentColorIndex];
...
graphicsPort->SetColor(&entry, colorSlot);
It reads the file's VGA PALETTE, lifts one RGB triple, and pushes it into a
hardware colour slot. No pixel of ADPAL ever reaches a framebuffer, so its
index-99 content cannot matter. The palette PAIR is a flash toggle
(paletteToggle alternates per frame), which is also what twoColorMode's
name comparison is for.
NEGATIVE. Recorded as a closed hypothesis rather than deleted, so it does not
get re-raised the next time somebody greps the sweep and sees ADPAL at 100%
out-of-range.
What this leaves: the in-mission exposure, if there is one, is NOT the
mappers. The *CRIT/*HEAT silhouettes still carry 1500-2500px of index 231
each, and which widget blits them is still unestablished. The boot cockpit
does not draw them, which is why it lights only 1030 pixels.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
5a465c2321 |
BT410 5.3.85: the six-bit hazard surveyed -- draw site named, 94 art files inventoried, a 1996 config typo found
Follow-up survey to 5.3.84, written into L4GREND.NOTES.md rather than left as
conversation.
THE DRAW SITE, NAMED. This closes a question that had been open across
several sessions ("which of the three DrawBitMapOpaque sites draws
sentinel-bearing sec art"). Answer: none of them. It is a DrawPixelMap8 --
BTL4GAU2.CPP:1524, the bgPixelMap widget, MoveToAbsolute(0,0) then a
full-image OPAQUE blit. BT_VIS_LOG places it at port=0, the six-bit head.
Every pixel of the colour head's background therefore goes through the
64-entry translation table, sentinels included.
ALSO PROVEN: no drawing operation contains the damage. The opaque inner loop's
Replace case is
*dest = (Word)((*dest & bitmask) | color);
which clears THIS port's bits and then ORs the whole table entry back --
including garbage in other ports' bit positions. Replace leaks exactly like
Or does.
ART INVENTORY. Swept all 428 PCX/PCC under GAUGE (which is all the art there
is -- nothing lives outside it). 94 carry pixels > 63: index 231 in 53 files
(sparse, 4-7%, the 172x217 avatar CRIT/HEAT silhouettes), 255 in 27 (the *HT
heat backgrounds), 226/228/229 as 1-2px specks in QJAK*, and 254 in 2.
BTSEC1.PCX is the only one whose out-of-range region is a SOLID RECTANGLE and
the only one proven to reach the screen -- 676px at (199,526)-(250,538),
rotation 270 mapping screen = (src_y, 479 - src_x), landing exactly on the
measured mask.
WHAT IS STILL OPEN, recorded as open rather than smoothed over:
The 354px cluster. Of the provoke control's 1030 lit pixels, 676 are the
bar; the remaining 33x32 cluster at (67,3)-(99,34) is NOT file-borne -- no
gauge file has a count near 354. So it is generated at runtime. The
obvious suspect CreateMutantPixelmap8 is ruled out: it is an inert stub
returning NULL, so recoloredMech never blits here. Unattributed; the fix
covers it regardless.
A possible second exposure, in-mission. L4GAUGE.CFG carries 1345
colour-mapper entries (656 cmCrit, 662 cmHeat, 27 cmArmor), BT_VIS_LOG puts
ALL of them on port 0, and their mode is ModeSecondaryCritical. They name
adpal/adpal2/heatpal, three of which the sweep flags -- ADPAL and ADPAL2 are
2x5 and ENTIRELY index 99. Whether that matters turns on whether a mapper
blits those pixels or merely reads their VGA palette, which nobody has
checked. Written down as a hypothesis with that caveat attached, NOT as a
finding. If the pixels are used as indices it would dwarf the bar and would
fire in-mission, which is where the operator sees the artifact.
The boot cockpit lighting only 1030 pixels says the avatar/mapper art is not
drawn on THAT screen. It says nothing about a live mission.
A SHIPPED DATA BUG, found in passing and deliberately NOT fixed:
L4GAUGE.CFG:2513-2514 name adpa12.pcc -- digit one for lowercase L. The
file does not exist; adpal2.pcc does. Two of 1345 lines, so the GAUSS and
AFC25 ammo-bin criticals have been loading a missing palette since 1996.
It is the original's bug and the shipped binary lives with it; the archive
is sacred, so we live with it too. Noted so the next person who sees a
warning about adpa12.pcc does not go hunting for a reconstruction defect.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
943219345a |
BT410 5.3.84: the yellow bar and the MFD bleed are ONE bug -- BTSEC1.PCX index 254, and the offender is named
Second operator report, same day: "what is that yellow orange artifact on the
display, it too shows up here and in bt411 but not in the original" -- a
solid vertical bar beside the RANGE readout on the colour head.
It is the same defect as 5.3.83's MFD bleed. Not a similar one: the same
676 pixels.
HOW IT WAS CAUGHT. 5.3.83 shipped a fix that could not be verified, because
the fix makes something INVISIBLE and the A/B rig's boot cockpit already
showed zero MFD extras -- the unfixed build looked perfectly clean. So this
commit adds a POSITIVE CONTROL instead of another argument:
BT_TRANS_PROVOKE fills the colour head's uninitialised
translationTable[64..255] with 0xFF00 -- every high-byte head bit --
rather than the zero the fix installs. Any draw that indexes the tail
then lights ALL the mono heads at once.
On the boot cockpit that lights exactly 1030 pixels: Eng1/Eng2/Eng3 +1030
each, Mfd1 +684, Mfd2 +704, Mfd3 +1074, Comm +676, against 0 extras with the
fix. The bug was firing the whole time. Our heap simply happened to hold
zeros in that tail -- the same luck the shipped binary has been having, which
is exactly why the operator sees the artifact and the A/B rig does not.
THE OFFENDER, NAMED. oormask.py renders which pixels those are and prints
their horizontal run lengths. The mask is a glyph cluster plus 52 runs of
13px -- a SOLID 13x52 BAR at screen (526,228)-(538,279). Solid means a
rectangle in the source art, so decode the art:
BTSEC1.PCX, the colour head's own 480x640 background, contains exactly
ONE out-of-range value in the entire image: index 254, exactly 676
pixels, a solid 52x13 rectangle at (199,526)-(250,538).
The sec port is configured at ROTATION 270, mapping source (x,y) ->
screen (y, 479-x). That puts the rectangle at screen x 526..538,
y 229..280. Measured mask: x 526..538, y 228..279. Same rectangle, to
the pixel.
ONE READ, TWO SYMPTOMS. translationTable[254] is never written (
BuildSecondaryTranslation fills only 1<<numberOfBits = 64 entries), and
DrawPoint ORs the result in unmasked. Whatever the heap left there decides
which symptom the operator sees:
low 6 bits set -> a coloured block on the COLOUR head, beside the RANGE
readout. The yellow-orange bar.
high 8 bits set -> garbage in the MFD / ENG / COMM planes. The bleed.
Both reports, one uninitialised int. It also explains the "not in the
original" asymmetry without needing the original to differ in code: it does
not differ, it is just getting zeros there. And it explains BT411 showing it
too -- both reconstructions inherit the read from the archive.
5.3.83's zero-fill therefore cures both, and turns a heap-lottery into a
guarantee. Still not DIRECTLY observed cured, because no rig we have was
showing the artifact to begin with; that honesty is recorded in the file
header rather than smoothed over.
ALSO IN: oormask.py, barbox.py, vis_provoke.conf, and a README section on
positive controls -- when a fix replaces garbage with a benign value, build
the variant that replaces it with a maximally LOUD value, because that turns
"I see no difference" into a number and separates "the fix works" from "this
screen never exercised the path".
METHOD NOTE, recorded in the README because it nearly cost the fix: three
grabs of one running instance score IDENTICALLY, so the within-boot noise
floor is zero -- and that is the wrong floor. Judging a rebuild needs the
ACROSS-boot floor (~6px on the MFD heads). A single bad grab, caught
mid-draw, read as a 2700px regression and nearly got a correct change
reverted.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
7c3d089d9c |
BT410 5.3.83: MFD-bleed hardening -- the uninitialised half of the 6-bit colour head's translation table
Operator report: "artifacts bleeding over from the radar display into the
MFDs, I have noticed this in the BT411 builds as well, but does not happen
in the original."
MECHANISM (proven from the archive source, not inferred):
The ten cockpit heads are not ten framebuffers -- L4GAUGE.CFG gives each
port a BIT MASK into ONE 16-bit buffer. The colour head is six bits
(sec, 0x003F); the MFDs are bits 8-15 of the same words. A colour-head
defect can therefore only ever surface as MFD garbage, which is exactly
the reported shape.
The two translation-table builders in L4VB16.CPP are asymmetric:
BuildSecondaryTranslation (:5419) writes ONLY the active-bit
combinations -- 1<<numberOfBits entries, so 64 of int
translationTable[256] for this port. Entries 64..255 are never
written.
BuildAuxiliaryTranslation (:5443) nests IncrementInactive inside
IncrementActive and so fills all 256 -- which is why the mono MFD
ports are safe by construction and must NOT be touched.
DrawPixelMap8 (:2708) indexes that table with an 8-BIT source pixel, and
DrawPoint ORs the result in UNMASKED. An out-of-range index therefore
ORs whatever the heap left at translationTable[n] straight into other
ports' planes. The colour head's own art carries out-of-range
sentinels: AVACRIT.PCC is 94% <= 63 with 6% at exactly 231, BTSEC1.PCX
99.8% <= 63 with 0.2% at 254. Transparent draws skip them; OPAQUE draws
put every one through the table.
That it depends on heap CONTENTS rather than drawing logic explains both
the intermittency and why BT411 shows it too: both reconstructions
inherit the same latent read from the archive.
THE CHANGE, via the sanctioned source410-shadows-CODE seam (CODE/ is
untouched): a verbatim copy of CODE/RP/MUNGA_L4/L4GREND.CPP with one
difference -- port construction builds a BTL4GraphicsPort, whose ctor
zeroes translationTable[1<<numberOfBits .. 255] for SECONDARY ports only.
The guard (bit_mask & 0xFF) != 0 is literally the test
BuildSecondaryTranslation itself asserts on. numberOfBits is a popcount
set in the base ctor, so the subclass body sees 6 and the later
SetSecondaryPalette -> BuildSecondaryTranslation refills 0..63 over the
top without disturbing the zeroed tail.
WHAT IS AND IS NOT PROVEN -- the honest scoreboard:
PROVEN non-regressive. Gauge A/B, three boots: unfixed x2 and fixed x1
score identically on every head (Mfd1 132/0, Mfd2 37/7, Mfd3 23/0,
Eng1/2/3 and Comm exact), against a ~6px across-boot noise floor
established by re-booting the same binary.
NOT PROVEN to cure the symptom. The A/B rig captures the boot/attract
cockpit, where the MFD heads already show ZERO extra pixels -- there is
no bleed there to remove. The operator's report is in-mission with the
radar live. This is a hardening change against a real uninitialised
read that matches the symptom exactly; the cure needs a mission-state
capture to confirm.
METHOD NOTE: the first "fixed" capture appeared to regress the MFDs by
~2700px and nearly got the change reverted. It was a bad grab caught
mid-draw -- rebuild-and-rerun matched the unfixed numbers exactly. The
two-agreeing-runs rule earned its keep; a single run would have thrown
away a correct change. Worth noting the noise floor that mattered was
ACROSS boots, not the within-boot floor measured first.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
b366da6f1e |
BT410 5.3.82: the MFD bleed -- mechanism found (uninitialised translation table)
Operator report: artifacts bleed from the colour/radar display into the MFDs, present in BT411 too, absent in the original. Almost certainly the same "intermittent plane-write artifact" that survived seven earlier hypotheses. The chain, every link read from the code: 1. The colour head is SIX BITS (L4GAUGE.CFG:4395, sec mask 0x003F), and the MFDs are bits 8-15 of the same 16-bit words -- one plane-packed framebuffer, which is why a colour-head defect surfaces as MFD garbage. 2. BuildSecondaryTranslation (L4VB16.CPP:5419) fills exactly 64 entries of int translationTable[256]; entries 64..255 are never initialised. 3. DrawPixelMap8 (L4VB16.CPP:2708) indexes that table with an 8-BIT source pixel -- no clamp. 4. DrawPoint/Replace (L4VB16.CPP:637) ORs the colour in UNMASKED; the writer assumes the caller already confined it to the port's bits. 5. The sec port's own art supplies out-of-range indices, and they are SENTINELS not colour data: AVACRIT.PCC is 94% <= 63 with 6% at exactly 231; BTSEC1.PCX is 99.8% <= 63 with 0.2% at 254. Transparent they are skipped; OPAQUE they index the uninitialised tail. Why the original is clean: translationTable lives in a heap-allocated port, and on the pod that allocation was almost certainly zeroed -- garbage reads as colour 0, black, invisible. Our heap has different history. That also explains the intermittency that defeated the earlier hypotheses: it depends on heap contents, not on drawing logic. Narrowed to three opaque call sites (BTL4GAU2.CPP:916/1159, BTL4GAUG.CPP:1953). The fix belongs in our tree via the source410-shadows- CODE rule -- zero the table's tail after BuildSecondaryTranslation -- which is worth doing regardless of which site is guilty, since it turns a heap-dependent artifact into a black pixel. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
ed218483ec |
BT410 5.3.81: the canopy question answered -- our wire is right, the bridge cannot decode it
Decoding the articulated run's fifodump settles it, and the answer is not in our code. 26133 vr_flush_dcs_artic records this run (was ~200), every one n=1 with a 12-byte body: [handle][sine][cosine]. Sample: handle 0x684, sine -0.04536 cosine 0.99897 = -2.60 deg, then -7.60 deg. That is the torso twist sweeping, authentically encoded -- our HingeRenderable emits exactly what the 1995 engine emits. vrboard's 0x1f parser accepts record widths 12, 2 and 5 floats but only applies the 12-float form: "joint sin/cos records: axis semantics unknown -- the flushed matrix stands (mech limbs won't articulate yet)". So the canopy's DCS never rotates in the bridge's scene while its camera picks the twist up by another route, and the two diverge -- the artefact. anim_abs reading 0 in the same run is innocent: the root only re-flushes when localToWorld changes, and this mech is deliberately stationary. The axis genuinely is not on the wire: dpl_SetDCSXAxis/YAxis/ZAxis all take (dcs, sin, cos) and all produce the same record (DPL_VPX.H:23-25), so the board carries it per node. A stream-only decoder cannot recover it. Two ways forward, both legitimate: tell the bridge the axis out of band (the .SKL's Type=hingex/y/z names it per node -- a handle->axis map at scene-build time, and this fixes the leg gait too), or use the engine's other hinge path (HingeRenderable::Execute is written both ways behind #if SINGLE_AXIS_HINGE; the #else branch writes a full matrix, producing the 12-float records the bridge already applies). Both are the shipped engine's own code. The reconstruction's articulation work is done and correct as of 5.3.80; what remains is a decoder gap in the viewing tool. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
07eb74d1f7 |
BT410 5.3.80: joint articulation is live -- 22 nodes, and the twist reaches the board
RecurseSKLFile now builds a joint renderable for any node whose page name resolves to a live skeleton Joint: HingeX/Y/Z -> HingeRenderable (watching Joint::GetHinge), Ball -> BallJointRenderable (watching GetEulerAngles), otherwise the static path. Each holds the rest offset in one DCS and the live rotation in a child DCS, and its Execute diffs the watched value and calls DPL_FLUSH_DCS -- the engine's own mechanism (L4VIDRND.CPP:1026+). The value comes from the mech's JointSubsystem via ResolveJoint, so sim and renderer read one source. Gated on BT_JOINTS while it proves out. [skl] video\max.skl -> 26 nodes, 1 objects, 1 eye, 22 articulated The bridge reported anim_abs=1 joints=0 twist=+0.00 before; it now reports joints=1 twist=-0.86, matching the game's [torso] twist=-0.856. With the mech stationary, frames that differed by 0.0% now differ by 62-80%. A crash it exposed: Mech::ResolveJoint passed segment->GetJointIndex() straight to GetJoint unchecked, and a segment with no joint reports -1 -- GetNthImplementation then indexes [base + -1*4] and dies (guest 00426A1D). Torso never hit it because it only asks for its own authored joint name; the walk asks for every page. Now bounds-checked against GetJointCount. Open: the canopy does not stay rigid in the view, though it and the eye hang off the same articulated node. Cancelling the bridge's cage compensation (CAGE_TWIST_SIGN=0) did not close it. Leading hypothesis: SetupCull builds worldToEyeMatrix from GetSegmentToWorld(siteeyepoint) -- the SIMULATION's segment transform -- independent of the render tree, so the canopy follows our render chain and the eye follows the sim's, and they diverge whenever one carries the twist and the other does not. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8769dae0ac |
BT410 5.3.79: the twist is live in the sim and absent from the wire
pod_render_twist.conf isolates it -- throttle and turn zero, torso sweeping, [sim] confirming spd=0, so only the torso joint can move the view. Five frames across a full sweep differ by 0-74 pixels: the view does not move. The same run's telemetry has the twist sweeping the full authored range (-0.895 to -0.022 rad, cmdL/cmdR alternating, joint resolved) and the bridge reports anim_abs=1, joints=0 -- only the root articulates. The cause: RecurseSKLFile builds each skeleton node's DCS with a baked matrix and flushes it once at build time, and nothing re-flushes a node when its joint angle changes. TorsoSimulation faithfully calls SetRotation() every frame, updating the Joint object, but no renderable reads that back into the node's dpl_DCS. The simulation twists; the board never hears. The missing piece is the per-frame joint renderable -- the engine's ChildOffsetRenderable family (L4VIDRND.CPP:930) exists for exactly this, and our RootRenderable already proves the pattern for the hull. That one brick unlocks the torso twist (and the cockpit eye panning with it, since the eye hangs off that chain), the leg gait, and weapon-pod aim. Until then the mech translates through the world as a rigid body -- which is what every frame so far has shown. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
3c289ff10a |
BT410 5.3.78: the torso twists from the button path; two buffer overruns fixed
BT_FORCE_TORSO sweeps the twist BUTTONS -- the members the streamed direct mappings write -- so the chain runs without a hand on the RIO. Live on the pod: twist sweeps -0.015 to -0.891 rad and back as cmdL/cmdR alternate, at the resource rate 0.873 rad/s, inside the authored +/-2.443 limits, with jointtorso resolved. End to end: attribute id -> direct-mapping destination -> TorsoSimulation integrator -> skeleton joint -> the canopy eye hanging off that chain. 5.3.68's table and 5.3.71's integrator both confirmed against real button semantics. Two buffer overruns, one of them mine. Reading TORSO.CPP back caught the ctor still zeroing dynamicsState[22] after I shrank the array to [16] by carving the command members out of its front: 24 bytes past the end of every Torso, on the heap, every mission, since 5.3.68. A sweep for the same shape across BT/BT_L4/MUNGA found a pre-existing one -- MECHMPPR.CPP zeroing reserved[22] against reserved[21]. Both now bound by ELEMENTS(). The rule for this tree: any array carved out of a reserve block must have its initialiser bound by ELEMENTS, because the carve is exactly the edit that silently invalidates a literal. Neither explains the residual host fault (it predates 5.3.68 at 3/3), but both were real corruption running in every mission -- and the Torso one was introduced by the very fix that cut the fault rate, which is worth remembering whenever a rate MOVES instead of going to zero. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e69d761fb8 |
BT410 5.3.77: the overrun drain loop is not the trigger either
serialnamedpipe's P_RX_BLOCKED path drains its whole backlog when the guest
hasn't read in time -- directserial's `while (doReceive());`. That looked
like the trigger, since a named pipe's backlog is unbounded where a real
port's is capped by the line rate. Bounded it to one byte (real UART overrun
semantics) and re-ran the conf that had faulted 3/3: FAULT 283s, FAULT 226s,
FAULT 204s. Killed.
The measurement that explains why it was never plausible: overruns during a
run are 1-3 per report period, because vRIO sends about a byte every 1-3ms --
the backlog is shallow and the loop had nothing to teleport. The 494/499
counts all land after the game exits. I read the code and inferred a burst
without measuring the queue depth it operates on.
Default restored to drain-all: it is the validated directserial behaviour the
RIO's rxpollus/rxburst tuning was calibrated against, and changing it on a
dead hypothesis would risk real-cockpit timing for nothing.
VPX_RX_OVERRUN_ONE=1 opts into the bounded form.
The roadmap now carries a fault ledger of every dead hypothesis so none get
re-run. What survives: a deterministic DPMI-host path walking a
{next,handler} chain into a node whose pointer is an unhooked IVT value,
entered under RIO interrupt load. Naming the owning routine needs a trace of
entries to host 0x66CF, clean vs faulting -- a DPMI32VM reversing
sub-project, now decoupled from the reconstruction (vRIO down = 100%
reliable, logged conf = ~70%).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d378148c14 |
BT410 5.3.76: THE COCKPIT, FROM THE PILOT'S SEAT
cockpit_max.png: the MAX_COP canopy seen from the eyepoint on the live pod --
window frames, A-pillars, the sill across the bottom, and the arena city out
through the glass at 30fps. The run's log line:
[skl] video\max.skl -> 26 nodes, 1 objects, 1 eye
Exactly the predicted shape: the X-variant carries the same 26-node chain as
the body skeleton, so eye composition and torso twist are identical, with ONE
object -- the canopy shell -- instead of nineteen body parts, and the eye on
jointeye inside it.
This closes the arc that began with a mech statue at the world origin and a
camera frozen beside it: entity -> RootRenderable -> skeleton under its DCS ->
cockpit variant for the inside view -> eye on the canopy joint. What a pilot
saw in 1996, rendered by the reconstruction through the emulated Division
board.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
edb1670900 |
BT410 5.3.75: the quiet conf faults 3/3 -- logging SUPPRESSES the fault
Hypothesis: our BT_* logging inflates the load window, the window is the exposure, so a quiet run should load faster and fault less. Both halves wrong. pod_render_quiet -- which differs from pod_render_rec by exactly the six BT_* logging vars -- faulted 3/3, same signature, against ~25-40% for the logged conf. That is the strongest evidence yet that the residual is a timing-sensitive HOST bug rather than anything about our data: writing to COM3 changes when interrupts land relative to the DPMI host's handler-chain walk, and nothing about our attribute tables or renderables could plausibly be modulated by whether we print to a serial port. Interrupt phase can. Two immediate consequences: pod_render_rec is the LESS fault-prone rig, so use it; and the 'shipped survives' baseline weakens further, since shipped also writes plenty to COM3 and so sits on the suppressed side of the same effect on top of its shorter window. The wall-clock numbers also correct an earlier assumption: 150-275s boot+load, not ~60s. The 60s figure was the load DRAIN in game ticks; DOS boot, two diagnose passes and card init dominate the rest. Next is clearly emulator-side: the host keeps no EBP frames (the frame walk returned one bogus entry), so naming the routine that owns the chain needs a breakpoint-style trace of entries to 0x66CF with the head node, diffing a clean run against a faulting one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e97862507c |
BT410 5.3.74: the faulting instruction decoded -- a handler-chain walk with a bad node
The SegPhys-corrected dump read real code at last (the earlier zeros were the
probe reading EIP as a bare linear address; CS=00FF has a base):
66D4 push 0 / push edi / push ebx / push esi
66D9 call dword near [ebx+4] <- faults
66DC cmp eax,0 / jz done
66E5 mov ebx,[ebx] / jmp loop
A linked-list walk with a per-node callback -- node = {next@+0, handler@+4},
four args, "handled" on nonzero. A DPMI exception/interrupt handler chain.
The arithmetic closes exactly: the read is DS_base + ebx + 4, and
7000FA64 - F000CA64 = 80003000, so DS base is 0x80003000 and the famous cr2
is that wrap. My earlier '[EBX+0x3004]' reading was fabricated from the cr2
alone -- there is no 0x3004 displacement, which is why the constant appeared
in no binary.
So: a chain node's NEXT pointer holds F000CA60, the value parked in unhooked
IVT slots -- the walk expected a terminator and got an interrupt vector, then
read its +4 as a function pointer. Identical registers on every catch, so
one deterministic path, entered often enough that our 60s load window catches
it 25-40% of the time while shipped's 15s window mostly does not.
Open: which routine owns the chain, and where the bogus next came from. The
probe now walks the EBP frame chain to name the caller.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
2ce3528f1c |
BT410 5.3.73: the bad pointer is the emulator's default-handler stub
The IVT scan on the second probe catch settles EBX: F000:CA60 is the emulator's default unhandled-interrupt callback -- the value filling every vector nobody hooked (it matched ~56 of them), while the hardware-IRQ block IVT[08-0F] holds the DPMI host's own 0D1F reflector stubs. The serial IRQs are properly hooked; the host read one of the UNHOOKED vectors' default value and probed [value + 0x3004] as a flat pointer, wrapping to the famous cr2. Also corrected: the first dump read ESP/EIP as bare linear addresses -- the 'stack' bytes were DOS conventional memory (they decode as 16-bit real-mode code), since SS=0107 has a nonzero base. SegPhys-corrected dumps are built and armed for the next catch. The zeros at the faulting EIP are consistent with DPMI32VM being a virtual-memory host that pages its own regions -- 'Reference to a page you don't own' is its pager's abort message for an address outside every region it owns. Nailed: deterministic host-side path; the serial stream correlates as exposure, not as trigger bytes; the bad read is [unhooked-vector-default + 0x3004]. Open: which vector, which host routine -- the corrected dumps should name the return chain on the next catch. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f2687a3146 |
BT410 5.3.72: the fault caught in the act; the inside view gets its cockpit
VPX_PF_WATCH (fork, cpu/paging.cpp): on a guest page fault at the watched linear address, dump guest registers, the last 64 serial RX deliveries with guest cs:eip at each, the code bytes at the faulting EIP and the stack top. serialnamedpipe's doReceive feeds the ring. Armed in podrun.sh and launch_pod.ps1. The first catch decoded the residual fault completely: CS:EIP 00FF:000066D4 in the DPMI host, EBX = F000CA60 -- an IVT entry read as a dword, segment F000 offset CA60, a BIOS default interrupt handler -- and the faulting access is [EBX+0x3004], whose 0x80000000 segment-base wrap gives exactly cr2 7000FA64. The host probes a word 0x3004 bytes past a real-mode vector value treated as a flat pointer: harmless for its own low-memory handlers, a fault for BIOS F000:xxxx defaults. The serial ring shows a steady 1-byte/1-3ms vRIO stream with nothing special at the fault -- the stream determines which vectors get walked, not the crash itself. The 0x3004 appears nowhere in DPMI32VM.OVL or 32RTM.EXE as an immediate, so the probe now also dumps code bytes at EIP; faulthunt.sh loops runs until the next catch. Shipped baseline streak: 4/4 clean -- consistent with exposure, not yet discriminating. The inside view now loads the COCKPIT skeleton: the fleet-wide X-variant naming convention (MAD->MAX etc., all 64 skeletons present) selects the same 25-joint chain with a single object -- max_cop.bgf, the MAX_COP canopy shell with the PUNCH-texel windows from the capture forensics. The donor names the same mechanism from the decomp side (inside = SkeletonType_A with '_cop' selection). Fallback to the body skeleton when no X file exists. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
bd1070110b |
BT410 5.3.71: torso button commands wired into the sim; residual fault rate measured
TorsoSimulation now carries the donor's digital command handling (binary @004b5cf0): elevate up/down, twist left/right with limit clamps, the centre-button recenter latch slewing home across frames, and the ramp machinery kept verbatim including the binary's punchline -- the shipped build unconditionally overwrites the ramp with 1.0f, authored dead weight preserved as the binary's shape. recenterActive carved from the dynamicsState reserve. With 5.3.68's table fix the RIO/TM twist buttons land in these members and move the torso, and the canopy eye rides the twist chain. Also corrected in the roadmap: the 5.3.70 TriggerState 'wart' was not a wart -- CheckFireEdge already compares fireImpulse's bit pattern as a signed int, sign-correct for button ints and floats alike, and the donor binds and types the member identically. Fault rate, measured: riostreak.sh ran four consecutive vRIO-live runs -- CLEAN/CLEAN/FAULT/CLEAN, faults only ever in the load window (ticks 42-53k), clean runs always past launch. Post-fix tally 8 runs / 2 faults (~25%; was 3/3 before the Torso fix). Our load window is ~60s where shipped's is ~15s: if the residual is a constant-rate hazard during load, shipped's expected per-run rate is only ~7% -- 'shipped survives' may be exposure luck, and the residual is most likely an emulator-fork serial/DPMI interaction rather than a game defect. Next probe: hook the fork's page-fault path on the fixed cr2 and dump recent serial-IRQ deliveries + guest cs:ip history, plus a shipped-exe streak for the baseline. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8b9c2e6b11 |
BT410 5.3.70: the full direct-mapping contract audited; the fault is not fully dead
The canopy-eye confirming run (5th post-Torso-fix run with vRIO live) faulted with the old signature during the load drain -- four clean runs then one fault. The two-runs rule keeps earning its keep. The BT_MAP_LOG audit now prints the whole resolved contract, not just NULLs. The L4 list carries 13 direct mappings: mapper stick/throttle/reverse/looks (ids 3/4/6/10-12), Torso torsoCenter (id 14, matching the donor), and TriggerState (id 19 = the known 0x13 binding) on six weapons. Every destination resolves and every size is right -- the 8-byte joystick write lands in an 8-byte ControlsJoystick, buttons are ints into ints. One type wart noted, not fixed: MechWeapon binds TriggerState onto fireImpulse, a Scalar -- a ButtonGroup direct writes an int bit-pattern, so a real trigger press stores 1.4e-45f. Size-safe, but the fire FSM will never read a real button press as a pull; fire worked in tests only via the BT_FORCE_FIRE hook, which writes a proper float. The RIO trigger needs an int home for TriggerState. Surviving fault theories, in current order: an emulator/DPMI-host serial-IRQ interaction whose probability guest timing merely modulates (the fault's constants -- fixed host EIP, fixed cr2, always during the load drain -- fit this better than the corruption story ever did, since the smash point never moved across builds where our BSS moved); a second corruption source outside the direct mappings; sampling noise. riostreak.sh is measuring the post-fix rate over four consecutive runs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d4f105f42d |
BT410 5.3.69: the cockpit eye sits in the canopy -- site handling in the walk
RecurseSKLFile now walks "site=" children. Sites get no draw component -- which is why the walk's 26 nodes matched the real capture all along: sites were never DCS nodes -- and siteeyepoint spawns the cockpit camera with the donor's exact construction (bt411 btl4vid.cpp:462, decomp FUN_004579a8): offset = the site's own local rest transform, parent = the site's PARENT joint's DCS, not the hull root. World orientation, torso twist and gait all arrive through chain composition. The root-DCS eye remains only as the zero-construction fallback for skeletons with no siteeyepoint. Measured on the live pod with the RIO streaming: [eye] cockpit eye on 'jointeye', 26 nodes / 19 objects / 1 eye, and the bridge camera sits at Y = hull + 7.7 -- the exact chain sum jointlocal 5.29 + hip 0.37 + torso 0.31 + jointeye 1.69. canopy_eye1.png is the pilot's view, the mech's own nose wedge visible below the sightline. Own-mech cage handling is the next layer. Also: eye_count threading through ReadSKLFile/RecurseSKLFile, and the [skl] summary now reports it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
8231bee58c |
BT410 5.3.68: THE LIVE-RIO FAULT IS FIXED -- it and the TM crash were one defect
The [map] audit named it in one run: subsystem 17 = Torso, attribute ids 12/13 unresolved. The donor's decompiled torso.hpp carries the authentic enum with binary offsets -- ids 3..15, including StickPosition(9), TorsoUp/Down/ Left/Right(10-13), TorsoCenter(14), MotionState(15). Our table stopped at seven entries with MotionState at id 9, on the strength of a comment claiming the rest were messages. The streamed control mappings bind BY ID as direct WRITE destinations, so the truncation produced two different crashes from one cause: in TM mode the button ids 12/13 resolved NULL (the boot write-fault the moment the joystick polled), and in RIO mode the analog id 9 resolved to our motionState StateIndicator -- every analog packet from a live vRIO wrote raw floats over a watcher-socketed object, and the corrupted chains walked into unmapped memory ~30-40s later. That was the 'intermittent' pod fault at the fixed DPMI-host address. vRIO down = no analog = no corruption, which is why the norio conf launched: every observation from the whole hunt drops out of this mechanism. Fix: the full 13-entry donor table; five int command members carved from the dynamicsState reserve (class size unchanged); ctor zeros them. Verified, two agreeing runs each way: TM mode launches with a clean audit and the Thrustmaster joystick driving the torso (stickY=0.907 -> torsoElev 0.349); RIO mode with vRIO STREAMING launches, zero faults, weapons cycling. pod_render_rec is no longer poisoned and the norio workaround is obsolete. The audit-before-the-archive-call pattern (BT_MAP_LOG walks the same streamed table CreateStreamedMappings consumes, naming what will not resolve) turned a two-day intermittent-corruption hunt into a one-run lookup. The streamed RES tables are binding contracts on our attribute enums. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f30acc6a94 |
BT410 5.3.67: pose_chase was upside down -- harness bug, wire exonerated
The operator caught it: the 5.3.65 image showed the mech hanging from its own ground-shadow plate (-y is up in the bridge world). The follow-up pinned down what was and wasn't wrong, and the reconstruction survives untouched. Right, by three independent wire-level proofs: MAD.SKL offsets are a coherent Y-up model (hip +5.29, knees descending, torso ascending -- anatomical only read +Y-up); the BGF vertex extents agree (foot meshes extend -1.0 below their ankle to the sole, torso +3.8 up to the missile pods); and the SHIPPED pod's wire carries the same signs (hip +6.21, knees negative, vehicle root a pure yaw with det +1). Our wire and shipped's are convention-identical -- no flip exists between game and board in either game. Wrong: the ad-hoc chase camera. It borrowed the aerial up-hint from the calibrated live-bridge world (where -y is up) and applied it to a raw scene-space render of Y-up wire data; the symmetric checkerboard ground disguised the inversion. The shipped capture rendered upside-down through the same harness, which is what cleared the wire. Both conventions now documented in the roadmap: wire/scene space is Y-up in both games; the live bridge presents -y-up via its calibrated mirror family (FP_RIGHT_SIGN / CAGE_TWIST_SIGN, COCKPIT-CAGE-NOTES.md). Tooling: pose_probe.py replaces the throwaway harness (updir=+1 scene space default, -1 for the bridge sense). pose_chase_upfixed.png is the corrected milestone image: the MadCat right side up -- chicken-walker legs, purple knee actuators, feet planted, shadow under the feet. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
1233a1d5f6 |
BT410 5.3.66: RIO-fault eliminations, and TM mode's own NULL-deref surfaced
Three clean eliminations on the live-RIO fault, all on the ack-fixed emulator: not the ack heuristic (reproduces after the fix), not the debug logging (pod_render_quiet faults identically with no BT_* env), and not the failed test-mode handshake (the quiet run's init PASSED and it faulted anyway). The Thrustmaster isolator -- L4CONTROLS=THRUSTMASTER so the game never opens the RIO while vRIO keeps streaming into serial1 -- was blocked by its own discovery: TM mode dies at boot in ControlsInstanceDirectOf<int>::Update+0x6, mov eax,[eax+0x1c] with EAX=0, at guest 00485E1A. That address resolves cleanly in our map and the loader names BTL4REC.EXE CODE 0x75E1A directly: a real defect in our controls wiring for the non-RIO path, and its own reconstruction brick. Standing facts for the hunt: our exe + live vRIO = fault in ~30-40s; shipped exe + the same stream on the same emulator = fine (two baselines). The ISR and LBE4ControlsManager are archive code; ours in the per-event path are the mapper (MECHMPPR) and the event/receiver plumbing. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
e846717283 |
BT410 5.3.65: THE MECH STANDS -- wire matrix convention proven from engine source
pose_chase.png: the MadCat assembled from its 19 parts, standing in the arena at its live articulated position. Feet, legs, torso, weapon pods, canopy plate. Not a collapsed stack at the origin, which is what every prior frame actually showed. The convention is no longer inferred -- it is read from the engine. RootRenderable's ctor writes an entity pose into a DCS with *(Matrix4x4*)dpl_GetDCSMatrix(myDCS) = localToWorld, so Matrix4x4::operator=(const AffineMatrix&) IS the wire layout: row-major, rotation in rows 0-2, zeros in column 3, translation in ROW 3 (entries 12/13/14). The old 3/7/11 choice rested on two errors, both corrected in the code comments: AffineMatrix does keep translation at 3/7/11, but because it is COLUMN-major -- citing it as precedent for a row-major layout read the storage and ignored the indexer (AFFNMTRX.HPP:99). And 'the last row made the maths blow up' came from the contaminated bisect; those crashes were the RIO fault. Rotation now composes through the engine's own Matrix4x4::operator=(const EulerAngles&) from the page's pitch/yaw/roll, and the node write uses the RootRenderable idiom (dpl_GetDCSMatrix + assign). Also banked: the chase-frame fifobridge variant -- camera framed on the articulated root instead of the arena bounds -- as the offline proof harness for any future pose question. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d143f833ab |
BT410 5.3.64: THE COCKPIT EYE IS LIVE -- root+eye composition, and the emulator
deadlock it exposed The camera gap root-caused and fixed through three layers: LAYER 1 (ours): "chain the engine after the skeleton" was a non-fix. The engine's MakeEntityRenderables only accepts Object/Rubble resources, so chaining it with a Skeleton printed "wrong video resource type" and built NOTHING -- the fifodump proved it: zero vr_flush_dcs_artic records. The real fix mirrors the engine's Mover composition (L4VIDEO.CPP:4795-4860) inside our mech case: a Dynamic RootRenderable (whose ctor seeds its DCS from localToWorld and whose Execute re-flushes on change -- the ONLY source of per-frame wire articulation), the skeleton hung UNDER its DCS via ReadSKLFile's new parent_dcs parameter, and for the inside view a DPLEyeRenderable on that root with the published EyepointRotation. Also explains the mystery monolith in the first frame: the skeleton was parked at the world origin with the camera inside its shins. LAYER 2 (1995 library, read from our own linked symbols): dpl_DrawSceneComplete = velocirender_frameack(0), and frameack's unsolicited-reply path prints "dpl error - unsolicited input during frame ack", adds the 1995 authors' own puzzled "flush artic??" when the stray action is 0x1f, and calls exit(9). Documented before it ever fires. LAYER 3 (emulator, the actual deadlock): with per-frame articulation flowing for the first time, the game stopped drawing at frame 232 -- one 0x1f burst per frame forever, no receives, no error. The VPX device feeds the frame ack only after 6 CONSECUTIVE empty polls and reset that counter on EVERY outputData write; per-frame articulation writes made the count unreachable, so dpl_DrawSceneComplete never went true. Fix: the reset is gated on !frame_outstanding -- once a draw is outstanding the only receive the game will do next is the frame ack. The iserver-drain concern the reset guarded is boot-time only, when no frame is outstanding. Verified: two agreeing runs of ours on the fixed emulator (611 draws / 222 artic batches in run 2 -- the 232 wall is gone), camera travelling with the mech (cam -329.7,0,60.5 -> -362.3,0,54.5 across 8s), anim_abs 0 -> 1 on the bridge, view now INSIDE the cockpit cage. Shipped binary re-run on the fixed emulator: launches and runs, no regression. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
af010c18c2 |
BT410 5.3.63: the mech gets both skeleton and engine renderables
MakeEntityRenderables no longer stops after the skeleton. It builds ours AND
chains DPLRenderer, which is what creates the RootRenderable and the
DPLEyeRenderable the camera follows. Verified over three runs:
[mer] 208 class=3001 view=1 res=1
[vid] type=1 file='mad.skl'
[skl] video\mad.skl -> 26 nodes, 19 objects
[chain] into DPLRenderer, skeleton=yes, stack at 0x12dde7
No fault, mission launches, sim runs. Worth noting separately: the engine's
eyepoint construction -- prime suspect for the page fault through most of a
session -- runs perfectly here, with the stack at 0x12dde7, right beside the
main loop's frame. Both are consistent with the corrected finding that the
fault is a live-RIO artifact at a fixed DPMI host address.
The camera still does not move: two frames six seconds apart differ by zero
pixels while [sim] shows the mech travelling from (487,20,288) to (493,20,301),
and the bridge reads cam (0.0, 10.0, 0.0) -- world origin plus eye height. So
the eye renderable was necessary but not sufficient.
Eliminated by reading our own code rather than guessing: the viewpoint entity
IS the mech (BTL4APP.CPP:354), the video renderer IS linked to it
(APP.CPP:1337), and SetupCull would Fail() loudly if the mech lacked
"siteeyepoint" or EyepointRotation -- neither Fail appears. That leaves the
path from SetupCull's worldToEyeMatrix to what the bridge reads off the wire,
which is bridge territory rather than reconstruction territory.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
3c42766f48 |
BT410 5.3.62: correction -- the fault address was never in our code
Every conclusion that named EulerAngles::operator=(const LinearMatrix&) as the
fault site was wrong, including the disassembly built on top of it. Three
independent proofs:
1. ROTATION.CPP is ours, so the function was instrumented directly -- print
this, &matrix and a local's address on the first twelve calls plus any wild
pointer. The run faulted with ZERO [euler] lines: never called.
2. After the probe was added, 0x66D9 disassembles to a two-byte conditional
jump (jnl) that touches no memory. The dump reports ErrCode 0004, a data
READ.
3. EIP 0x66D9 and cr2 0x7000FA64 are identical to the byte across every build
this session, including builds where the code at that offset changed
completely. A fault in our CODE segment moves when the code moves. This
one does not -- it lives in the DPMI host, which loads low and never gets
relinked.
The map lookup was a coincidence: our _TEXT is section-relative from 0, so its
offsets overlap the host's low addresses, and the map will name a function for
any small number. Before resolving a fault address against the map, confirm it
belongs to our segment -- cheapest check is whether it moves on relink.
What is actually true, by same-binary A/B:
vRIO DOWN, serial1 still a namedpipe -> mission LAUNCHES
vRIO UP, same conf, same binary -> FAULT
The trigger is a LIVE RIO stream, not an unplugged cable. The emulator reports
continuous serial1 RX overruns, the fault lands wherever the app happened to be
(terrain one run, the mech's inside view another), and the shipped binary
survives the identical conditions. A fault at a fixed host address, at
arbitrary points in the app, requiring an interrupt-driven stream, is an
interrupt re-entrancy signature rather than a wild pointer.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
d5ae9c4851 |
BT410 5.3.61: the camera never moves -- our mech case skips the engine's eyepoint
Two frames captured seconds apart are pixel-identical while [sim] shows the mech travelling from (180.4, 10, -358.7) to (-604.8, 0, -352.4), with the bridge title reading cam (0.0, 10.0, 0.0) throughout. The scene renders, the eyepoint is never driven. Our own MakeEntityRenderables explains it. When ReadSKLFile finds a Skeleton entry it sets handled = True and the chain to DPLRenderer::MakeEntityRenderables is skipped -- but that is the call which builds the RootRenderable and the DPLEyeRenderable the camera follows (L4VIDEO.CPP:4548-4566). We hand the board a skeleton with no root and no eye. That also closes the loop on the fault: the eyepoint construction we skip is the same code that faults when it runs. The mech case must end up doing BOTH, so the EulerAngles::operator=(const LinearMatrix&) fault has to be understood before the mech can be seen from its own cockpit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
9c959e6891 |
BT410 5.3.60: first rendered frame from a running reconstructed mission
emulator/render-bridge/first_mission_frame.png -- arena city, textured buildings, horizon, 28fps, captured from the GL bridge while our build ran a live mission with the mech walking. The skeleton went in with it: [skl] video\mad.skl -> 26 nodes, 19 objects. Corrects yesterday's RIO framing. I called serial1 an unplugged cable. Wrong twice: VRio.App is running on this rig (the dev-rig default since 7/17), and the emulator log recorded real byte counts on that port -- RX overruns of 471, 503, 528 -- which an unconnected pipe cannot produce. The port was LIVE and streaming during every crashing run. So the defect is not that we mishandle a disconnected port, it is that we mishandle a live RIO stream, which is worse: every production pod has a real RIO. The shipped binary on the same rig logs 'lost RIO analog request', keeps sending characters, and runs the mission; ours logs 'RIO never came back from test mode!' and dies partway through the load. The overrun counts say the guest is not draining the port fast enough. L4RIO.CPP's handshake is authentic archive code, but MechRIOMapper (BT/MECHMPPR.CPP) is ours -- that is the thread. pod_render_norio.conf stays the way to get a running mission today, but it is a workaround: it disables the cockpit's primary input device. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
475ae3106e |
BT410 5.3.59: A MISSION RUNS -- the RIO serial port was blocking every load
pod_render_norio.conf is pod_render_rec.conf with one line changed, serial1=disabled instead of the vrio named pipe. Same binary, same egg. With it, our build launches and the mech walks: [launch] state=2 minPriorityEmpty=1 ticks=85316 [queues] p0=- p1=- p2=BUSY p3=- p4=- BTL4Application::RunMissionMessageHandler Turning Plasma Score Display On [sim] pos=(180.466,10,-358.703) yaw=-0.275605 spd=14.4 [sim] pos=(251.959,10,-250.498) yaw=-0.89103 spd=14.4 No fault, 290 log lines and counting. This is the first time the reconstruction has reached a running mission on the pod. The hang was never starvation, a deadlock, or a refill loop -- it was VOLUME, the first candidate I listed and then talked myself out of. 530 renderer events queue at priority 0 during load; BackgroundTasks::Execute runs exactly one task per call round-robin over seven tasks, and the 1ms frame budget is always blown here so no extra background passes happen. That is ~9 events per real second, so a load legitimately takes ~60s. The fault arrived at ~40s, before the backlog could clear. Every 'the queue never drains' reading was really 'the process dies before it can'. The disproof was already in hand: the post tally climbed to 530 and went FLAT, which means nothing was refilling. p0=BUSY with nextReady=1 is equally consistent with a large finite backlog still draining, and I read it as refill. Why the RIO port: the pipe has no server attached, which should behave as an unplugged cable. The emulator log shows steady serial1 RX overruns, our build prints 'RIO never came back from test mode!', and the SHIPPED binary prints 'lost RIO analog request' and launches anyway on this identical rig. So an unplugged RIO is survivable and our handling of it is not -- a real defect, not a rig artifact, since a pod with an unplugged RIO cable should still boot. The wild reference at EulerAngles::operator=(const LinearMatrix&)+0x19 is still unexplained, but it is now reproducible on demand by enabling serial1 rather than being a coin flip. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d3b332a62a |
BT410 5.3.58: the load gate is blocked by a REFILL loop on priority 0, not starvation
Per-priority occupancy, measured every second until the fault: [queues] p0=BUSY p1=- p2=BUSY p3=- p4=- nextReady=1 That single line kills the two leading theories. p3/p4 empty means nothing above priority 0 is competing, so the pump is free to serve it -- no starvation. nextReady=1 means PeekAtNextEvent always has a READY event, so priority 0 is not holding a timed event whose alarm never comes due -- no deadlock. What remains is refill: priority 0 is replenished as fast as the pump drains it, ~143 events per simulated second. It is also fully deterministic. Two runs of the same binary reported pump=352/1499 at frame 1001 and 495/2643 at frame 2002, identical to the byte. So 'crashes about half the time' was never a race -- it was comparing runs that differed in binary or conf. Only InterestManager::PostRendererEvent posts at priority 0 (it maps every renderer event there while the app is not RunningMission), so naming the message names the flooder. Application::Post now tallies priority-0 posts by message ID and reports the busiest three. Notably, all three producers of NotifyOfNewInterestingEntity are entity-CREATION paths, so if that ID dominates then something is creating entities in a loop -- which would explain the unbounded growth behind the fault as well as the hang. pod_render_noskl.conf now carries the same probes, so one instrumented binary can be run with the skeleton walk on and off. Walk-off runs are the only configuration of ours known to reach 'Turning Plasma Score Display On'. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4cf09917ce |
BT410 5.3.57: the crash is deterministic, and the real blocker is the load gate
The previous commit's attribution was wrong and is corrected in the roadmap. WRONG: 'the crash is the skeleton walk's' rested on 0 crashes in 2 runs with the walk disabled -- a 25% coin flip presented as evidence. The oldest preserved dump has no [skl] line and ends on the old 'couldn't figure out how to MakeEntityRenderables' fallback, so it predates the walk and faults identically. WRONG: 'it lands at different points each time'. Every dump carries the same numbers to the byte (cr2 7000FA64, EIP 66D9). What varies is how far the log gets, not where the fault is. Resolving the address through btl4opt.map names it exactly: EulerAngles::operator=(const LinearMatrix&)+0x19, a read through the matrix reference. A binary scan finds all three call sites of that operator pass a stack local, so the 1.79GB pointer still has no static explanation -- that is now its own open item rather than a guess. Two theories killed cleanly by the new BT_STACK_LOG probe and a binary scan: ESP drift (drift=0 over 2002 frames; exactly one callee-cleans function exists in the whole binary) and an undersized stack (our PE and the shipped one have identical 1MB/8K geometry). What the probe found matters more: the application is parked in state=2, which is LoadingMission, not WaitingForLaunch. Priority 0 is where the interest manager queues renderer events during load, the gate needs that priority empty, and it never empties -- so the mission never launches and the renderer holds a blank screen by design. The fault arrives ~30s into that wait. Runs that DO launch never print a single [launch] line. So 'crashes half the time' and 'hangs during load' are one event seen twice. A 15x disagreement between two clocks looked like a reconstruction slip -- BTL4.CPP passes GetTicksPerSecond() where ApplicationManager wants a frame rate. Checked against the shipped binary before touching it: same instruction sequence, same kind of static float pushed. Authentic. Documented so nobody 'fixes' it. Also swept every subsystem DefaultData against its real C++ base. Fifteen chain past their immediate parent, but fourteen skip only classes that add no handlers and no attributes, and no class's attribute-ID base disagrees with its index chain -- so there are no gap slots there. The one real defect: Generator is a HeatSink but chained to Subsystem::MessageHandlers, so it ignored every ToggleCooling message. Fixed (compiles next build). Tooling: podrun.sh stages the build over BTL4REC.EXE, exports the host-side VPX board env -- without which the run dies at the iserver handshake rather than merely rendering nothing -- and archives every run's log, marking it -CRASH when it faulted. Before this the only preserved dump was an accident, on a rig where each run costs four minutes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
f54e1a3c19 |
BT410 5.3.56: the intermittent pod crash is the skeleton walk's -- attributed by same-binary A/B
Added BT_NO_SKL, which skips the skeleton build so ONE executable can be run with and without it. That is the only clean way to test an intermittent fault, and it settles attribution: walk enabled ~50% of runs die (page fault at 66D9) at DIFFERENT points walk disabled 0 crashes in 2 runs shipped exe 0 crashes in 2 runs, same rig and conf So it is our new code, not the rig, and the varying fault point points at memory corruption surfacing later rather than a bad instruction in the walk. Hypotheses killed by reading the private library headers: the matrix array is the right size (dpl_MATRIX is float32[4][4]); and neither s_dplobject nor the common dpl_node header retains an object name, weakening the dangling-string theory (a private name cache inside the library is still possible and is the best surviving suspect). Also checked something that would have been much worse than a crash: ~NotationFile REWRITES its file when dirty, and these are the game's shipped .SKL data files. MAD.SKL is untouched, and ReadSKLFile now carries the engine's own Verify(!IsDirty()) guard. Next tests are listed cheapest-first in the roadmap: keep the NotationFile alive, copy the Object= names, count board allocations, and re-check whether MakeEntityRenderables runs twice per mech. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
d66c2b0812 |
BT410 5.3.55: the mech skeleton is built -- .SKL walked into a dpl_DCS tree
BTL4VideoRenderer answers MechClassID: walks the video-object chain the way the engine does, hands every L4VideoObject::Skeleton entry to ReadSKLFile, and recurses the .SKL into a dpl_DCS tree with geometry instanced onto it. Verified repeatedly on the live pod: '[skl] video\mad.skl -> 26 nodes, 19 objects', with no 'wrong video resource type' complaint and no load failures. Those counts are exactly what the file declares (25 joint= entries + root, 19 Object= entries). Corrects the earlier success criterion in this file, which said 22 instances by reading the reference capture's 'instance x22' against DZoneCount=22. Damage zones are not geometry -- 28 dzone= tags spread across 19 objects. Translations are written to matrix[3]/[7]/[11], MUNGA's own AffineMatrix layout. It walks cleanly but no frame has been seen WITH the mech yet, so the slot choice is recorded as unconfirmed. Rotation stays identity by design: every base-pose angle in MAD.SKL is 0 or ~1e-3, so translation alone assembles the model and isolates one convention at a time. AND A CORRECTION I have to flag loudly: I earlier concluded from single runs that non-identity translations crashed the pod, 'isolated' it, and 'confirmed' the alternative also crashed. That was all noise. The same binary re-run gives walk / crash / walk / crash -- the known intermittent plane-write defect is now firing on ~half of pod runs and lands at different points each time, which is precisely what made it look deterministic. Never accept a single pod run as evidence on this rig; require two agreeing runs. That defect is now the top of the list: a run must survive both the skeleton build and the launch to render anything, which at ~50% is a coin flip on a four-minute cycle. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
fac559bc31 |
BT410 5.3.54: skeleton-walk API inventory complete; staged plan with pixel-free checks
Confirmed every call the .SKL walk needs -- the NotationFile accessors and the MakeEntryList/GetFirstEntry/GetNextEntry idiom (copy DPLReadINIPage's objectpath loop), and the dPL side (NewDCS/SetDCSMatrix/AddDCSToDCS/ AddDCSToScene/SetDCSZone/FlushDCS, LoadObject/NewInstance/SetInstanceObject/ AddInstanceToDCS/FlushInstance). On the matrix: a DCS flush body is [remote][type_check][node][64 bytes] = a 4x4 of float32. Decoding a real BT capture shows a near-identity with a single 10.0 term in the last row, suggesting row-major with translation in row 3 -- but the rows print shifted by one word against a true identity, so the decoder offset is suspect and the convention is recorded as NOT yet proven. Flagged rather than guessed. Staged the work so each half is verifiable without looking at pixels: build the tree with identity matrices first and prove the structure by wire counts (26 DCS + 22 instance flushes, which MAD.SKL's JointCount=25/DZoneCount=22 and the dpl3-revive reference capture agree on), then add real transforms and fix the convention by watching which slot moves. Also noted the one API still to check first: the skeleton reaches us as a non-Object L4VideoObject::ResourceType, so branch on that enumerator rather than re-deriving the filename. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
a6612d5f75 |
BT410 5.3.53: full spec for the mech-skeleton brick, with an exact success criterion
Everything needed to write the .SKL -> dpl_DCS path is now known and written down: the file format (each page = one node with a local transform, an optional .bgf, damage-zone tags and its child joint pages), the dPL call set (NewDCS / SetDCSMatrix / AddDCSToDCS / LoadObject / NewInstance / AddInstanceToDCS / FlushDCS), and the algorithm matching RPL4VID.HPP's ReadSKLFile + RecurseSKLFile, including where DPLJointToDCSTranslator picks up the joint->DCS mapping afterwards. One unknown is flagged honestly: no surviving source calls dpl_SetDCSMatrix (the RP game file that did is missing, like BT's), so the matrix convention must be derived from the dpl3-revive protocol spec and DPLTYPES.H. The success criterion is exact and needs no pixels: the dpl3-revive reference capture of a real pod decodes as 26 DCS flushes and 22 instance flushes, and MAD.SKL declares JointCount=25 (+root = 26) with DZoneCount=22. A correct walk of that one file should reproduce those counts on the wire. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
15d509c8a8 |
BT410 5.3.52: frame rate measured properly -- the missing mech skeleton is the real gap
The pod updates every few seconds. Measured by sampling a head window every
2s (emulator/render-bridge/headrate.py): OURS changed in 5 of 12 samples,
SHIPPED in 1 of 12. So the reconstruction is not slower than the shipped
binary -- the emulated board is just expensive. Caveat recorded with it:
ours was being driven by the throttle hooks while the shipped exe ignores
them and sat parked, so treat those as same-order, not a win.
What did cost us 3x was mine: BT_MECH_LOG/BT_LAUNCH_LOG write per-frame lines
and DEBUG_STREAM=cout is redirected to COM3, so every one goes through an
emulated serial port. Turning them off took the wire from 480 to ~1480
bytes/sec. pod_render_quiet.conf is the conf to use for timing work.
And a correction to my own earlier framing: wire bytes/sec is NOT a frame
rate. Shipped pushes ~8900 B/s against our ~1480 and the difference is
CONTENT, not speed -- shipped submits the mech and we do not:
OURS: L4VIDEO.cpp wrong video resource type for object mad.skl
SHIPPED: (no such line)
The mech's model resource is a SKELETON. The engine's default
MakeEntityRenderables accepts only Object/Rubble and rejects anything else
with exactly that message, because skeletons are the GAME renderer's job. So
our arena renders but the MECH IS ABSENT, and with it the cockpit interior --
one unimplemented path explaining both the missing model and the lower wire
volume.
Next brick is now precisely scoped: answer MechClassID by reading the .SKL
notation pages into a dpl_DCS hierarchy and hanging the per-node .BGF
geometry off it.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
09192367c0 |
BT410 5.3.51: THE 3-D WORLD RENDERS -- reconstruction to emulated board to a picture
emulator/render-bridge/first-3d-frame.png: the arena from the pod -- sky, horizon, ground, the arena structures along the skyline -- at ~31fps on the VelociRender bridge. The whole chain now works in our build: MakeVideoRenderer -> BTL4VideoRenderer over DPLRenderer -> board boot -> Renderer::LoadMission -> DPLReadEnvironment for the art paths -> 40/40 arena objects loaded -> mission launch -> per-frame submission over the wire. The mistake worth remembering: the run was never stalled after InitializePlayerLink. I called that from a 150-second sample and it was just too short. CheckLoadMessageHandler reposts every second and will not advance until the min-priority event queue drains, and with the video renderer in the mission that takes minutes. A BT_LAUNCH_LOG trace showed the queue draining and then both RunMissionMessageHandler calls, the plasma display, and the first sensor tick -- matching the shipped binary line for line. The renderer deliberately holds a blank screen until RunningMission (L4VIDEO.CPP:5100), so 'black' was the app still loading, not a rendering bug. Zero unbuildable entities, zero geometry failures. The frame is static only because the mech is parked with no input, and the camera in the bridge title is the bridge's own viewer, not the game's eyepoint. Also banked the launch-gate trace (BT_LAUNCH_LOG) behind an env var. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
060078e4ae |
BT410 5.3.50: geometry loads -- the bring-up no-op was suppressing the base
DPLRenderer::LoadMissionImplementation (L4VIDEO.CPP:6007) is not empty: it
calls DPLReadEnvironment (which opens L4DPLCFG/btdpl.ini and sets the dPL
object/material/texmap paths from the objectpath= entries) and then
LoadNameBitmaps. The 'bring-up no-op' installed earlier therefore did not do
nothing -- it REPLACED that, so the loader never learned where the art lives
and every dpl_LoadObject returned NULL.
I justified that no-op by pointing at VideoRenderer's bare Tell and
GaugeRenderer's identical one; neither is our base. Check the ACTUAL base
before overriding in this engine and chain it unless there is a reason not
to. DPLReadEnvironment is private to DPLRenderer, so chaining is the only
way a game renderer can reach it at all.
before: 40x 'couldn't load object', 'NULL instance', run ends,
wire ~24 bytes/sec
after: 0 failures, run continues, wire ~12 KB/s (1.17MB climbing),
bridge live at 88fps
Proved it was ours and not the rig by running the SHIPPED exe under the same
conf: it loaded every object. That A/B is cheap and is the right first move
whenever the pod misbehaves.
Still black: the bridge camera sits at its default (0,10,0), so the wire
carries state and object loads rather than a populated scene. The mech and
arena entities still need MakeEntityRenderables bodies.
Rig improvements, both from the user: nosound is fine on the SLOW clock (it
is the FAST SOS clock that needs the AWE32), which saves ~4 min and ~300MB of
audio taps per run; and serial3=file + '> COM3' gives a live unbuffered log,
so a run that does NOT crash is finally readable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
||
|
|
ed7843b89f |
BT410 5.3.49: first renderable answer -- BTPlayer, and the pod stops complaining
BTL4VideoRenderer now overrides MakeEntityRenderables. Class 3035 resolves to BTPlayerClassID (the BT enum block starts at 3000 in VDATA.HPP), and a player carries no graphics -- exactly what the engine already does for its own PlayerClassID with an empty case. Everything else chains to DPLRenderer, so the override can only add answers. On the pod: the 'couldn't figure out how to MakeEntityRenderables' complaint is gone and the run no longer exits, it keeps running. It still renders nothing -- the wire fifodump grows at ~24 bytes/sec, keep-alive rather than a frame stream -- which is expected, since only the player has been answered and the mech and arena still have no renderables. Banked a practical problem worth fixing before the next session: a pod run that does NOT crash produces no readable log, because the conf redirects the game's stdout and DOS buffers it. Every readable log this session came from a run that crashed and had its buffer flushed by the fault handler. The RC.TXT marker (a separate command) survives that, and the com3/serial3=file trick from the emulator notes would give live unbuffered output -- worth wiring in, otherwise progress is invisible precisely when things are going well. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
32023a918f |
BT410 5.3.48: watcher ladder complete; the engine names the next brick
With every authored watcher resolving, the pod run reaches the mission, draws the FULL COCKPIT with the board booted and audio running, and exits GAME-RC=0 with no Fail. It builds no 3-D scene, and the engine says why: Entity 1:1 class3035 couldn't figure out how to MakeEntityRenderables L4VIDEO.cpp couldn't load object sky.bgf (then the whole arena) NULL instance VideoRenderer::MakeEntityRenderables (VIDREND.CPP:231) is the BOTTOM of a virtual chain -- its comment says so outright -- and BTL4VideoRenderer does not override it, so every entity falls through, no scene graph is built, and the geometry that would hang off it never loads. CORRECTS an earlier note in the roadmap: I recorded the first full-rig run as proving '.BGF loading was never a reconstruction problem, only a missing bridge', because the complaint count was zero. It was zero because the run Failed at a watcher long before reaching geometry. Now that it gets there, the loads are attempted and fail -- for want of renderables, not the bridge. Next brick is the real btl4vid body, shape pinned by the surviving sibling header RPL4VID.HPP: override MakeEntityRenderables, plus ReadSKLFile / RecurseSKLFile to walk the .SKL skeleton pages into a dpl_DCS hierarchy and load the .BGF geometry per node. Renderable content from the BT411 donor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |
||
|
|
4168f1ad4d |
BT410 5.3.47: the whole authored watcher set is published -- cockpit renders in the pod
Ladder rungs, each run-verified on the pod rig: ReportLeak (HeatSink), GeneratorOn (Generator -- member existed, publication missing), ConfigureActivePress (MechSubsystem -- likewise), AmmoState + FireCountdownStarted + the rest of the AmmoBin table. THE PINNED RANGE IS NOW THE AUTHENTIC SHAPE. Publishing ReportLeak on HeatSink and ConfigureActivePress on MechSubsystem shifts the chain down two, and MechWeapon's bridge to its binary-pinned PercentDone (0x12) collapses from THREE guessed pads to ONE -- which is exactly the arithmetic the shipped string pool predicts (MechSubsystem 1, HeatableSubsystem 3, HeatSink 6, PoweredSubsystem 5). Three pads of guesswork replaced by a real attribute and a real base-class row. HeatableSubsystem and Torso rebase onto MechSubsystem::AttributeIndex; MechControlsMapper does not (it derives from Subsystem directly). Types were chosen deliberately this time, per the AudioWatcher families the engine instantiates (Motion / Hinge / Scalar / StateIndicator): every *State name resolves to a StateIndicator, ReportLeak is a Scalar. RESULT: the pod run now gets past every authored watcher and DRAWS THE FULL COCKPIT -- sensor cluster, myomers, cooling, the weapon panels, kills/deaths -- with the board booted and audio running. It then exits without flushing its redirected stdout, so the exit reason is not yet known; next step is a run with the redirect removed so the console is readable. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> |