Files
TeslaRel410/RIOv4_2-ANALYSIS.md
T
CydandClaude Fable 5 e7869a6d04 RC1 acceptance PASS: E0000305 after check — edit 6 verified on both paths
Over-threshold check exchange repaints and re-renders a live E0 readout
(not the stale 04000000). The +2 drift in $3184 among 65 status frames
refines the counter model: $3184 = started timeout-retry sequences
(give-ups are the exhausted subset); timeout-recovered cycles exit via
the success teardown, NAK-recovered ones via $DA2F ($3185). All
observations from both bench sessions fit. Mash spot-check remains
before christening RIO 4.3.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-19 16:46:53 -05:00

549 lines
28 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# RIO v4.2 firmware — protocol wedge analysis
Reverse-engineering of `RIOv4_2.bin` (Toshiba TMP68HC11, AM27C512) to find
the board-side cause of the "reply path wedges under stress, button-press
revives it" fault. Disassembly by `disasm_6811.py`
`RIOv4_2.disasm.asm`. **Research only** — the fix below is proposed, not
yet burned or tested (no spare EPROM on hand). Validate on hardware with
the `RIO_TAP` mash test before trusting.
Addresses are CPU = file offset (EPROM at `$C000-$FFFF`; reset `$FFFE`
`$C000`). RAM lives at `$20xx-$31xx`.
## How the serial protocol is structured
- **SCI interrupt** (`$FFD6``$D630`): `JSR $D634; RTI`. `$D634` runs BOTH
workers every interrupt: `JSR $D6EA` (RX) then `JSR $D887` (TX). So the
transmitter is poked after every received byte, not only on TX-empty
interrupts.
- **RX ISR** `$D6EA`: reads SCSR/SCDR, stores the byte at `$3172`, then
`LDX $292F; JMP $00,X` — dispatches through a **state-handler pointer**
at `$292F`. Handlers classify bytes (`$D717`: `FE`=RESTART, `FF`=IDLE,
`FC`/`FD`=game ACK/NAK, `$82`=analog request, table lookup at `$3144`),
accumulate the body + checksum (`AND $7F`), and on a complete packet run
the ACK/NAK decision at `$D81F`.
- **TX ISR** `$D887`: if TDRE, send a pending ACK (`$316F``$FC`) or NAK
(`$3170``$FD`), else dispatch through the TX state pointer `$2D3B`
(`$D8C2` ring-drain → `$D90E` reply/retry machine). When idle it disarms
the TX interrupt (SCCR2 `#$2C`, TIE off) at `$D918`; the enqueue routine
`$D63B` re-arms it (SCCR2 `#$AC`, TIE on) at `$D664`.
## The wedge: an orphaned "reply-in-progress" latch (`$2521`)
`$2521` = "an analog reply is in progress." The analog-request handler
gates on it:
```
D74F CMPB #$82 ; analog request from the game
D753 LDAA #$01
D755 STAA $2520 ; arm reply generation
D758 TST $2521 ; already replying?
D75B BNE $D77A ; YES -> D77A: CLR $2520, drop this request
```
So while `$2521` is set, **every analog request is silently dropped**.
The latch is set when a reply is generated:
```
D847 JSR $C5EC ; build the analog reply
D84C STAA $2521 ; reply-in-progress = 1
```
and is cleared in only three places: power-on init (`$C0A3`), a host
reset/init command handler (`$C686`), and the reply **success** teardown
(`$DA00`). The success teardown is reached at `$D9C1` when the game ACKs
the reply, and clears the latch — but only conditionally:
```
DA21 LDAA $2522 ; did the $87 reply byte actually start sending?
DA24 CMPA #$01
DA26 BNE $DA2E ; if not, skip the clears <-- fragile
DA28 CLR $2521
DA2B CLR $2522
```
`$2521` is set the instant the reply is *generated* (`$D84C`), but `$2522`
is set only once the `$87` command byte *starts transmitting* (`$D8FD`).
**The leak** is the retry-exhausted give-up path, which is *separate* from
the success teardown. When the game fails to ACK a reply, `$D90E`/`$D9BE`
retries up to 4 times, then gives up:
```
D9D5 LDAB #$FE ; give up: send RESTART
D9D7 STAB $102F ; SCDR
D9DA INC $317A
D9DD JMP $DA2F ; teardown -- but DA2F never touches $2521
```
`$DA2F` resets the TX pointers and calls `$D5F2` (a debug-counter
formatter that does *not* clear the latch), then returns. **`$2521` is
left set forever.** From then on every `$82` analog request is dropped at
`$D758` → the board is mute to analog while its RX/event path stays fully
alive.
### Why a button press / new game revives it
The only mid-run code that clears `$2521` is the host command handler at
`$C669-$C689` (it clears `$2520`/`$2521`/`$2522` plus a raft of state).
That runs for a host-level reset/init command — exactly what the game
sends at game-start / on the player's opening button actions. Mid-mission
button-mashing sends no such command, so the leaked latch stays stuck
until the next game-start reset. This matches the field observation
precisely: the board goes mute under stress and only a new-game/button
resync brings analog back.
### Why mash stress triggers it
Button-event traffic floods the link while the board is mid-analog-reply;
the reply's ACKs collide/drop, the 4-retry budget exhausts, and the
give-up path (`$DA2F`) fires — leaking the latch. Light traffic rarely
exhausts the retries, so it's a stress-only fault. Two different USB
adapters showed the identical stall because the defect is in the board,
not the transport — consistent with this being firmware, not timing.
## Proposed fix (minimal, in-place; UNTESTED)
Clear `$2521` on *every* reply teardown, not just the `$2522`-gated
success path. Two edits, no code-size change, 8 KB of free ROM exists at
`$DFF0-$FFBF` for the stub:
1. **Give-up path** — redirect its teardown through a stub that clears the
latch first. At `$D9DD` change `JMP $DA2F` (`7E DA 2F`) →
`JMP $DFF0` (`7E DF F0`), and place at `$DFF0`:
```
DFF0 7F 25 21 CLR $2521
DFF3 7F 25 22 CLR $2522
DFF6 7E DA 2F JMP $DA2F
```
2. **Success path** — make the clear unconditional (belt-and-suspenders,
covers an abort before `$87` is sent). Replace `$DA21-$DA2D` (13 bytes)
in place:
```
DA21 7F 25 21 CLR $2521
DA24 7F 25 22 CLR $2522
DA27 01 01 01 01 01 01 01 (NOP x7)
DA2E 39 RTS (unchanged)
```
Rationale: `$2521` means "a reply is in progress"; any path that tears
down reply state must release it. There is no case where you reset the
reply machine yet want the latch to stay set, so unconditional clearing is
safe. This is the board-side analogue of the game-side "make collisions
harmless" patches (BTL4OPT v2-v4) — instead of widening a timing window it
removes the latch leak entirely.
### Patched binary — built & statically verified (2026-07-04)
`make_patch.py` applies both edits to `RIOv4_2.bin` (asserting the exact
original bytes at each site first) → **`RIOv4_2_patched.bin`**
(sha256 `3fc8170caf60e2580641724ff995176c93c4f2e706f31487beded8233142493f`,
23 bytes changed). Re-disassembling it (`RIOv4_2_patched.disasm.asm`) and
diffing against the original confirms the change is confined to exactly
three regions with no downstream desync:
- `$D9DD` `JMP $DA2F` → `JMP $DFF0`
- `$DFF0` new stub: `CLR $2521 ; CLR $2522 ; JMP $DA2F`
- `$DA21` `CLR $2521 ; CLR $2522 ; NOP×7` (RTS at `$DA2E` intact)
Flash `RIOv4_2_patched.bin` directly to the W27C512 (DIP-28). This is
static verification only; dynamic proof still needs the burned chip.
### 31250-baud variant (2026-07-17) — `RIOv4_2_patched_31250.bin`
`make_patch.py --baud31250` adds **one byte** to the wedge patch: the SCI
init at `$D62A` (`LDAA #$30 ; STAA $102B BAUD`) becomes `LDAA #$02`.
BAUD `$30` = prescale ÷13, divider ÷1 → 2 MHz E-clock / (16·13) = 9615
("9600"); `$02` = prescale ÷1, divider ÷4 → 2 MHz / (16·4) = **31250
exactly** (3.3× faster; analog exchange ~15 ms → ~4.6 ms; 96-lamp burst
~0.4 s → ~0.12 s; every collision window shrinks 3×). The BAUD write
confirms the 8 MHz crystal / 2 MHz E-clock — 19200/38400 are unreachable.
sha256 `9f866cf353d04906e1d3e3847b6375eaae251c987b656f68f093e61ba1bd545b`,
24 bytes changed (the 23 wedge bytes + `$D62B: 30→02`). Re-disassembly
diff vs the classic patched image shows exactly the one operand line.
**Caveats:** 31250 is non-standard — FTDI-class USB adapters make it
exactly, classic 16550 UARTs cannot; the **native games still speak
9600**, so a 31250 chip is bench/RIOJoy-only until Firestorm/Red Planet
get BTL4OPT-style baud patches; and the 2 MHz HC11's RX ISR has ~640
cycles/byte at this rate (fine) — do NOT be tempted to `$01`/62500 or
`$00`/125 k without cycle-counting the `$D630` ISR worst path first.
PC side: `SerialPortTransport` takes a baud parameter and the
monitor/mash tools accept `--baud 31250`. Validate the wedge patch at
9600 FIRST (one variable at a time), then A/B this chip with
`--mash COM1 300 --label patched-31250 --baud 31250`.
### Validation plan (when a chip is available)
Burn the two edits to a W27C512, socket it (preserve the original AMD
chip), then run the `RIO_TAP` two-handed 8-button mash test. Expect: no
permanent analog mute; any collision self-recovers without a game-start
reset. Compare dropout counts to the 2026-07-03/04 baseline taps.
**Instrumented harness** (2026-07-17): the mash protocol is mechanized in
`tools/RioSerialMonitor` —
dotnet run --project tools/RioSerialMonitor -- --mash COM1 300 --label patched
It disables the app's own >5 s reset-recovery (a wedge must stay observable),
echoes lamps on every press (lamp/reply collisions are the trigger), and logs
to `riomash-<label>-<stamp>.log`: analog-gap histogram + longest gaps, WEDGE
events with beep/banner and a self-recovered vs button-revived classification
(button within 300 ms of resume = the unpatched revival signature), and the
board's own RestartCount/AbandonCount/FullBufferCount before/after with the
delta (7-bit wrap-aware). Exit 0 = no wedge, 1 = wedge seen. Run once with
`--label baseline` on the original chip (control: prove the test still bites),
then `--label patched`; the fixed-layout summary blocks diff directly.
`--mash --selftest` runs a scripted in-memory board that fakes a
button-revived wedge — use it to sanity-check the alarm before a session.
## Firmware memory map (as decoded so far)
| addr | meaning |
|---|---|
| `$292F` | RX state-handler pointer (`JMP $00,X` dispatch) |
| `$2D3B` | TX state-handler pointer |
| `$2D34/$36/$38` | TX ring read/write/aux pointers (ring `$2932-$2D31`) |
| `$2520` | reply gate (analog request pending) |
| `$2521` | **reply-in-progress latch — the wedge** |
| `$2522` | `$87` analog-reply-byte-sent flag |
| `$316C/$6D` | game ACK / NAK received |
| `$316E` | unknown-command seen |
| `$316F/$70` | ACK / NAK pending to send (→ TX ISR) |
| `$3172` | last received byte |
| `$3173/$74/$75` | ACK/NAK/wait retry counters (limit 4) |
| `$317A/$7B` | RESTART / IDLE keep-alive counters |
| `$3184/$85` | give-up / error diagnostic counters |
| `$3186` | RX overrun flag (set at `$D701`, **never read** — not the cause) |
| `$102D/$2E/$2F` | SCCR2 / SCSR / SCDR (HC11 SCI) |
## Bench validation results (2026-07-18) — PATCH CONFIRMED
Rig: real RIO board, FTDI FT232R on COM1, `RioSerialMonitor --mash`
(app auto-recovery disabled, lamp echo on). Logs in `testlogs/`.
| | baseline (original AMD chip) | patched (W27C512) |
|---|---|---|
| startup version/check exchange | **wedged at 0.62s** (reply/reply collision) | survived |
| analog replies | 6 total, then dead for 300s | 2419 @ 152s, steady ~16/s |
| mash | 951 presses (peak 460/min) | 1315 presses (peak 635/min) |
| worst analog gap | infinite — no recovery in 300s despite 951 presses | 0.58s, self-recovered |
| wedge events | 1, unresolved (power cycle required) | 0 |
Notes: the baseline wedge didn't even need the mash — the tool's own
version+check request colliding with the analog poll killed the reply
path instantly, and (contrary to the earlier RIO-NOTES observation)
button presses did NOT revive it. The patched chip's single 0.58s gap
is the intended failure mode: collision → teardown → latch cleared by
the $DFF0 stub → self-recovery. Patched run cut at ~155s by operator
(finger fatigue); given the baseline's 0.62s time-to-wedge, 2.5 min of
heavier mash is decisive. Remaining (cabinet): overnight idle soak +
a native-game session. Both check snapshots also reported a board/lamp
fault item — persistent across chips, likely a real tired lamp; inspect
separately.
## 31250-baud bench run (2026-07-18) — patch holds; retry window is byte-scaled
`testlogs/riomash-patched-31250-20260718-003114.log` (127s mash, FTDI
COM1 @ 31250, latency timer already 1ms): **zero wedges**, but
framing=5655, analog replies at 215% of poll slots (duplicates), and
AbandonCount=65 — a reply-retry storm. Reading: the board's reply
ACK-wait window scales with BYTE TIME, not wall clock. At 9600 the
window (~5+ms) hides USB turnaround entirely (zero framing in both
9600 runs); at 31250 it shrinks to ~1.6ms, which USB cannot reliably
beat, so most replies retransmit and 65 exhausted all 4 retries.
Two conclusions:
1. **Brutal confirmation of the wedge patch**: 65 trips through the
give-up path in 127s, every one self-recovered ($DFF0 stub). The
unpatched firmware would have latched dead on the first.
2. **31250 needs one more firmware tweak to be clean**: widen the
reply-retry wait (the pacing/limit around the $D90E/$D9BE retry
machine) so the window stays >=10ms of wall clock at 31250. Until
then: cabinets run the 9600 patched chip (validated clean); 31250
stays a bench branch. Note the duplicate replies also mean the raw
reply rate overstates throughput — unique samples are still capped
by the host's 55ms poll; harvesting the speed needs a faster poll
AFTER the retry window is fixed.
### 31250 v2 (2026-07-18) — `RIOv4_2_patched_31250v2.bin` — widened ACK-wait
Root cause of the v1 retry storm CONFIRMED in the disassembly: the reply
ACK-wait loop at `$D9E0` self-clocks in **byte times** — each tick sends an
IDLE (`$FF`) keep-alive and re-enters on that byte's TX-complete interrupt,
so the `CMPA #$04` limit means ~5 byte times of grace: ~5.2ms at 9600
(USB ACK wins; zero framing in both 9600 runs) vs ~1.6ms at 31250 (USB
loses; framing 5655 / Abandon 65 on the v1 bench run). Those keep-alive/
RESTART bytes landing mid-packet are exactly the host-side framing resyncs.
`make_patch.py --baud31250 --widen-ackwait` → one more byte,
`$D9E7: $04 → $28` (40 ticks ≈ 12.8ms at 31250, clears FTDI worst case
with margin). The NAK-retry limit ($3175) is event-counted, untouched.
sha256 `420d4cfc6b513651687982a70db2daeecda7bd00a5324321e272f35b95dca753`,
25 bytes vs original; re-disassembly diff vs the v1 31250 image is exactly
the one CMPA operand line. Expected on the bench: framing 5655 → ~0,
Abandon 65 → ~0, wedges 0 → 0.
### 31250 v2 bench results (2026-07-18) — CLEAN. Byte-clock theory confirmed.
`testlogs/riomash-patched-31250v2-20260718-110853.log` (192s mash):
| metric | v1 @31250 | v2 @31250 (ACK-wait $28) |
|-------------------|----------------|--------------------------|
| framing resyncs | 5655 (~45/s) | **54 (~0.3/s, -99.4%)** |
| Abandon delta | 65 | **+0** |
| Restart delta | 2 | **+0** |
| analog vs polls | 215% (dupes) | **88.6% (no dupes)** |
| wedge events | 0 | **0** |
| worst mid-run gap | 0.09s (steady) | 0.14s (tail <100ms) |
The 88.6% delivery ratio equals the healthy 9600 patched profile (87.5%)
— the deficit is the $D758 in-flight drop gate, identical at both rates,
i.e. structural, not a v2 artifact. Counters perfectly flat: the board
never entered retry escalation in 192s. FastRIO is now *clean*, not
just survivable. To actually harvest the bandwidth, the next dial is
host-side: drop the 55ms analog poll interval (the board now answers in
~4.6ms of wire time). Cabinet adoption still gated on game-side baud
patches; 9600 patched chip remains the production standard meanwhile.
### 125000 bench results (2026-07-18) — CPU wall found; RX overrun confirmed
`testlogs/riomash-patched-125000-20260718-125310.log` (103s): zero wedges
(all gaps self-recovered <=0.52s even here), but the board's RX drowns:
NAK=1082 (~2/3 of the ~1650 lamp commands arrived checksum-corrupt —
overrun-dropped bytes; lamps visibly dead), analog delivery 60.6% (2-byte
requests survive better than 4-byte packets), both status exchanges
failed. Direction asymmetry nails the cause: board->PC was near-perfect
(framing 5), so the MAX232 leg is fine at 125k — the 2MHz HC11 RX ISR
(~160 E-cycles/byte, 1-byte buffer) is the wall, exactly as the cycle
estimate predicted. Verdict: 125000 NOT viable. Sweet-spot candidate:
`RIOv4_2_patched_62500.bin` (320 cycles/byte, MAX232 in-spec, ACK-wait
80 ticks = 12.8ms), sha256 below, awaiting bench.
### 62500 bench results (2026-07-18) — the ladder is complete
`testlogs/riomash-patched-62500-20260718-130333.log` (132s): analog
pristine (88.1% delivery = the structural drop-gate ratio; worst mid-run
gap 0.14s; counters flat; zero wedges; status exchanges fine) but
**NAK=69 of ~1060 lamp commands (~6.5%)** — 4-byte bursts still outrun
the RX ISR occasionally, and fire-and-forget lamp writes turn each loss
into a visible stuck-bright / missed-bright lamp. The ISR's worst-case
path therefore sits right at ~320 E-cycles.
**The RIO speed ladder (2 MHz E-clock, all with wedge fix):**
| rate | bytes OK inbound | verdict |
|---|---|---|
| 9600 | all | production (native-game compatible) |
| 31250 v2 | all (NAK 0.3%) | CLEAN — FastRIO recommended |
| 62500 | 2-byte yes, 4-byte ~6.5% loss | usable ONLY with host NAK-retransmit |
| 125000 | ~40-65% loss | not viable — CPU wall |
Host-side NAK-driven retransmit (the protocol's intended use of ACK/NAK;
RioSerialLink currently fire-and-forgets outbound) would make 62500
lamp-solid (1 retry -> ~0.4% residual, 2 -> ~0.03%) and is good
robustness at any rate. Until then, 31250 v2 is the honest FastRIO.
### 62500 final verdict (2026-07-19) — SHELVED. Chord-driven RX overrun.
Three rounds of host-side hardening at 62500 (evidence in `testlogs/`,
implementations in RioJoy.Core RioSerialLink):
1. **NAK-race resend** — resends tracked NAKs 1:1 (321/321) yet lamps
still glitched: under bursts the NAK arrives after a newer command is
already "latest", so the wrong packet was resent; total shreds never
NAK at all.
2. **Stop-and-wait** (one command in flight; ACK/NAK/50ms-timeout;
retransmit the SAME packet, limit 2) — still glitched: reply
resolution was type-blind, so an analog reply in USB transit from the
previous poll falsely confirmed the next lamp command.
3. **Typed resolution** (a reply only resolves its matching request;
ACK/NAK safe type-blind because the board's TX ISR prioritizes them
ahead of reply data; 10ms settle after budget-exhausted drops) —
improved, but **chorded presses still glitch even at slow press
rates**.
The chord observation closes the case: a chord makes the board scan,
queue and transmit N button events at once — worst-case ISR latency —
exactly when our 4-byte lamp replies arrive at 160us/byte into a 2-byte
RX buffer. Corruption at 62500 is chord-shaped, not rate-shaped; no
host-side protocol can prevent the board losing bytes it never latched,
and retries can land in the same busy window. Fixing it for real means
firmware ISR restructuring (or sequence numbers) — out of scope.
**Final ladder verdict: 31250 v2 is the FastRIO standard** (3.3x stock,
clean without any host-side compensation — NAK 5 in 192s of mash).
The stop-and-wait + typed-resolution machinery stays enabled at every
rate as delivery insurance (it hardens 9600 production cabinets against
ordinary electrical noise, and its retries are near-zero on clean links).
62500 images remain in the repo for future firmware work.
## FINAL CERTIFICATION (2026-07-19) — 31250 v2 ships
Operator-verified at 31250 v2 with stop-and-wait active: **zero lamp
misses/hangs** through the slow-chord test (the exact scenario that
condemned 62500) and a 392-presses/min mash.
`testlogs/riomash-31250v2-final`: analog 89.0% (structural), **NAK = 0**,
25 resends — all timeout-triggered, i.e. the retry machinery silently
healing occasional INBOUND response loss (29 host-side framing events),
which is exactly the insurance it exists to provide. Counters flat,
zero wedges, gaps <100ms except run-start/stop artifacts.
`testlogs/riomash-31250v2-poll20` (--poll 20): delivered ~31 Hz analog
(2036 replies/65s) vs the legacy 18 Hz — 1.7×. The 62.6%-of-slots figure
is a host artifact, not board saturation: the poll loop delays AFTER each
awaited exchange (~6-9 ms), so the real cadence is interval+exchange.
Absolute-schedule pacing would approach 40+ Hz; ~94% of actually-issued
polls were answered. Future host tweak, not firmware.
### Production configuration
| deployment | chip | config.json |
|---|---|---|
| native-game cabinets | `RIOv4_2_patched.bin` (9600) | defaults (RioBaudRate 9600, AnalogPollMs 55) |
| RIOJoy-only cockpits (FastRIO) | `RIOv4_2_patched_31250v2.bin` | RioBaudRate 31250, AnalogPollMs 20-25, FTDI-class adapter |
AppConfig.RioBaudRate + AnalogPollMs are plumbed through RioCoordinator;
stop-and-wait delivery is always on (near-zero cost on clean links).
Firmware images, hashes, disassemblies, the mechanized mash harness, and
all eleven bench logs live in this directory. The original AMD chip is
retired, preserved unmodified.
## Error-handling inventory (v4.2, unpatched baseline)
Complete map of what the board does — and pointedly does not do — when
things go wrong. Complements the wedge analysis above; display-side
detail in [`docs/hardware/display-board-1408.md`](../docs/hardware/display-board-1408.md).
### Host-facing (protocol)
- **Inbound packet checksum** (7-bit sum, verified at the `$D81F`
decision): good packet → queue **ACK `$FC`** (`$316F`); bad → **NAK
`$FD`** (`$3170`). The TX ISR sends pending ACK/NAK ahead of reply
data (`$D8A6`). Control bytes `$FE` RESTART / `$FF` IDLE.
- **Outbound reply delivery**: replies must be ACKed by the game; the
retry machine (`$D90E`/`$D9BE`) retransmits — **5 times** (bench
2026-07-19; the earlier code reading of "4" was wrong) — then gives
up: sends `$FE` RESTART, `INC $317A`, `INC $3184` (give-up counter),
teardown at `$DA2F` (`INC $3185` — teardowns: ANY reply cycle that
needed at least one retransmission ends here, including NAK-recovered
ones, so `$3185` is the counter behind a lone cockpit `E0000001`).
The give-up teardown is where the `$2521` wedge leak lives (see
above; fixed in the patched images). Bench-observed await behavior:
the ACK wait arms only after the COMPLETE reply frame is sent
(mid-frame ACK/NAK/RESTART is ignored); a game NAK of the completed
frame triggers exactly one immediate counted retransmit; once the
first retry timeout lapses the cycle runs blind to give-up — a late
ACK cannot rescue it.
- **RX overrun: unhandled.** The RX ISR reads SCSR/SCDR but never acts
on the overrun flag — dropped bytes just corrupt the frame and
surface as checksum NAKs. This is why the 125000-baud bench run
produced NAK=1082 rather than any explicit overrun signal.
### Internal
- **TX-ring overflow** (`$D63B` enqueue): if the ring would overtake
the drain pointer, the byte is silently dropped, `INC $3187`, and the
E0 diagnostic display refreshes. The host is not told.
- **TX dispatcher sanity trap** (`$D8C2`): each outgoing command byte
is looked up in an 8-entry table at `$3158`; no match ⇒ internal
corruption ⇒ `JMP $DAB2` — the **`dEAd` crash screen**, an infinite
loop rendering `dEAd` + RAM `$0000/$0001` on the 8-digit display.
Permanent until reset.
- **Unused interrupt vectors** → `NOP; BRA self` stubs at
`$DB07-$DB3D`: silent hang, no display, no recovery.
- **Error counters** `$3173-$317B` (ACK/NAK/retry/give-up bookkeeping)
and the three displayed ones `$3187/$3184/$3185` — rendered as
`E0 <ring-ovf> <give-ups> <teardowns>` by `$D5F2` at every increment.
### Hardware (RIO_1407 sheet 2)
- **MAX690 supervisor**: power-on/brownout reset via the PFI
comparator. The front-panel RESET toggle (SW1) and the optocoupled
remote-reset input both work by pulling PFI low — a "commanded
power-fail". **FAULT LED = PFO**, so it lights during any such event.
- **No working watchdog.** WDI has no off-sheet tag (never reaches the
CPU) and the firmware contains no strobe code — so software hangs
(`dEAd`, vector stubs, the reply wedge) persist until a human or the
game PC resets the board. This is corroborating evidence for the
bench-observed behavior: a wedged board stays wedged indefinitely.
- **RUN LED = buffered `AS*`** (address strobe): it indicates "clock
alive", not "software alive" — it stays lit even inside the `dEAd`
loop. TX/RX LEDs are buffered `SER_OUT`/`SER_IN`.
### E0-display threshold patch (edit 5, `--e0thresh[=N]`)
Stock firmware has **no threshold**: the first increment of any of the
three displayed counters repaints the cockpit display from `F0000000` to
the `E0` readout and it never reverts — so a single benign reply
give-up (e.g. one lost ACK healed by host-side stop-and-wait) leaves a
permanent `E0000001`. `make_patch.py --e0thresh=N` (default 5) gates the
renderer: `$D5F4`'s `LDX #$2038` becomes `JMP $E000`, and a 30-byte cave
at `$E000` (erased region, clear of the `$DFF0` wedge stub) compares
`$3187`/`$3184`/`$3185` against N — all below ⇒ exit through the
routine's own epilogue (`$D61D`, registers restored, display untouched);
any at/above ⇒ resume the render at `$D5F7`. Counters still accumulate
regardless, so the diagnostic history is intact — the display just stays
quiet until real trouble (N events since power-on).
Built variants (wedge fix + threshold 5):
| image | config | sha256 (first 12) |
|---|---|---|
| `RIOv4_2_patched_e0t5.bin` | 9600, native-game compatible | `9c21ac7199fb` |
| `RIOv4_2_patched_31250v2_e0t5.bin` | 31250 + widened ACK-wait (FastRIO) | `b43032b016d7` |
**Bench-verified 2026-07-19 (9600 image, burned): PASS.** Driven by
`RioSerialMonitor --e0test` (DTR reset → clean handshake → four
NAK-then-ACK sub-threshold teardown cycles → one unACKed give-up
cycle): the display held `F0000000` through the handshake and all four
sub-threshold teardowns — stock paints `E0…` at the first — and flipped
exactly at the 5th teardown to **`E0000105`** (`$3187`=00 overflows,
`$3184`=01 give-up, `$3185`=05 teardowns), confirmed by eye on the
cockpit display. Mash regression (120s, lamps on,
`testlogs/riomash-patched-e0t5-9600-20260719-154259.log`): **PASS**
0 wedges, 236/236 presses, 88.0% analog fill, NAK 0, board counters
flat. The 31250v2_e0t5 image is the same patch on the certified
FastRIO base and is expected to behave identically.
**Post-run display `04000000` — the CheckRequest self-test (corrected
finding).** The first read of this (keypad A-9-E chord entering test
mode) was wrong — disproved on the bench: the display is static (axis
movement does not change it) and version-only exchanges leave it alone.
The real mechanism is the **CheckRequest (`0x80`) handler at `$C5A6`**:
a full built-in self-test that sets the test-display flag `$2421`
(`$C5B1`), sends `TestModeChange 0x8C` to the host (`$C70E`), runs a
lamp pattern (`$D34A`), scans the pods, sweeps all five encoder
channels (`$CC53`, ending on channel 4 — whose `cc00hhll` frame is the
last thing rendered), enqueues the fault reports, clears the flag
(`$C5E1`), notifies `0x8C` again — and **returns without repainting the
display**. Every status check therefore leaves a stale
`04 00 <ch4 count>` snapshot (idle: `04000000`) on the digits. The
board is fully healthy afterwards; no reset is needed. Proper fix:
**edit 6** (`make_patch.py --checkrepaint`) — a 10-byte cave at `$E020`
hooked from the handler's final notify (`$C5E4`): `JSR $C70E` (original
notify), `JSR $CB7C` (repaint `F0000000`), `JSR $D5F2` (re-render the
E0 readout, which the edit-5 gate suppresses below threshold), `RTS`.
Built into **`RIOv4_3rc1.bin`** (with edits 1-2 and 5; sha256
`dc59bd51cae3…`); prior generations moved to `archive/`.
Why `04000000` was never seen in period (user-verified by re-slotting
the ORIGINAL stock chip, 2026-07-19): stock leaves the same stale frame
after every check, but the ungated E0 readout papers over it at the
first counter event — `04000000` flashes momentarily, then `E000000x`
as errors accrue and the reply-latch wedge trips under load. The wedge
fix plus the E0 threshold gate are what exposed the stale frame.
**RC1 acceptance (burned 2026-07-19): PASS on both edit-6 paths.**
Below threshold: a version+check exchange leaves the display at
`F0000000`. Over threshold: the full `--e0test` sequence reproduced the
certified behavior (held `F0000000` through four sub-threshold events,
flipped at the 5th teardown), and the closing CheckRequest — 65 status
frames — left the display at **`E0000305`**: repaint + live E0
re-render, not the stale `04000000`. The `$3184`=03 (vs predicted 01)
pins down the counter semantics one notch further: two of the 65 check
frames hit a timeout-retry that our late ACK then rescued — so
**`$3184` counts started timeout-retry sequences** (give-ups are the
subset that exhaust all 5 retransmits), and a timeout-recovered cycle
exits via the success teardown (no `$3185`), while NAK-recovered cycles
tear down through `$DA2F` (`$3185`). Every observation from both bench
sessions fits this model. Remaining for the RIO 4.3 christening: mash
spot-check.