BT410 5.3.74: the faulting instruction decoded -- a handler-chain walk with a bad node

The SegPhys-corrected dump read real code at last (the earlier zeros were the
probe reading EIP as a bare linear address; CS=00FF has a base):

    66D4  push 0 / push edi / push ebx / push esi
    66D9  call dword near [ebx+4]     <- faults
    66DC  cmp eax,0 / jz done
    66E5  mov ebx,[ebx] / jmp loop

A linked-list walk with a per-node callback -- node = {next@+0, handler@+4},
four args, "handled" on nonzero.  A DPMI exception/interrupt handler chain.

The arithmetic closes exactly: the read is DS_base + ebx + 4, and
7000FA64 - F000CA64 = 80003000, so DS base is 0x80003000 and the famous cr2
is that wrap.  My earlier '[EBX+0x3004]' reading was fabricated from the cr2
alone -- there is no 0x3004 displacement, which is why the constant appeared
in no binary.

So: a chain node's NEXT pointer holds F000CA60, the value parked in unhooked
IVT slots -- the walk expected a terminator and got an interrupt vector, then
read its +4 as a function pointer.  Identical registers on every catch, so
one deterministic path, entered often enough that our 60s load window catches
it 25-40% of the time while shipped's 15s window mostly does not.

Open: which routine owns the chain, and where the bogus next came from.  The
probe now walks the EBP frame chain to name the caller.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Cyd
2026-07-29 18:34:22 -05:00
co-authored by Claude Fable 5
parent 2ce3528f1c
commit e97862507c
5 changed files with 651 additions and 644 deletions
@@ -1634,3 +1634,43 @@ serial stream's presence but not with any particular byte; the bad read =
host routine. Next catch carries corrected stack/code dumps; if those name
a return chain, the routine falls out. Rate note: 4 faults in the last 10
vRIO-live runs -- hovering near the 25-40%% band, still load-window-only.
--------------------------------------------------------------------------------
THE FAULTING INSTRUCTION, DECODED: A HANDLER-CHAIN WALK WITH A BAD NODE
--------------------------------------------------------------------------------
The SegPhys-corrected dump finally read real code (the earlier zeros were the
probe reading EIP as a bare linear address -- CS=00FF has a base). At the
fault:
66CF ... <- loop top
66D4 push 0
66D6 push edi ; 00017574
66D7 push ebx ; F000CA60 <- the node
66D8 push esi ; 00017640
66D9 call dword near [ebx+4] <- FAULTS
66DC cmp eax,0
66DF jz 6E81 ; a handler claimed it -> done
66E5 mov ebx,[ebx] ; next node
66E7 jmp 66CF ; loop
That is a LINKED-LIST WALK with a per-node CALLBACK: node = {next @ +0,
handler @ +4}, four args pushed, "handled" if the handler returns nonzero.
Textbook DPMI exception/interrupt handler chain.
The arithmetic closes exactly: the read is DS_base + ebx + 4, and
0x7000FA64 - 0xF000CA64 = 0x80003000, so DS has base 0x80003000 and the
famous cr2 is just that wrap. (My earlier "[EBX+0x3004]" reading was
fabricated from the cr2 alone -- there is no 0x3004 displacement; the
instruction is a plain +4. That is why the constant appeared in no binary.)
So the defect is: a chain node's NEXT pointer holds 0xF000CA60, the value
DOSBox-X (and a real BIOS, comparably) parks in unhooked IVT slots -- the
walk expected a terminator and got an interrupt vector. The host then reads
that value's +4 as a function pointer and dies. Every catch has identical
registers (EAX=1 EBX=F000CA60 ECX=0 EDX=FFFFFFFF, same ESP/EBP), so this is
one deterministic path, entered often enough that a 60-second load window
catches it ~25-40%% of the time and a 15-second one (shipped) mostly does not.
Still open: WHICH routine owns the chain and where the bogus next came from.
The probe now walks the EBP frame chain to name the caller; faulthunt.sh is
looping for the next catch.