Golden-model conformance suite (C1-C4 CONFORMANT) + GPU-compute retarget plan

igc_conformance.py pins the reference behaviour any port (RTL/GPU/C) must
reproduce: the DUMP packet parse, a synthetic raster, the bars pipeline
(texu==x+2 from the real cap7 tile programs), and the fxtest effect
primitives (coverage + z pinned) -- fixtures committed under conformance/.
GPU-RETARGET.md maps the DOSBox-renderer path: EMC tile as compute shader
(pixel=thread), 4 milestones, conformance-gated.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Cyd
2026-07-19 10:21:47 -05:00
co-authored by Claude Opus 4.8
parent cc7c95d738
commit 5d29442ecb
5 changed files with 206 additions and 0 deletions
+70
View File
@@ -0,0 +1,70 @@
# GPU-compute retarget — the DOSBox-renderer path (2026-07-19)
Companion to FPGA-RECONSTRUCTION.md. This is the plan for shipping the decoded
IGC/EMC model as a real-time renderer behind the existing DOSBox seam — the
practical answer to "use this as the renderer for the pod."
## Why GPU compute
The EMC array is 8,192 bit-serial PEs per 64×128 tile. On a GPU that is one
thread per pixel executing the same decoded op stream — a literal 1:1 mapping,
no hardware, ships as software next to the renderer we already deploy.
The original array ran ~40 MHz bit-serial; any integrated GPU exceeds the
whole frame's work by orders of magnitude.
## Architecture
```
DOSBox (game, unmodified)
│ wire protocol (known) ── existing seam: render-bridge/live_bridge
emu860 geometry stage ── the i860 firmware, C/JIT port for realtime
│ per-draw DMA stream + payloads (the coefficient programs)
IGC translator (host) ── igc_exec.parse -> flat op records
│ SSBO: op stream + payload floats + tile list
compute shader "EMC tile" ── pixel = thread; fields in registers/LDS:
│ texz/texu/texv/zbuf as uint20, r/g/b as uint8, enable as bool;
│ ops = a switch over the named table (TREE*, MEM*, CPY, sweeps, SCA*)
readout pass ── ramp(texu) / rgb24 -> RGBA8 frame texture
present (the existing renderer window / vr_readpixels reply)
```
## What ports 1:1 from the golden model
- The op table (data) and parser — host-side, unchanged.
- Tile semantics — each `Tile.run` branch becomes a shader case; the pixel
memory bit-fields become packed uints (no bit-serial emulation needed:
field-level semantics are already proven equivalent by the conformance
suite).
- Readout — ramp LUT as a small texture.
## Conformance gate
`igc_conformance.py` (C1-C4, committed fixtures) is the acceptance bar: the
shader path must reproduce C2-C4 outputs exactly (C1 is host-side parse).
Add a C5 full-frame image hash (bars + fxtest frames) once texel sampling
lands.
## Realtime budget (per frame)
- 52 tiles × ~100-3000 op-steps × 8,192 px ≈ 1-10 G pixel-ops worst case —
a compute-shader trifle at 60 fps on integrated graphics.
- Geometry stage: the real wall. emu860 in Python ≈ 600K steps/s; a frame is
~10-50M steps → need the C/JIT step-core (planned; the i860 ISA semantics
are fully pinned by emu860 + MAME cross-validation). Interim: pre-captured
streams replay at full speed for validation.
## Milestones
1. **M1 shader tile**: WGPU/GL compute kernel executing the named ops;
passes C2-C4. (small)
2. **M2 frame loop**: DMA-stream walker + tile dispatch + readout; renders
the bars + fxtest fixtures. (small-medium)
3. **M3 texel path**: TXDN stream -> texture sampling (needs the remaining
spec work — the io-port fetch protocol). (medium, spec-gated)
4. **M4 live seam**: behind live_bridge in place of / beside the GL renderer;
real-time gate = the C step-core for emu860. (the big one, independent)
## Open spec items feeding this (tracked in IGC-ENCODING-DERIVATION.md)
- op-0x2c edge exact semantics (winding rule) — affects M1 fidelity.
- TXDN io-port texel stream — gates M3.
- The texz coordinate-seed op + the real ramp tables — purity for M2 images.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+136
View File
@@ -0,0 +1,136 @@
"""igc_conformance.py -- the golden-model conformance suite.
Any port of the IGC/EMC model (RTL, GPU compute, C) must reproduce these
outputs bit-for-bit (where pinned) from the committed fixtures in
conformance/. The fixtures are REAL captured data from the firmware running
on emu860 (see IGC-ENCODING-DERIVATION.md for provenance).
Cases:
C1 DUMP reference packet parses to the exact named instruction sequence.
C2 Synthetic triangle rasterizes to the exact lit-pixel count.
C3 Bars pipeline: the cap7 bench per-tile programs (sends_cap7.pkl), with
the documented texz seed, compute texu == global x (+2) per pixel.
C4 fxtest effect primitives (fx_program.pkl) rasterize: 9 primitives drawn,
pinned z-buffer coverage + centroid.
Run: python igc_conformance.py (exit 0 = conformant)
"""
import os, sys, pickle, struct
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import igc_exec
HERE = os.path.dirname(os.path.abspath(__file__))
FIX = os.path.join(HERE, 'conformance')
def f32(w):
return struct.unpack('<f', struct.pack('<I', w & 0xffffffff))[0]
def c1_dump_reference():
ref = [0x100,
0x3ae94200, 0xc1900000, 0x41200000, 0xc4840000,
0x3ae94200, 0xc1400000, 0x41200000, 0xc34c0000,
0x3ae94200, 0x41f00000, 0xc1a00000, 0x44a50000,
0x44ea2120, 0x80000000, 0x80000000, 0x3f7fffff,
0x438a4320,
0x3daa4352, 0x3e849bae, 0xbdd70a3b, 0xc1976c7f,
0x3daa435a, 0xbea24ddd, 0x3d4ccccc, 0x421a3788,
0x41aa4362, 0x80000000, 0x80000000, 0x3f7fffff]
ins, unk = igc_exec.parse(ref)
names = [x[0] for x in ins]
assert names == ['SETENABS', 'TREEltZERO_L3', 'TREEltZERO_L3', 'TREEltZERO_L3',
'MEMltTREE_L3', 'TREEintoMEM_L0', 'TREEintoMEM_L3',
'TREEintoMEM_L3', 'TREEintoMEM_L3'], names
assert not unk
return 'C1 dump-reference parse: PASS (%d instrs)' % len(ins)
def c2_synthetic_triangle():
tri = [('SETENABS',),
('TREEgeZERO_L3', 0, 2, 1.0, 0.0, -8.0),
('TREEgeZERO_L3', 0, 2, -1.0, 0.0, 40.0),
('TREEgeZERO_L3', 0, 2, 0.0, 1.0, -20.0)]
t = igc_exec.Tile()
t.run(tri)
lit = sum(t.enab)
assert lit == 3564, lit
return 'C2 synthetic triangle: PASS (%d enabled)' % lit
def c3_bars_pipeline():
d = pickle.load(open(os.path.join(FIX, 'sends_cap7.pkl'), 'rb'))
pls = {(a, sz): list(w) for h, (a, sz, w) in d['payloads'].items()}
order = [(0x8015000, 4), (0x8015020, 69), (0x8015260, 33), (0x8015380, 41)]
t = igc_exec.Tile(ox=256, oy=0)
for y in range(igc_exec.TILE_H):
for x in range(igc_exec.TILE_W):
t._wr(t.mem[x + y * igc_exec.TILE_W], 32, 20, (x + 256) & 0xfffff)
for key in order:
ins, unk = igc_exec.parse(pls[key])
t.run(ins)
# texu must equal global x + 2 (the sweep's constant offset) at sampled cols
for x in (0, 16, 32, 48):
u = t._rd(t.mem[x], 57, 20)
assert u == 256 + x + 2, (x, u)
return 'C3 bars pipeline: PASS (texu == x+2 across the tile)'
def c4_fxtest_primitives():
d = pickle.load(open(os.path.join(FIX, 'fx_program.pkl'), 'rb'))
prog = d['prog']
def rd(a): return prog.get(a, 0)
setups = [a for a in sorted(prog)
if rd(a) == 0x100 and ((rd(a + 4) >> 8) & 0xff) == 0x2c]
assert len(setups) == 10, len(setups)
# rasterize coarsely on a 104x64 grid (1/8 scale) and pin coverage
W, H = 104, 64
drawn = 0
covered = 0
zwins = 0
zbuf = [[-1e30] * W for _ in range(H)]
for a in setups:
edges = []
p = a + 4
while ((rd(p) >> 8) & 0xff) == 0x2c and len(edges) < 6:
edges.append((f32(rd(p+4)), f32(rd(p+8)), f32(rd(p+12))))
p += 16
zp = None
q = p
for _ in range(40):
w = rd(q)
if ((w >> 8) & 0xff) == 0x21:
zp = (f32(rd(q+4)), f32(rd(q+8)), f32(rd(q+12)))
break
q += 4
if zp is None:
continue
any_px = False
for gy in range(H):
for gx in range(W):
x, y = gx * 8, gy * 8
vs = [A * x + B * y + C for A, B, C in edges]
inside = all(v >= 0 for v in vs) or all(v < 0 for v in vs)
if not inside:
continue
any_px = True
covered += 1
z = zp[0] * x + zp[1] * y + zp[2]
if z > zbuf[gy][gx]:
zbuf[gy][gx] = z
zwins += 1
if any_px:
drawn += 1
assert drawn == 10, drawn
assert covered == 5706, covered
assert zwins == 5119, zwins
return 'C4 fxtest primitives: PASS (10 drawn, coverage %d, zwins %d)' % (covered, zwins)
if __name__ == '__main__':
results = []
for case in (c1_dump_reference, c2_synthetic_triangle, c3_bars_pipeline,
c4_fxtest_primitives):
results.append(case())
print('\n'.join(results))
print('CONFORMANT')