Files
firestorm/MW4COMPARE/tools/decompile/verify_smallmodel.py
T
83478b7666 Add MW4COMPARE: .mw4 decompiler toolchain and V4H comparison harness
Tooling built to recover editable source for six 'Mech chassis that exist
in the parallel FS_Build_V4H build but not in this repo. Reverse-engineers
every compiled record type in the .mw4 package format back to the .data /
.instance / .subsystems / .damage / .contents / .torso / .engine /
.armature sources the content pipeline consumes.

Nothing here is wired into the game build. It is a standalone analysis
harness run from Linux.

Package format
--------------
"#VBD" container. Directory records are [len][name][FILETIME][origSize]
[storedSize][offset], payload base at dword 0x0C. A record is stored raw
when storedSize == origSize, otherwise LZW (9->12-bit LSB-first codes,
256=clear, 257=EOF, dict from 258), per Database.cpp:451.

GameModel records are flat /Zp4 structs following the C++ inheritance
chain Entity(0) -> Mover(28) -> MWObject(80) -> Vehicle(664) -> Mech(756),
1636 bytes total. CreateMessage records follow Replicator -> Entity ->
Mover -> MWMover -> MWObject -> Vehicle -> Mech from start=16 (the
undeclared Connection__Message header), ending at 341 and padded to 344.

tools/decompile/
----------------
  datamap.py        header-driven layout engine; CHAIN + ANCHORS
                    {Vehicle:664, Mech:756} assert the struct offsets
  mw4msg.py         CreateMessage reader/walker
  data.py           .data      constants.py  define/table symbol resolution
  damage.py         .damage    contents.py   .contents
  smallmodel.py     .torso + .engine         instance.py  .instance
  armature.py / armature_parts.py  .armature + armaturedata/armaturevideo
  assembly.py       joint hierarchy renderer
  make_generic_doll.py  builds generic MFD/Radar damage dolls
  verify_*.py       per-type round-trip verifiers

Verified round-trip across all 64 shared chassis:
  .armature      2938/2976 pages     .subsystems  7579/7585 keys
  .data map      6071/6071 values    .data trip   8291/8306 keys
  .damage        6605/6605 keys      .contents    7480/7480 keys
  .torso+.engine 1280/1280 keys      .instance     896/896 keys, 64/64 pages
  armature_parts 1202/1202 .data, 1149/1202 .video

Layout-discovery lessons (documented in DECOMPILING.md)
-------------------------------------------------------
- Never let a field map be discovered by the values that verify it. A
  value-matching pass reported 4288/4288 while mis-assigning 34 keys. The
  map was rebuilt from header declaration order, anchored on uniquely
  resolved fields.
- Read the factory, not the data. 12 .data fields and 5 Torso fields are
  declared plain Stuff::Scalar but multiplied by Radians_Per_Degree in
  Mech_Tool.cpp:889 / Torso_Tool.cpp.
- Strip typedefs before walking a header. A stray `typedef int AttributeID;`
  masked a missing ClassID - two 4-byte errors cancelling out, caught only
  by the ANCHORS assertion.
- A verifier that silently narrows its own input reports success. Braced
  blocks must be hidden before splitting pages, replacing both CR and LF,
  because a `Shadow={...}` block contains a line reading `[shadow]` and
  splitlines() also splits on bare CR.
- NSWIZZLE is undefined, so the #else branch is live and orders members
  differently. bool is 1 byte; char x[MaxStringLength] is 256.
- V4H carries stale Mech IDs (their Atlas is 5, ours 6), so 64 of 65 shared
  chassis are off by one; --retarget-ids emits $(M_<Chassis>)/$(IDS_<Chassis>).

reports/ holds generated diffs. The two ~5 MB manifest-*.tsv intermediates
are gitignored; regenerate everything with run-comparison.sh.

Co-authored-by: Claude Opus 5 (Anthropic) <noreply@anthropic.com>
Co-authored-by: GitHub Copilot <copilot@github.com>
2026-08-08 16:47:57 -05:00

105 lines
3.9 KiB
Python

#!/usr/bin/env python3
"""Round-trip verifier for the .torso and .engine decompilers.
Regenerates both files for every chassis and compares key order and values
against the authored source. `$(SYMBOL)` and its numeric expansion are treated
as equal, since the record stores only the resolved float.
python3 verify_smallmodel.py [--show KIND CHASSIS]
"""
import argparse, collections, glob, os, re, sys
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
import smallmodel
REC = "/home/rich/Repositories/FS_Ours_extracted/core/mechs"
SRC = "/home/rich/Repositories/firestorm/Gameleap/mw4/Content/Mechs"
NUM = re.compile(r'^-?(?:\d+\.?\d*|\.\d+)$')
def resolve(value, symbols):
"""$(NAME) -> its numeric value, so symbol and literal compare equal."""
m = re.fullmatch(r'\$\((\w+)\)', value.strip())
if m:
return symbols.get(m.group(1).lower(), value.strip().lower())
return value.strip().lower()
def canon(value, symbols):
v = resolve(value, symbols)
return f"{float(v):.5g}" if NUM.match(str(v)) else str(v)
def symbol_values(path):
"""-> {symbolNameLower: numericString}"""
out = {}
if os.path.exists(path):
txt = re.sub(r'//[^\n]*', '', open(path, encoding="latin-1", errors="replace").read())
for name, val in re.findall(r'^\s*!(\w+)\s*=\s*([-\d.]+)\s*$', txt, re.M):
out[name.lower()] = val
return out
def source_kv(path):
txt = re.sub(r'//[^\n]*', '', open(path, "rb").read().decode("latin-1"))
return [(m.group(1), m.group(2).strip())
for m in re.finditer(r'^([A-Za-z]\w*)=([^\r\n]*)', txt, re.M)]
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--show", nargs=2, metavar=("KIND", "CHASSIS"))
args = ap.parse_args()
t = collections.Counter()
bad = collections.Counter()
examples = []
order_bad = []
for kind, spec in sorted(smallmodel.SPECS.items()):
symbols = symbol_values(spec["defines"])
for d in sorted(glob.glob(REC + "/*")):
ch = os.path.basename(d)
src = [p for p in glob.glob(f"{SRC}/*/*.{kind}")
if os.path.basename(p).lower() == f"{ch.lower()}.{kind}"]
if not src or smallmodel.record(d, kind) is None:
continue
t[f"{kind} files"] += 1
want = source_kv(src[0])
got = smallmodel.decompile(d, kind)
if args.show and args.show[0] == kind and args.show[1].lower() == ch.lower():
sys.stdout.write(smallmodel.emit(kind, got))
return
if [k.lower() for k, _ in want] != [k.lower() for k, _ in got]:
order_bad.append((kind, ch, [k for k, _ in want], [k for k, _ in got]))
continue
for (wk, wv), (_gk, gv) in zip(want, got):
t["keys"] += 1
if canon(wv, symbols) == canon(gv, symbols):
t["ok"] += 1
else:
bad[f"{kind}.{wk}"] += 1
if len(examples) < 12:
examples.append((kind, ch, wk, wv, gv))
for kind in sorted(smallmodel.SPECS):
print(f"{kind:8s} files : {t[kind + ' files']}")
print(f"keys compared : {t['keys']} exact: {t['ok']} wrong: {t['keys'] - t['ok']}")
if order_bad:
print(f"\nkey order mismatches: {len(order_bad)}")
for kind, ch, w, g in order_bad[:3]:
print(f" {kind} {ch}\n want {w}\n got {g}")
if bad:
print("\nvalue mismatches:")
for k, n in bad.most_common(15):
print(f" {k:34s} x{n}")
print("\nexamples (kind, chassis, key, source, decoded):")
for e in examples:
print(f" {e[0]:7s} {e[1]:14s} {e[2]:22s} want={e[3]!r:22s} got={e[4]!r}")
if __name__ == "__main__":
main()