Files
firestorm/MW4COMPARE/tools/classify-survivors.py
T
83478b7666 Add MW4COMPARE: .mw4 decompiler toolchain and V4H comparison harness
Tooling built to recover editable source for six 'Mech chassis that exist
in the parallel FS_Build_V4H build but not in this repo. Reverse-engineers
every compiled record type in the .mw4 package format back to the .data /
.instance / .subsystems / .damage / .contents / .torso / .engine /
.armature sources the content pipeline consumes.

Nothing here is wired into the game build. It is a standalone analysis
harness run from Linux.

Package format
--------------
"#VBD" container. Directory records are [len][name][FILETIME][origSize]
[storedSize][offset], payload base at dword 0x0C. A record is stored raw
when storedSize == origSize, otherwise LZW (9->12-bit LSB-first codes,
256=clear, 257=EOF, dict from 258), per Database.cpp:451.

GameModel records are flat /Zp4 structs following the C++ inheritance
chain Entity(0) -> Mover(28) -> MWObject(80) -> Vehicle(664) -> Mech(756),
1636 bytes total. CreateMessage records follow Replicator -> Entity ->
Mover -> MWMover -> MWObject -> Vehicle -> Mech from start=16 (the
undeclared Connection__Message header), ending at 341 and padded to 344.

tools/decompile/
----------------
  datamap.py        header-driven layout engine; CHAIN + ANCHORS
                    {Vehicle:664, Mech:756} assert the struct offsets
  mw4msg.py         CreateMessage reader/walker
  data.py           .data      constants.py  define/table symbol resolution
  damage.py         .damage    contents.py   .contents
  smallmodel.py     .torso + .engine         instance.py  .instance
  armature.py / armature_parts.py  .armature + armaturedata/armaturevideo
  assembly.py       joint hierarchy renderer
  make_generic_doll.py  builds generic MFD/Radar damage dolls
  verify_*.py       per-type round-trip verifiers

Verified round-trip across all 64 shared chassis:
  .armature      2938/2976 pages     .subsystems  7579/7585 keys
  .data map      6071/6071 values    .data trip   8291/8306 keys
  .damage        6605/6605 keys      .contents    7480/7480 keys
  .torso+.engine 1280/1280 keys      .instance     896/896 keys, 64/64 pages
  armature_parts 1202/1202 .data, 1149/1202 .video

Layout-discovery lessons (documented in DECOMPILING.md)
-------------------------------------------------------
- Never let a field map be discovered by the values that verify it. A
  value-matching pass reported 4288/4288 while mis-assigning 34 keys. The
  map was rebuilt from header declaration order, anchored on uniquely
  resolved fields.
- Read the factory, not the data. 12 .data fields and 5 Torso fields are
  declared plain Stuff::Scalar but multiplied by Radians_Per_Degree in
  Mech_Tool.cpp:889 / Torso_Tool.cpp.
- Strip typedefs before walking a header. A stray `typedef int AttributeID;`
  masked a missing ClassID - two 4-byte errors cancelling out, caught only
  by the ANCHORS assertion.
- A verifier that silently narrows its own input reports success. Braced
  blocks must be hidden before splitting pages, replacing both CR and LF,
  because a `Shadow={...}` block contains a line reading `[shadow]` and
  splitlines() also splits on bare CR.
- NSWIZZLE is undefined, so the #else branch is live and orders members
  differently. bool is 1 byte; char x[MaxStringLength] is 256.
- V4H carries stale Mech IDs (their Atlas is 5, ours 6), so 64 of 65 shared
  chassis are off by one; --retarget-ids emits $(M_<Chassis>)/$(IDS_<Chassis>).

reports/ holds generated diffs. The two ~5 MB manifest-*.tsv intermediates
are gitignored; regenerate everything with run-comparison.sh.

Co-authored-by: Claude Opus 5 (Anthropic) <noreply@anthropic.com>
Co-authored-by: GitHub Copilot <copilot@github.com>
2026-08-08 16:47:57 -05:00

90 lines
3.2 KiB
Python

#!/usr/bin/env python3
"""
classify-survivors.py - split a pruned extracted tree into NEW vs DIFFERS.
python3 classify-survivors.py <prunedExtractedDir>
Run after prune-identical.py. Everything still present is something we do not
have byte-for-byte; this reports which of two reasons applies:
NEW no file at that entry path exists on our side at all
DIFFERS we have that entry path, but the bytes differ
"Our side" means our extracted packages OR our Content/ source tree, matched
case-insensitively - the same notion of "have it" the pruner uses.
Writes _classified.tsv next to the tree and prints a summary.
"""
import sys, os, collections
OURS_EXTRACTED = "/home/rich/Repositories/FS_Ours_extracted"
OUR_SOURCE = "/home/rich/Repositories/firestorm/Gameleap/mw4/Content"
def index_paths(root, skip_top=()):
out = {}
for dp, dirs, fs in os.walk(root):
if dp == root:
dirs[:] = [d for d in dirs if d.lower() not in skip_top]
for f in fs:
p = os.path.join(dp, f)
out[os.path.relpath(p, root).replace("\\", "/").lower()] = p
return out
def package_of(rel):
parts = rel.split("/")
head = parts[0].lower()
if head in ("core", "props", "textures"):
return parts[0], "/".join(parts[1:])
if head in ("maps", "missions", "variants") and len(parts) > 2:
return "/".join(parts[:2]), "/".join(parts[2:])
if head == "pilots" and len(parts) > 3:
return "/".join(parts[:3]), "/".join(parts[3:])
return parts[0], "/".join(parts[1:])
def main():
if len(sys.argv) < 2:
sys.exit(__doc__)
root = sys.argv[1]
merged = os.path.join(root, "_merged")
ours = index_paths(OURS_EXTRACTED, skip_top={"_merged"})
ours_entries = set()
for rel in ours:
ours_entries.add(package_of(rel)[1])
src_entries = set(index_paths(OUR_SOURCE))
rows, summary = [], collections.Counter()
for dp, _dirs, fs in os.walk(root):
if dp == merged or dp.startswith(merged + os.sep):
continue
for f in fs:
rel = os.path.relpath(os.path.join(dp, f), root).replace("\\", "/")
if rel.startswith("_"):
continue
pkg, entry = package_of(rel)
e = entry.lower()
verdict = "DIFFERS" if (rel.lower() in ours or e in ours_entries
or e in src_entries) else "NEW"
rows.append((verdict, rel, pkg, entry))
summary[(verdict, pkg.split("/")[0])] += 1
with open(os.path.join(root, "_classified.tsv"), "w", encoding="utf-8") as fh:
fh.write("verdict\tpath\tpackage\tentry\n")
for r in sorted(rows):
fh.write("\t".join(r) + "\n")
tot = collections.Counter(v for v, _, _, _ in rows)
print(f"survivors: {len(rows)} NEW={tot['NEW']} DIFFERS={tot['DIFFERS']}\n")
print(f"{'package group':16s} {'NEW':>7s} {'DIFFERS':>8s}")
groups = sorted({g for _v, g in summary})
for g in groups:
print(f"{g:16s} {summary[('NEW', g)]:7d} {summary[('DIFFERS', g)]:8d}")
print(f"\n-> {os.path.join(root, '_classified.tsv')}")
if __name__ == "__main__":
main()