Tooling built to recover editable source for six 'Mech chassis that exist
in the parallel FS_Build_V4H build but not in this repo. Reverse-engineers
every compiled record type in the .mw4 package format back to the .data /
.instance / .subsystems / .damage / .contents / .torso / .engine /
.armature sources the content pipeline consumes.
Nothing here is wired into the game build. It is a standalone analysis
harness run from Linux.
Package format
--------------
"#VBD" container. Directory records are [len][name][FILETIME][origSize]
[storedSize][offset], payload base at dword 0x0C. A record is stored raw
when storedSize == origSize, otherwise LZW (9->12-bit LSB-first codes,
256=clear, 257=EOF, dict from 258), per Database.cpp:451.
GameModel records are flat /Zp4 structs following the C++ inheritance
chain Entity(0) -> Mover(28) -> MWObject(80) -> Vehicle(664) -> Mech(756),
1636 bytes total. CreateMessage records follow Replicator -> Entity ->
Mover -> MWMover -> MWObject -> Vehicle -> Mech from start=16 (the
undeclared Connection__Message header), ending at 341 and padded to 344.
tools/decompile/
----------------
datamap.py header-driven layout engine; CHAIN + ANCHORS
{Vehicle:664, Mech:756} assert the struct offsets
mw4msg.py CreateMessage reader/walker
data.py .data constants.py define/table symbol resolution
damage.py .damage contents.py .contents
smallmodel.py .torso + .engine instance.py .instance
armature.py / armature_parts.py .armature + armaturedata/armaturevideo
assembly.py joint hierarchy renderer
make_generic_doll.py builds generic MFD/Radar damage dolls
verify_*.py per-type round-trip verifiers
Verified round-trip across all 64 shared chassis:
.armature 2938/2976 pages .subsystems 7579/7585 keys
.data map 6071/6071 values .data trip 8291/8306 keys
.damage 6605/6605 keys .contents 7480/7480 keys
.torso+.engine 1280/1280 keys .instance 896/896 keys, 64/64 pages
armature_parts 1202/1202 .data, 1149/1202 .video
Layout-discovery lessons (documented in DECOMPILING.md)
-------------------------------------------------------
- Never let a field map be discovered by the values that verify it. A
value-matching pass reported 4288/4288 while mis-assigning 34 keys. The
map was rebuilt from header declaration order, anchored on uniquely
resolved fields.
- Read the factory, not the data. 12 .data fields and 5 Torso fields are
declared plain Stuff::Scalar but multiplied by Radians_Per_Degree in
Mech_Tool.cpp:889 / Torso_Tool.cpp.
- Strip typedefs before walking a header. A stray `typedef int AttributeID;`
masked a missing ClassID - two 4-byte errors cancelling out, caught only
by the ANCHORS assertion.
- A verifier that silently narrows its own input reports success. Braced
blocks must be hidden before splitting pages, replacing both CR and LF,
because a `Shadow={...}` block contains a line reading `[shadow]` and
splitlines() also splits on bare CR.
- NSWIZZLE is undefined, so the #else branch is live and orders members
differently. bool is 1 byte; char x[MaxStringLength] is 256.
- V4H carries stale Mech IDs (their Atlas is 5, ours 6), so 64 of 65 shared
chassis are off by one; --retarget-ids emits $(M_<Chassis>)/$(IDS_<Chassis>).
reports/ holds generated diffs. The two ~5 MB manifest-*.tsv intermediates
are gitignored; regenerate everything with run-comparison.sh.
Co-authored-by: Claude Opus 5 (Anthropic) <noreply@anthropic.com>
Co-authored-by: GitHub Copilot <copilot@github.com>
65 lines
2.3 KiB
Python
Executable File
65 lines
2.3 KiB
Python
Executable File
#!/usr/bin/env python3
|
|
"""
|
|
diffindex.py - case-insensitive package/entry diff between two mw4index manifests.
|
|
|
|
python3 diffindex.py <A.tsv> <B.tsv> [labelA] [labelB]
|
|
|
|
Per package it prints:
|
|
* name-set diff - which source assets exist on each side (the reliable signal)
|
|
* blob multiset - stored-byte md5 counts, names ignored (a rough upper bound
|
|
on how much payload differs; see the caveat in mw4index.py)
|
|
|
|
Packages present on only one side are reported as A-ONLY / B-ONLY.
|
|
"""
|
|
import sys, collections
|
|
|
|
|
|
def load(path):
|
|
names = collections.defaultdict(set)
|
|
blobs = collections.defaultdict(collections.Counter)
|
|
sizes = collections.defaultdict(dict)
|
|
with open(path, encoding="latin-1") as fh:
|
|
for line in fh:
|
|
line = line.rstrip("\n")
|
|
if not line:
|
|
continue
|
|
pkg, _rid, name, dlen, _rlen, h = line.split("\t")
|
|
pkg = pkg.lower()
|
|
key = name.lower().replace("\\", "/")
|
|
names[pkg].add(key)
|
|
blobs[pkg][h] += 1
|
|
sizes[pkg][key] = int(dlen)
|
|
return names, blobs, sizes
|
|
|
|
|
|
def main():
|
|
if len(sys.argv) < 3:
|
|
sys.exit(__doc__)
|
|
la = sys.argv[3] if len(sys.argv) > 3 else "A"
|
|
lb = sys.argv[4] if len(sys.argv) > 4 else "B"
|
|
an, ab, asz = load(sys.argv[1])
|
|
bn, bb, bsz = load(sys.argv[2])
|
|
|
|
for pkg in sorted(set(an) | set(bn)):
|
|
if pkg not in an:
|
|
print(f"== {pkg}: {lb}-ONLY PACKAGE ({len(bn[pkg])} entries)")
|
|
continue
|
|
if pkg not in bn:
|
|
print(f"== {pkg}: {la}-ONLY PACKAGE ({len(an[pkg])} entries)")
|
|
continue
|
|
aonly = sorted(an[pkg] - bn[pkg])
|
|
bonly = sorted(bn[pkg] - an[pkg])
|
|
shared = sum((ab[pkg] & bb[pkg]).values())
|
|
print(f"== {pkg}: {la}={len(an[pkg])} {lb}={len(bn[pkg])} | "
|
|
f"names: common={len(an[pkg] & bn[pkg])} {la}_only={len(aonly)} {lb}_only={len(bonly)} | "
|
|
f"blobs: identical={shared} {la}_uniq={sum(ab[pkg].values()) - shared} "
|
|
f"{lb}_uniq={sum(bb[pkg].values()) - shared}")
|
|
for n in aonly:
|
|
print(f" +{la} {n} ({asz[pkg][n]}B)")
|
|
for n in bonly:
|
|
print(f" -{lb} {n} ({bsz[pkg][n]}B)")
|
|
|
|
|
|
if __name__ == "__main__":
|
|
main()
|