Per-priority occupancy, measured every second until the fault: [queues] p0=BUSY p1=- p2=BUSY p3=- p4=- nextReady=1 That single line kills the two leading theories. p3/p4 empty means nothing above priority 0 is competing, so the pump is free to serve it -- no starvation. nextReady=1 means PeekAtNextEvent always has a READY event, so priority 0 is not holding a timed event whose alarm never comes due -- no deadlock. What remains is refill: priority 0 is replenished as fast as the pump drains it, ~143 events per simulated second. It is also fully deterministic. Two runs of the same binary reported pump=352/1499 at frame 1001 and 495/2643 at frame 2002, identical to the byte. So 'crashes about half the time' was never a race -- it was comparing runs that differed in binary or conf. Only InterestManager::PostRendererEvent posts at priority 0 (it maps every renderer event there while the app is not RunningMission), so naming the message names the flooder. Application::Post now tallies priority-0 posts by message ID and reports the busiest three. Notably, all three producers of NotifyOfNewInterestingEntity are entity-CREATION paths, so if that ID dominates then something is creating entities in a loop -- which would explain the unbounded growth behind the fault as well as the hang. pod_render_noskl.conf now carries the same probes, so one instrumented binary can be run with the skeleton walk on and off. Walk-off runs are the only configuration of ours known to reach 'Turning Plasma Score Display On'. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
96 lines
3.3 KiB
Bash
96 lines
3.3 KiB
Bash
#!/bin/bash
|
|
#
|
|
# podrun -- stage the current build and start one pod run, keeping the log.
|
|
#
|
|
# podrun.sh [conf-basename] [tag]
|
|
#
|
|
# podrun.sh # pod_render_rec.conf, our build
|
|
# podrun.sh pod_render_shp # the shipped exe, for A/B
|
|
# podrun.sh pod_render_rec noskl
|
|
#
|
|
# WHY THIS EXISTS
|
|
#
|
|
# Two traps have each already cost a round of wrong conclusions on this rig:
|
|
#
|
|
# 1. STALE BINARY. The confs run BTL4REC.EXE out of ALPHA_1/REL410/BT,
|
|
# while the build writes build410/btl4opt.exe. Forgetting to copy means
|
|
# silently measuring the previous binary. This script always stages and
|
|
# always prints the timestamp it staged, so the evidence is in the log.
|
|
#
|
|
# 2. OVERWRITTEN LOG. serial3 truncates podlog.txt on every launch, so the
|
|
# only preserved crash dump on this rig was an accident. Runs are four
|
|
# minutes and the defect is intermittent, so a lost log is a lost run.
|
|
# Every run is archived to logs/<tag>-<n>.txt when the NEXT one starts.
|
|
#
|
|
# The rule this rig taught: an intermittent fault needs TWO AGREEING RUNS
|
|
# before you believe anything. One run is a story, not evidence.
|
|
#
|
|
set -e
|
|
|
|
BR=/c/VWE/TeslaRel410/emulator/render-bridge
|
|
IMG=/c/VWE/TeslaRel410/ALPHA_1/REL410/BT
|
|
SRC=/c/VWE/TeslaRel410/restoration/build410/btl4opt.exe
|
|
CONF=${1:-pod_render_rec}
|
|
TAG=${2:-$CONF}
|
|
|
|
mkdir -p "$BR/logs"
|
|
|
|
# archive whatever the last run left behind before serial3 truncates it
|
|
if [ -s "$BR/podlog.txt" ]; then
|
|
N=$(ls "$BR/logs" 2>/dev/null | wc -l)
|
|
PREV="$BR/logs/$(printf '%03d' "$N")-prev.txt"
|
|
cp "$BR/podlog.txt" "$PREV"
|
|
if grep -q "page you don't own" "$PREV"; then
|
|
mv "$PREV" "${PREV%-prev.txt}-CRASH.txt"
|
|
echo "archived previous run (CRASHED) -> $(basename "${PREV%-prev.txt}-CRASH.txt")"
|
|
else
|
|
echo "archived previous run -> $(basename "$PREV")"
|
|
fi
|
|
[ -s "$BR/dosboxlog.txt" ] && cp "$BR/dosboxlog.txt" "$BR/logs/$(printf '%03d' "$N")-dosbox.txt"
|
|
fi
|
|
|
|
if tasklist //FI "IMAGENAME eq dosbox-x.exe" 2>/dev/null | grep -qi dosbox-x; then
|
|
echo "killing running dosbox-x"
|
|
taskkill //F //IM dosbox-x.exe > /dev/null 2>&1 || true
|
|
sleep 2
|
|
fi
|
|
|
|
case "$CONF" in
|
|
*shp*)
|
|
echo "shipped side -- no staging"
|
|
;;
|
|
*)
|
|
[ -f "$SRC" ] || { echo "no build at $SRC -- run build410.sh link"; exit 1; }
|
|
cp -f "$SRC" "$IMG/BTL4REC.EXE"
|
|
echo "staged BTL4REC.EXE <- btl4opt.exe $(ls -la "$SRC" | awk '{print $5" bytes "$6" "$7" "$8}')"
|
|
;;
|
|
esac
|
|
|
|
#
|
|
# HOST-SIDE BOARD ENV. These belong to the VPX device inside DOSBox, not to
|
|
# the guest, so they must be exported here. Launching without them does not
|
|
# merely disable rendering -- the iserver handshake fails outright with
|
|
# "Protocol error : length 65535 too big" and the run dies during boot.
|
|
#
|
|
WORK="$LOCALAPPDATA/Temp/vwe-pod"
|
|
mkdir -p "$WORK"
|
|
export VPXLOG="$WORK/vpxresp.txt"
|
|
export VPX_RESPOND=1
|
|
export VPX_RENDER=1
|
|
export VPX_EXPLODE=1
|
|
export VPX_DUMPDIR="$WORK"
|
|
export VPX_NOMAIN=1
|
|
export VPX_CLEAR=0,0,0
|
|
export RIO_TAP="$WORK/riotap.txt"
|
|
|
|
rm -f "$BR/podlog.txt"
|
|
cd /c/VWE/TeslaRel410/emulator/src/src
|
|
#
|
|
# Keep the emulator's own stdout/stderr. The game's stdout comes back over
|
|
# COM3, but the DPMI host's exception report and any device-side complaint go
|
|
# to the emulator, and those were being thrown away.
|
|
#
|
|
./dosbox-x.exe -conf "$BR/$CONF.conf" > "$BR/dosboxlog.txt" 2>&1 &
|
|
echo "launched $CONF (tag=$TAG) -- pid $!"
|
|
echo "watch: tail -f $BR/podlog.txt"
|