1. UDP ENDPOINT HIJACK. `from_host` in a UDP envelope is the sender's OWN claim,
and the relay used it directly to refresh that host's downstream endpoint --
so ANY datagram claiming host N silently stole host N's traffic (a zombie pod
from a previous round, a stale NAT mapping, or anyone who guessed a host id).
The victim simply stopped receiving on a channel that still looked healthy.
The claim is now bound to the identity we actually authenticated: the IP of
that host's live TCP game connection. The PORT is deliberately not checked --
it moves on a NAT rebind, which is the whole reason the endpoint map refreshes
per datagram -- and a genuine IP change cannot happen without the TCP
connection breaking and re-registering, so a legitimate pod is never rejected.
Rejections are counted (udp_spoofed) and logged once per offending (host, IP).
Known limit: two pods on one machine share an IP, so this cannot separate
them; same-machine trust is assumed.
2. THE EGG ACK COULD BE MISSED ENTIRELY -- a silent, unrecoverable wedge. The
pod's ACK is the launch gate's ONLY signal, and it was detected by parsing a
single recv() at fixed offset 0 behind a `len(data) >= 24` test. On a real
network (loopback hid both cases) the 28-byte ACK arriving SPLIT -- 16 bytes
then 12 -- was dropped by that test and never looked at again, and two
COALESCED messages meant only the first was read. A missed ACK means that
seat never counts toward the gate, so the round can never be released.
_console_read now buffers per connection and walks every complete frame
(16 + messageLength); an implausible length drops the connection WITH A REASON
rather than mis-reading that pod all night, since a stream protocol cannot be
resynced by guessing. Bounds: CONSOLE_MSG_MAX 64K, CONSOLE_INBUF_MAX 1 MiB.
3. REMOTE-OPERATOR MODE WAS HALF A CONSOLE. LAUNCH and END MISSION could never
enable (both conditions required `console_proc is not None`, which the remote
path never sets), the pilot lights were permanently blank (SessionMonitor was
built with an empty tag list, so `seated` was always 0), and Launch-local was
refused by the same guard even though the code below it already built
BT_RELAY from the remote host. All three enable conditions now accept EITHER
channel, and the monitor ADOPTS THE ROSTER from the relay's
"roster: N pilot(s) -> hostIDs [...]: [...]" line -- which the relay replays to
every newly AUTHed operator -- so a remote operator gets real lights and a
real seat count.
VERIFIED
* scratchpad/test_relay_net.py (new, 17 checks): spoofed datagram cannot move
an endpoint and is counted; a NAT rebind on the same IP still is honoured; an
unregistered host id still dropped; the ACK is found whole, split in two,
split three ways mid-header, and when hiding behind another message;
a garbage length drops the conn; the roster line is adopted, maps a
subsequent SEATED onto a real tag, lifts `seated` off 0, and re-feeding it
preserves known state.
* Live 2-pod rig: real pods ACKed through the new reassembly path (2/2 ready,
zero desync drops), the mission launched and ran, UDP flowed throughout
(93 rx / 85 tx, udp-known [2,3]) with ZERO false spoof rejections, then a
clean StopMission + round RESET.
* scratchpad/test_relay_rearm.py still 19/19; checkctx CLEAN.
Docs updated to match (context/operator-console.md gains reassembly and
anti-hijack sections; the guide no longer claims remote mode is crippled).
Remaining open items are now recorded in the topic's frontmatter: mesh mode's
operator buttons are inert by construction (no stdin reader in that path -- left
alone, mesh self-launches), _log_launch_readiness can cry STALL on a healthy
launch-with-whoever, and a straggler's late FIN is still misread as a pod dying
mid-load.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
19 KiB
id, title, status, source_sections, related_topics, key_terms, open_questions
| id | title | status | source_sections | related_topics | key_terms | open_questions | ||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| operator-console | The Operator Console + Relay (btoperator.py / btconsole.py) | living | tools/btoperator.py; tools/btconsole.py (module docstring is authoritative on the wire); context/multiplayer.md §relay/§operator console; docs/OPERATOR_GUIDE.md |
|
|
|
The Operator Console + Relay
The system-operator station: build a mission, run a session, watch pods arrive, launch rounds. This is the piece the 1995 archive lost and we rebuilt.
Sysop-facing instructions live in
docs/OPERATOR_GUIDE.md. This file is the engineering description — how it is built and where it breaks.
READ THIS FIRST — there are TWO programs and they are not interchangeable
The single most common mistake (made by me, twice, after a context compaction):
| File | What it is | How it runs |
|---|---|---|
tools/btoperator.py |
The operator console. A PySide6 GUI — the window the sysop actually uses to build missions, watch pods and press LAUNCH. | python tools/btoperator.py (from the repo root, no args) |
tools/btconsole.py |
The headless console/relay. All the wire logic: egg hand-out, seat assignment, the launch pair, the mission clock, the TCP+UDP frame router. No UI. | The GUI spawns it as a QProcess; also runnable standalone from the CLI |
"Launch the console" means tools/btoperator.py. Backgrounding btconsole.py is not it —
it has no window, and with no stdin the operator commands are dead. The division is deliberate:
all protocol logic stays headless so the CLI, the GUI and the test rigs share one tested
implementation; the GUI is only UI + process management.
The GUI drives the relay two ways: stdin (launch, stop, rearm) for a local child, and
the control port for a remote relay. Both paths reach the same handlers.
Topology and ports
Pods dial OUT to the relay (this is what makes internet play work through NAT — no inbound
ports needed at the player's end). With console port P (default 1500):
| Port | Protocol | Purpose |
|---|---|---|
P /tcp |
console protocol, relay-terminated | streams the egg to each pod, then the RunMission launch pair and StopMission |
P+1 /tcp |
envelope-framed game relay | the frame router between pods ({int32 route; uint32 length} + payload) |
P+1 /udp |
envelope-framed unreliable channel | the 60 Hz update-record fan-out ({int32 route; int32 fromHost; uint32 seq} + frame) |
P+7 /tcp |
operator control channel | AUTH <secret> then launch / stop / rearm / ping / get / set (CONTROL_PORT_OFFSET = 7) |
| udp/15999 | LAN discovery | answers probes so pods can use BT_RELAY=auto |
Pod-side env contract: BT_RELAY=<host>:<P> (or auto for LAN discovery) and
BT_SELF=<their [pilots] entry> — or BT_SELF=auto, which makes the relay assign the next free
roster seat. players/join.bat sets these.
Route table (tools/btconsole.py:276-286)
>= 2 unicast to that hostID (relay->client: the SENDER's hostID)
-1 broadcast to every registered pod except the sender
-2 HELLO payload: int32 magic 'BTR1', int32 myHostID
-3 PEER_UP relay->client
-4 PEER_DOWN relay->client
-5 UDP HELLO-ACK (sets the client's udpUp)
-6 SEAT_REQUEST client->relay: assign me a free roster seat
-7 SEAT_ASSIGN relay->client: int32 hostID + NUL-terminated tag
-8 SEAT_FULL relay->client: no free seats
-9 MATCHLOG client->relay: upload (filename NUL + bytes); saved to
matchlogs/<peer>_<name>, cap 8 MB
-10 READY client->relay: mission load complete (the READY light)
-11 REJOIN relay->client: abandon this round and rejoin
The roster, seats and identity
The egg's [pilots] list is the roster (parse_egg_roster, btconsole.py:372). Entry i
becomes hostID 2 + i — FIRST_GAME_HOST_ID = 2, because CONSOLE_HOST_ID = 1 is
FirstLegalHostID and belongs to the console itself. A pilot entry's text is its tag
(e.g. 10.99.0.1:1502); the tag's position in the received egg is how a pod derives its own
host id, which is why a roster trim remaps seat ids and everything re-keys by position.
- Seat requests (route -6) let a walk-up pod claim a free seat and carry its callsign + mech.
- Reservations last
SEAT_RESERVE_SECONDS = 60. - Reclaim by identity: a returning player (same IP + callsign) gets their old seat back
(
seat_identity, btconsole.py:1258-1276). Watch for this in the log asseat RECLAIMED by identity.
Launch with whoever connects: the operator's LAUNCH can _trim_roster down to the seats that
actually have live beacons or claims, so a 8-seat egg runs a 3-player round.
The round lifecycle (the part that used to strand the operator)
session start
-> egg HELD until len(console_conns) >= len(roster) [btconsole.py:565-578]
-> _release_eggs() (or _release_eggs(present) -> _trim_roster) [:627]
-> pod ACK per pod (clientID 0, msgID 4) -> _pods_ready()
-> READY (route -10) per pod: mission load complete
-> operator LAUNCH manual mode: waits for the operator [_tick_launch]
-> RunMission #1 then #2 after LAUNCH_STEP_SECONDS = 4s launches_sent 1 -> 2
-> mission clock the egg's [mission] length, +STOP_GRACE = 2s [_tick_stop]
-> StopMission clock expiry or operator End Mission
-> round end pods exit, join.bat relaunches them, they reconnect
-> RE-ARM so the next LAUNCH works <-- see below
Timings: LAUNCH_SETTLE_SECONDS = 20 (egg → first launch; pods must reach WaitingForLaunch or the
handler Fail()s), LAUNCH_STEP_SECONDS = 4, STOP_GRACE_SECONDS = 2,
ABORT_SETTLE_SECONDS = 8 (egg-release hold after a round abort).
The all-ready gate
Firing a mission while a pod is still LOADING drowns that pod's load in the running mission's update streams (same event queue) — the operator, who joins last, once loaded 5+ minutes. So the RunMission pair is held until every registered pod has sent READY. Escape hatches: the holdouts are announced every 10 s, and a 180 s cap force-fires so one wedged pod cannot hold the night hostage.
Re-arm: why LAUNCH used to do nothing after a round [FIXED 2026-07-26]
The failure the operator reported: "after a game ends I push LAUNCH and nothing happens until I reset the session or reboot the console." Their own counter-observation is the tell: with a stable group it worked fine repeatedly.
The manual LAUNCH branch is guarded on launches_sent == 0, but a finished mission leaves it at
2. The only automatic paths back to 0 were both all-or-nothing:
_check_launch_gate— needs every seat of the last round to re-ACK_maybe_reset_round— needs every pod to be gone
A games night lives between those: most pods rejoin, one player closes their window. And the
"not all seats filled" diagnostic sat inside the launches_sent == 0 guard, so in exactly that
state nothing printed at all. The GUI agreed: its LAUNCH button is gated on
not monitor.launched, and launched was cleared only by the two relay lines those same gates
emit — so the button was greyed out precisely when wedged. Only a session restart recovered,
because a fresh relay process means fresh state.
Now: an explicit operator LAUNCH is itself sufficient authority to start a new round.
_rearm_for_new_round() clears the round state and restores the template roster/egg while
keeping seats, beacons and connections. Plus a rearm/newround command and a Re-arm
button, so recovery never needs a restart.
The trap to remember if you touch this: launch_requested deliberately stays set across the
RunMission pair, so re-arming on launches_sent >= 1 alone also matches the normal state
between #1 and #2 — that resets the pair mid-flight and re-releases eggs every tick, a re-arm
storm that kicks every pod into an identity-resync loop. The re-arm therefore also requires
launch_at is None (nothing in flight). Caught on the rig, 2026-07-26.
The post-abort dead lobby [FIXED 2026-07-26]
A distinct wedge with the same symptom, live-proven in content/operator_relay.log on
2026-07-25 (≈5 minutes of a night lost). _abort_round sets eggs_released = False before the
surviving pods' sockets close, so every later _maybe_reset_round hits its own
if not self.eggs_released: return and the template restore never runs: roster and
expected_ids stay trimmed to last round's attendance and egg_path stays the trimmed egg. All
three automatic release paths compare against len(self.roster), so with N−1 pods back nothing
releases, walk-ups get ROSTER FULL, and the relay sits dead.
_rearm_for_new_round's template restore is therefore unconditional — an earlier cut gated it
on eggs_released, which made Re-arm useless in the one state that most needs it. It also clears
round_hold_until so an abort's settle window is not inherited.
Aborts happen for an ordinary reason: nothing on the wire distinguishes "a straggler's round-N FIN arriving late" from "a pod died mid-load", so a slow or remote pod closing ~18 s after StopMission is read as the latter.
The stop latch
_tick_stop returns early while launches_sent < 2, so a stop_requested set between rounds
used to sit latched and then StopMission the next mission in the very tick it launched — which from
the operator's seat is also "launch did nothing". It is now consumed and announced while idle.
Liveness (the "reboot the console" half) [FIXED 2026-07-26]
Nothing used to notice a pod that vanished without closing its TCP: a sleeping laptop, a
dropped Wi-Fi link, a killed process, a NAT idle-timeout. There was no SO_KEEPALIVE and no idle
deadline anywhere, so the dead connection sat in console_conns with acked=True forever,
which (a) permanently blocked _maybe_reset_round (it requires no conn acked) and (b) held a
phantom seat in the launch count. Only restarting the relay cleared it.
Now: keepalive on every accepted socket (_enable_keepalive), a last_seen stamp on every read
and on every inbound UDP datagram, and _reap_dead_peers() with DEAD_PEER_SECONDS = 180.
The reaper is deliberately narrow, and the narrowness is load-bearing. It runs only in the
active staging window — egg released, mission not yet running — and skips console pads with no egg
sent and game conns that have not HELLO'd. The first cut was gated only on "no mission running",
which armed it for the entire between-rounds wait, and that is a period when a pod is
legitimately byte-silent: its seat beacon is write-only for the process lifetime
(L4NET.CPP: "the relay ignores its silence", and the pod sets its own keepalive on it), and its
console pad has no egg yet so it cannot ACK. A real night showed 12-minute and 5-minute gaps
between rounds, so a 180 s deadline would have dropped healthy players and forced their clients
to relaunch — strictly worse than the ghost the reaper exists to remove. Caught in review before it
ever ran live. Regression guards: scratchpad/test_relay_rearm.py cases 5b/5c.
Half-open sockets are TCP keepalive's job; this deadline is only the backstop for platforms where the keepalive tunables do not take.
Console-protocol reassembly [FIXED 2026-07-26]
The pod's egg ACK is the launch gate's only signal, and it used to be detected by parsing a
single recv() at fixed offset 0 with a len(data) >= 24 test. Two ways that lost an ACK on a real
network (loopback hid both): the 28-byte ACK arriving split (16 bytes then 12) was dropped by
the length test and never looked at again, and two coalesced messages meant only the first was
read. A missed ACK means that seat never counts, so the round can never be released — a silent,
unrecoverable wedge. _console_read now buffers per connection and walks every complete frame
(16 + messageLength), and an implausible length drops the connection with a reason rather than
mis-reading that pod for the rest of the night (a stream protocol cannot be resynced by guessing).
Bounds: CONSOLE_MSG_MAX = 65536, CONSOLE_INBUF_MAX = 1 MiB.
UDP endpoint anti-hijack [FIXED 2026-07-26]
from_host in a UDP envelope is the sender's own claim, and the relay used it directly to
refresh that host's downstream endpoint. So any datagram claiming host N silently stole host N's
traffic — a zombie pod from a previous round, a stale NAT mapping, or anyone who guessed a host
id — and the victim simply stopped receiving on a channel that still looked healthy. The claim is
now bound to the identity we actually authenticated: the IP of that host's live TCP game
connection. The port is deliberately not checked (it moves on a NAT rebind, which is the whole
reason the endpoint map refreshes per datagram), and a genuine IP change cannot happen without the
TCP connection breaking and re-registering, so this never rejects a legitimate pod. Rejections are
counted (udp_spoofed) and logged once per offending (host, IP) pair.
Limitation worth knowing: two pods on the same machine share an IP, so this cannot tell them apart — same-machine trust is assumed.
Also: every pod-facing send used to flip the socket to blocking with no timeout on the relay's
single-threaded selector loop, so one wedged peer could freeze the whole relay (no launch tick,
no UDP fan-out, no control channel). All sends now go through _send_all_guarded
(SEND_TIMEOUT_SECONDS = 10, drops the peer on failure).
Modes
- Relay (internet) —
btconsole.py --relay <port> <egg>; pods dial out. The normal mode. - Mesh (LAN legacy) —
btconsole.py <egg> <host:port> ...; the console dials OUT to each pod's-netport. The authentic 1995 cabinet direction. - Remote relay — the GUI drives someone else's relay over the control port instead of spawning a child. Known broken: the LAUNCH button can never enable in this mode.
The console must STAY CONNECTED. A disconnect trips the pod's console-loss path, which also closes its game listener — an engine bug, not ours.
Logs and evidence
| File | Notes |
|---|---|
content/operator_relay.log |
the relay's own stdout, teed by the GUI. Truncated at each session start — so a Start Session destroys the evidence of the wedge it is being used to clear. Copy it first when diagnosing. |
matchlogs/<peer>_<name> |
per-pod match reports, auto-uploaded on route -9 at each clean round end (relay modes only; a mid-round crash loses it) |
content/join.log |
the pod's own log; never uploaded, only on the player's machine |
Every relay log line is wall-clock stamped (_StampedOut) so post-mortems can measure delays.
The GUI's monitor regexes all use .search(), so the prefix is transparent to them.
Mode-specific traps (verified 2026-07-26 by reading btoperator.py)
These are not theory — each was read out of the code, and each makes a control look live while doing nothing:
- MESH mode: the operator buttons are inert.
_relay_sendreports ">> operator LAUNCH sent" whenever the child process is Running, but the stdin command thread only exists inrelay_main(btconsole.py) — the mesh path (btconsole.main) has no stdin reader at all. So LAUNCH / END MISSION / Re-arm are silently discarded in mesh mode.Apply mission settingsis likewise inert there: the mesh console reads the egg bytes once at startup. Mesh also never prints "WAITING FOR OPERATOR", so the LAUNCH button cannot even enable. - REMOTE-operator mode — FIXED 2026-07-26. LAUNCH and END MISSION could never enable (both
conditions required
console_proc is not None, which the remote path never sets), the roster lights were permanently empty (SessionMonitor(True, [])→ no tags →seatedalways 0), andLaunch local instanceswas refused by the same guard even though the code below it already builtBT_RELAYfrom the remote host. Now: all three enable conditions accept either channel (console_procorremote_link), and the monitor adopts the roster from the relay'sroster: N pilot(s) -> hostIDs [...]: [...]line, which the relay replays to every newly AUTHed operator — so a remote operator gets real pilot lights and a real seat count. Start sessionsilently overwrites the egg on disk (including re-rasterized callsigns), with no prompt and no backup.Applywrites only the six[mission]keys (map/time/weather/scenario/temperature/length). Roster edits needSave; a seat-count change is refused by the relay withegg file roster changed N->M seats -- IGNORED (restart the session to resize).- The Color dropdown offers every colour in the RES vehicletable, but only
Black/Brown/Crimson/Green/Grey/Tan/Whiteactually paint — anything else silently renders GRAY (the relay warns). A "valid" mission can still produce a grey mech. operator_secret.txtis opened by a bare relative name by the relay, whose working directory is inherited from whatever directorybtoperator.pywas started in (the GUI never callssetWorkingDirectoryon the relay child). Start from the repo root and it will not be thecontent\operator_secret.txtthe tooltips point at.- After the relay dies on its own,
_console_finishedleavesconsole_procnon-None, soLaunch local instancesstays enabled and will spawn clients against a dead relay.
Gotchas / things that look like bugs
egg HELD (n/N pods present)is normal staging, not an error — the egg waits so every pod's copy carries every walk-up seat pref.*** WARNING: ... EMPTY seat(s) ... the mission will STALLcan fire misleadingly after a trim, because_log_launch_readinesscountsby_hostagainst the (remapped) roster. Do not act on it blindly.- A relay-side crash leaves the GUI looking alive —
_console_finisheddoes not clearconsole_proc. Commands now report honestly instead of logging ">> sent" into the void. - Session restart is the big hammer: it re-spawns the relay, so it clears every latch. That is why it always "fixed" things — and why it hid the real defects for so long.
- A
--manual-launchrelay waits for the operator; without it the launch pair fires automatically once all pods ACK.
Key Relationships
- Uses: content-archives (the EGG is the mission + roster) · multiplayer (what the relayed frames mean: master/replicant, dead reckoning)
- Feeds: multiplayer · steam-networking (the alternate wire seam) · pod-hardware
- Sysop instructions:
docs/OPERATOR_GUIDE.md