This directory contains the implementation for Scalable Collision Avoidance
in Large Low Earth Orbit Constellations. It is built on the local OrbitZoo
fork, but thesis-specific components live under src/orbitzoo/thesis/ so the
generic orbital-dynamics library remains reusable.
The planned system uses discrete MAPPO under Centralized Training with Decentralized Execution (CTDE):
local k-neighbor observation -> shared actor -> one maneuver per satellite
full training-only constellation state -> centralized critic -> learning signal
The policy decides when and how to avoid a conjunction. Returning to the nominal orbit afterwards is out of scope and left to standard station-keeping; the drift avoidance causes, and the delta-v a return would cost, are measured in evaluation.
The actor will use seven actions: no-op, prograde, retrograde, radial-out, radial-in, cross-track positive, and cross-track negative. The critic is used only while training; the deployed/evaluated policy uses the shared actor and each satellite's local observation only.
src/orbitzoo/ reusable OrbitZoo code
src/orbitzoo/thesis/ thesis-specific source code
configs/ versioned JSON experiment configurations
tests/ automated tests
runs/ generated metrics, metadata, and TensorBoard logs (ignored)
checkpoints/ generated model files (ignored)
Every run must save:
config.json: the exact experiment configuration;environment_info.json: device, platform, Python, and PyTorch details;metrics.csvand TensorBoard data when training is added;- model checkpoints when MAPPO is added.
orbitzoo.thesis.runtime.select_device() selects CUDA when available, then
Apple Metal (MPS), otherwise CPU. The choice is recorded for each run.
The initial configuration is configs/mappo_toy.json.
It uses 16 agents, seven discrete actions, and the calibrated neighborhood size
k = 1 and decision interval of 120 seconds (see
calibration findings).
Its maneuver is 0.5 m/s per action at up to 7 N, as selected by the
maneuver sizing study.
The discrete MAPPO implementation is validated independently of orbital physics. See MAPPO.md for its CTDE design and rollout API. The deterministic toy environment validates end-to-end shared-policy learning before the policy is connected to orbital simulation.
The maneuver contract defines the discrete action IDs, their RSW thrust directions, finite-burn execution, and delta-v accounting.
The collision-avoidance environment connects that contract to OrbitZoo propagation, deterministic conjunction screening, rewards, episode termination, and diagnostics.
The environment is verified end to end on Orekit: action directions, burn displacement, fuel and delta-v accounting, determinism, termination, and the full reward table (see verification). Bodies are now created at the configured initial epoch, so real calendar epochs propagate correctly.
The training loop (oz train) trains the shared policy on the
environment, with time-limit bootstrapping, per-update metrics, TensorBoard logs,
checkpoints, and exact resume. Its scenario source is pluggable; only the
development fixture exists until the training scenarios are built.
Evaluation (oz evaluate) compares trained checkpoints with
no-op and rule-based Clohessy–Wiltshire baselines on identical held-out episodes.
The scalability evaluation (oz scale) runs the frozen actor
against the full TLE catalog, sweeping catalog size and agent count, and reports
conjunctions, maneuver-induced secondary conjunctions, delta-v, and per-stage cost.
Maneuver sizing (oz size-maneuvers) derives the per-action
delta-v and minimum thrust from the reference conjunctions.
Visualization (oz watch) flies a checkpoint through one
held-out episode in the headed OrbitZoo viewer, optionally recording an mp4.
Training scenarios are generated per episode from real
orbits and real close-call geometry, in a 16 → 64 → 150 agent curriculum.
Training trials record the short runs that chose the reward
weights and exploration settings used by the curriculum, and
training results record the adopted curriculum run, its seed
repeat, and its evaluation against the no-op and rule-based baselines.
Multi-threat benchmark measures where the actor beats the
rule and by how much; scalability results measure how often
that situation occurs, the per-agent cost to 10,000 agents, and the fuel penalty;
fuel trade-off retrains the curriculum at four delta-v penalties
and shows the margin depends on the burns a cheaper policy stops making;
collinear baseline measures the traditional along-track
heuristic that the methodology names.
Ablations remove locality in both directions — selecting the
neighbour by proximity, and removing selection entirely — and show the policy fails
either way, against a control model trained under the ablations' own protocol.
Convergence reports the fifth evaluation measure for every
variant, and shows it separates a variant that never trained from one that trained
cleanly and learned to do nothing.
The actor now receives fixed-width, threat-ranked local observations containing
k relative-neighbour blocks with explicit padding masks. The critic receives
a separate, deterministic full-system training state. This makes the deployed
actor input independent of constellation population size.
The versioned calibration configuration defines the
catalog inputs, propagation window, candidate values, deterministic seeds, and
passing thresholds that will be used to select k and the decision interval
before MAPPO training. Its strict catalog loader now validates two-line and
three-line TLE input, records NORAD IDs and UTC epochs, applies the configured
freshness cutoff, and joins optional object metadata. Validated calibration data
models provide SI-unit Cartesian frames, deterministic agent selections,
ID-based conjunctions and neighbor rankings, combination metrics, and final
recommendations without coupling reference results to the runtime safety model.
The propagation layer runs a
60-second pass whose spatial index contains the full catalog but is queried only
from selected-agent positions. It excludes self and catalog-only pairs,
deduplicates agent-agent pairs, merges conservative candidate encounter windows,
and produces 10-second SGP4 TEME states only for the involved pairs. This avoids
both all-pairs Python screening and a full day of 10-second catalog states while
preserving SI units and deterministic ordering.
Agent population selection now operates on the post-altitude-filter payload
pool. Explicit PCG64 permutations make the 16-, 64-, and 256-agent populations
nested and reproducible for every calibration and validation seed, with the
selected NORAD IDs stored in a validated versioned manifest.
The largest nested populations across all seeds can be screened once, and each
smaller population reuses the resulting windows through an ID-based filter.
Fine pair trajectories are converted into versioned reference conjunction
truth using bounded TCA refinement between 10-second samples. Coarse false
positives are removed at the configured safe separation, collision status is
derived from catalog radii, overlapping detections are deduplicated, and nested
agent selections reuse the shared truth through deterministic ID filtering.
Candidate decision timelines now share one set of propagation epochs. At each
epoch, vectorized selected-agent-to-catalog calculations produce deterministic
threat rankings without evaluating irrelevant catalog-only pairs. The maximum
configured neighborhood is streamed once and filtered for nested populations
and smaller k values.
One joint evaluator now consumes those ranking frames and updates every
applicable neighborhood prefix and decision schedule simultaneously. It matches
ranked pairs to future reference events within the screening horizon, records
first visibility, and derives the number of actionable decisions remaining
before TCA without rerunning propagation for any (k, delta t) combination.
Detailed detections are aggregated into the complete per-sample combination
grid, including explicit zero-event cases. Calibration and validation counts
are pooled separately before recall and timely-detection rates are calculated,
and versioned passing thresholds produce an auditable result for every
candidate pair without averaging percentages across unequal samples.
Final selection now considers calibration evidence only, preferring the
smallest passing neighborhood and then the largest passing decision interval.
The exact selected pair is audited on held-out validation data without fallback
selection, producing separate calibration, validation, and final acceptance
statuses.
The oz calibrate command now orchestrates the complete offline workflow once,
publishes deterministic versioned evidence in an atomic run directory, reports
stage progress and runtime, refuses accidental overwrite, and distinguishes a
calibration no-pass from held-out validation rejection through its exit status.