Multi-agent reinforcement learning for autonomous collision avoidance in low Earth orbit, trained and evaluated on real two-line element catalogue data.
One shared policy is flown by every satellite. Each one sees only its single most threatening neighbour and decides every 120 seconds whether to burn. No satellite communicates with another and no priority rule breaks ties, so any coordination between them is emergent rather than designed. Because the observation is a fixed width regardless of how many satellites exist, the cost of running the policy per satellite does not grow with the constellation.
| Path | Contents |
|---|---|
src/orbitzoo/thesis/environments/ |
the collision-avoidance environment, observations and rewards |
src/orbitzoo/thesis/scenarios/ |
the scenario generator, built from real orbits and real close calls |
src/orbitzoo/thesis/training/ |
the MAPPO training loop and curriculum |
src/orbitzoo/thesis/evaluation/ |
baselines, metrics and the frozen benchmark |
src/orbitzoo/thesis/scalability/ |
the catalogue-scale evaluator |
src/orbitzoo/cli/ |
the oz command line that drives all of the above |
docs/ |
design, methods and results |
configs/ |
every experiment configuration |
python3.11 -m venv .venv && .venv/bin/pip install -e .
.venv/bin/python -m pytest -q --ignore=tests/test_interface_connections.py
.venv/bin/oz --helpoz provides train, evaluate, scale, calibrate and size-maneuvers.
docs/README.md is the index. Start with THESIS_IMPLEMENTATION.md for the architecture, repository layout and current status.
| Section | Contents |
|---|---|
| design/ | how the environment, policy and scenarios work |
| methods/ | how to run training, evaluation, calibration and the scale sweeps |
| results/ | what each study measured, with a summary |
The orbital dynamics, the Orekit and SGP4 propagation backends, the 3D interface, and the base MARL scaffolding come from OrbitZoo. ATTRIBUTION.md lists every upstream file, verbatim or modified.