Code for a multi-agent particle environment used in the paper "Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments"
-
Updated
Apr 24, 2020 - Python
Code for a multi-agent particle environment used in the paper "Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments"
Various projects related to contingent planning under partial observability
SymDer: Symbolic Derivative Approach to Discovering Sparse Interpretable Dynamics from Partial Observations
Solving pursuit-evasion problems on graphs using Reinfocement Learning and GNNs
Online Replanning in Belief Space for Partially Observable Task and Motion Problems
Official PyTorch implementation of POEM (Partial Observation Experts Modelling) as introduced in the paper Contrastive Meta-Learning for Partially Observable Few-Shot Learning
Minimal control-theoretic toy models for studying stability and failure modes under partial observability.
Blog post on safety constraint learning
Python library for studying elimination-based learning in deterministic MDPs under state aliasing
Hybrid planning agent for decision-making under partial observability (PG-WMA)
Partially Observable Monte Carlo Planning algorithm (POMCP)
UWM-JEPA: a JEPA world model with a density-matrix latent and learned unitary predictor for imagining hidden continuations under partial observability.
The AIS planner source code + benchmarks
Bachelor's thesis: reproducible multi-seed MiniGrid benchmark of PPO, A2C, DQN and RecurrentPPO under partial observability — memory, intrinsic motivation (RND, ICM, RIDE, NovelD) and curriculum learning.
Operator-geometric framework for metabolic identifiability and metabolite panel design under partial observability using Human-GEM, proteomics-informed reaction weighting, and spectral geometry.
Priced information-gathering and abstention evaluation in a synthetic partially observed city.
MATLAB code for the paper "Assimilative Causal Inference".
意思を持つテトロミノたち — 各落下ピースが盤面スナップショット画像だけを見て収まり先を決めるマルチエージェント・テトリス PoC。ローカル VLA 対応、自律/A2A/独裁の協調モード比較、マシン性能ベンチマーク付き。
Defender-side cyber incident simulator. An abstract state machine, not real infrastructure, and reinforcement-learning policies, not LLM agents. Host status is hidden behind noisy alerts, the attacker improvises, and a web console lets you play the same scenario by hand and compare your score to a trained agent's.
Comparing a full-map A* oracle with partially observable knowledge-based and genetic-policy agents in Wumpus World.
To associate your repository with the partial-observability topic, visit your repo's landing page and select "manage topics."