A Survey on End-to-End Autonomous Driving Training from the Perspectives of Data, Strategy, and Platform
Welcome to this curated collection of papers on End-to-End Autonomous Driving (E2E-AD), aimed at researchers, engineers, and enthusiasts in the field of autonomous driving systems. This repository provides a comprehensive selection of papers, focusing primarily on the training methods and ecosystems that drive the development of intelligent autonomous vehicles.
In particular, we focus on the Data-Strategy-Platform framework for E2E-AD systems, offering insights into:
- Data Layer: Addressing data collection, coverage, and governance practices.
- Strategy Layer: Exploring imitation learning, reinforcement learning, and generative approaches.
- Platform Layer: Understanding scalable training infrastructures and cloud-edge collaborations.
Each paper in this repository has been selected for its relevance and contribution to the field, and we hope it serves as a valuable resource for anyone working in or learning about autonomous driving technology.
| 🎉 2026.05.20 | Our survey was accepted by IEEE Transactions on Intelligent Transportation Systems (TITS). |
| 🔄 2026.05.22 | Repository updated with the latest E2E-AD training ecosystem papers and resources. |
We warmly welcome pull requests and suggestions for adding new papers, benchmarks, datasets, and useful resources.
If you find this repository helpful, please consider citing our survey, starring this repository ⭐, and sharing it with the community.
- Survey paper
- Research Papers
- Datasets & Benchmarks
- Other Awesome Lists
- Citation
| Title | Abstract | Year | Project |
|---|---|---|---|
| Closed Loop Dynamic Driving Data Mixture for Real-Synthetic Co-Training | DetailsProposes AutoScale, a closed-loop data engine that optimizes real-synthetic driving data mixtures with scene representations, cluster reweighting, retrieval, training, and evaluation feedback. |
arXiv 2026 | |
| 4DLidarOpen: An Open 4D FMCW Lidar Dataset for Motion-Aware Autonomous Driving | DetailsIntroduces a multi-modal 4D FMCW LiDAR dataset with point-wise velocity, multi-LiDAR and camera streams, 3D boxes, tracks, and benchmarks for detection, BEV flow, forecasting, and planning. |
arXiv 2026 | |
| XWOD: A Real-World Benchmark for Object Detection under Extreme Weather Conditions | DetailsBuilds an extreme-weather object-detection benchmark with real traffic images across rain, snow, fog, flooding, tornado, wildfire, and other adverse conditions. |
arXiv 2026 | |
| ScenePilot-Bench: A Large-Scale Dataset and Benchmark for Evaluation of Vision-Language Models in Autonomous Driving | DetailsProvides a first-person driving VLM benchmark built on ScenePilot-4K, evaluating scene understanding, spatial perception, motion planning, safety reasoning, and regional generalization. |
arXiv 2026 | Dataset |
| VR-Drive: Viewpoint-Robust End-to-End Driving with Feed-Forward 3D Gaussian Splatting | DetailsUses feed-forward 3D Gaussian Splatting to synthesize viewpoint-robust driving observations, improving end-to-end policy robustness under camera pose and viewpoint changes. |
NeurIPS 2025 | Project |
| WOD-E2E: Waymo Open Dataset for End-to-End Driving in Challenging Long-tail Scenarios | DetailsAdapts Waymo Open Dataset into an end-to-end driving benchmark emphasizing challenging long-tail scenes for planning-oriented policy evaluation. |
arXiv 2025 | Project |
| SimScale: Learning to Drive via Real-World Simulation at Scale | DetailsScales real-world simulation for closed-loop end-to-end driving by converting large-scale logs into interactive training environments for policy learning and evaluation. |
CVPR 2026 Oral | Project / Code |
| SynAD: Enhancing Real-World End-to-End Autonomous Driving Models through Synthetic Data | DetailsStudies synthetic-data augmentation for real-world end-to-end driving models, targeting better robustness and generalization under data scarcity and long-tail scenarios. |
ICCV 2025 | |
| CoVLA: Comprehensive Vision-Language-Action Dataset for Autonomous Driving | Details |
WACV 2025 | Project |
| Argoverse 2: Next generation datasets for self-driving perception and forecasting | Details |
arXiv 2023 | Project |
| WOMD-Reasoning: A Large-Scale Dataset for Interaction Reasoning in Driving | Details |
ICML 2025 | Code / Project |
| nuscenes: A multimodal dataset for autonomous driving | Details |
CVPR 2020 | Project |
| One million scenes for autonomous driving: Once dataset | Details |
arXiv 2021 | Project |
| Scalability in perception for autonomous driving: Waymo open dataset | DetailsIntroduces the Waymo Open Dataset for scalable autonomous-driving perception research, including synchronized LiDAR, camera, labels, and benchmarks for detection and tracking. |
CVPR 2020 | Project |
| Zenseact Open Dataset: A Large-Scale and Diverse Multimodal Dataset for Autonomous Driving | DetailsIntroduces a large-scale multimodal autonomous-driving dataset with diverse European driving scenes and annotations for perception, prediction, and planning research. |
ICCV 2023 | Project |
| Scaling out-of-distribution detection for real-world settings | Details |
PMLR 2022 | Code |
| SHIFT: a synthetic driving dataset for continuous multi-task domain adaptation | DetailsIntroduces a synthetic driving dataset for continuous domain adaptation across weather, time, and scene changes, supporting multiple perception tasks. |
CVPR 2022 | Project |
| V2x-vit: Vehicle-to-everything cooperative perception with vision transformer | Details |
ECCV 2022 | Code |
| Deepaccident: A motion and accident prediction benchmark for v2x autonomous driving | DetailsProvides a V2X benchmark for accident and motion prediction, focusing on safety-critical cooperative driving scenarios. |
AAAI 2024 | |
| Bdd100k: A diverse driving dataset for heterogeneous multitask learning | Details |
CVPR 2020 | Project |
| The apolloscape dataset for autonomous driving | Details |
CVPR 2018 | Project |
| Tumtraf v2x cooperative perception dataset | Details |
CVPR 2024 | Project |
| Tumtraf intersection dataset: All you need for urban 3d camera-lidar roadside perception | DetailsPresents an urban roadside camera-LiDAR dataset for 3D perception at intersections, supporting infrastructure-side cooperative perception research. |
ITSC 2023 | Code / Project |
| V2v4real: A real-world large-scale dataset for vehicle-to-vehicle cooperative perception | DetailsIntroduces a real-world V2V cooperative perception dataset with synchronized multi-vehicle sensing for studying collaboration under realistic communication and viewpoint constraints. |
CVPR 2023 | Project |
| Rope3d: The roadside perception dataset for autonomous driving and monocular 3d object detection task | Details |
CVPR 2022 | Project |
| Cooperative perception for 3D object detection in driving scenarios using infrastructure sensors | Details |
TITS 2020 | |
| Lumpi: The leibniz university multi-perspective intersection dataset | Details |
IV 2022 | Project |
| An automated driving systems data acquisition and analytics platform | Details |
TRC 2023 | |
| S-nerf++: Autonomous driving simulation via neural reconstruction and generation | Details |
TPAMI 2025 | Code |
| ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration | DetailsBuilds driving world models for scene reconstruction through online restoration, improving reconstruction quality and temporal consistency for simulation and data generation. |
CVPR 2025 | |
| Scene reconstruction techniques for autonomous driving: a review of 3D Gaussian splatting | Details |
AIR | |
| Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving | Details |
CVPR 2024 | Code / Project |
| Scenecontrol: Diffusion for controllable traffic scene generation | Details |
ICRA 2024 | |
| Diffscene: Diffusion-based safety-critical scenario generation for autonomous vehicles | Details |
AAAI 2025 | |
| Simulation-based reinforcement learning for real-world autonomous driving | Details |
ICRA 2020 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| CLOVER: Closed-Loop Value Estimation & Ranking for End-to-End Autonomous Driving Planning | DetailsRanks candidate plans with closed-loop value estimation, improving planning selection by considering downstream interactive outcomes rather than only open-loop trajectory error. |
arXiv 2026 | Code |
| Beyond Imitation: Learning Safe End-to-End Autonomous Driving from Hard Negatives | DetailsImproves imitation-based driving by mining hard negative behaviors and training the policy to avoid unsafe actions in challenging scenarios. |
arXiv 2026 | |
| Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling | DetailsModels ego-centric scene causality jointly with driving policy learning, aiming to separate causal factors from spurious correlations for safer end-to-end planning. |
arXiv 2026 | |
| Temporal Sampling Frequency Matters: A Capacity-Aware Study of End-to-End Driving Trajectory Prediction | DetailsAnalyzes how temporal sampling frequency interacts with model capacity in trajectory prediction, offering guidance for data construction and policy training. |
arXiv 2026 | |
| Future-Aware End-to-End Driving: Bidirectional Modeling of Trajectory Planning and Scene Evolution | DetailsJointly models future scene evolution and ego trajectory planning, using bidirectional interactions between prediction and planning to improve driving decisions. |
NeurIPS 2025 | Code |
| Perception in Plan: Coupled Perception and Planning for End-to-End Autonomous Driving | DetailsCouples perception and planning in a single end-to-end framework so planning supervision can shape perception features toward driving-relevant scene understanding. |
AAAI 2026 | |
| Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning | DetailsUses historical prediction and future planning jointly, bridging temporal context and planning targets for stronger end-to-end driving performance. |
CVPR 2025 | Code |
| Don't Shake the Wheel: Momentum-Aware Planning in End-to-End Autonomous Driving | DetailsIntroduces momentum-aware planning to reduce unstable steering and improve trajectory smoothness while preserving planning accuracy. |
CVPR 2025 | Code |
| Planning-oriented autonomous driving | Details |
CVPR 2023 | Code |
| Transfuser: Imitation with transformer-based sensor fusion for autonomous driving | Details |
TPAMI 2022 | Code |
| Safety-enhanced autonomous driving using interpretable sensor fusion transformer | Details |
CoRL 2022 | Code |
| Reasonnet: End-to-end driving with temporal and global reasoning | Details |
CVPR 2023 | Code |
| Ppad: Iterative interactions of prediction and planning for end-to-end autonomous driving | Details |
ECCV 2024 | Code |
| Multi-modal fusion transformer for end-to-end autonomous driving | Details |
CVPR 2021 | Code |
| Think twice before driving: Towards scalable decoders for end-to-end autonomous driving | Details |
CVPR 2023 | Code |
| Learning from all vehicles | Details |
CVPR 2022 | Code |
| Neat: Neural attention fields for end-to-end autonomous driving | Details |
ICCV 2021 | Code |
| Learning to steer by mimicking features from heterogeneous auxiliary networks | Details |
AAAI 2019 | Code |
| Driving on Registers | DetailsInvestigates register-like internal representations for end-to-end driving, improving how models store and use scene context for planning. |
arXiv 2026 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving | DetailsCombines cognitive reasoning distillation with physical action learning, using reinforcement learning to align high-level foresight with low-level control. |
arXiv 2026 | |
| MAPLE: Latent Multi-Agent Play for End-to-End Autonomous Driving | DetailsIntroduces latent multi-agent play to train end-to-end driving policies against interactive agents, improving robustness in socially complex traffic. |
arXiv 2026 | |
| MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning | DetailsUses transformer-based 3D affordance representations with reinforcement learning to improve robust urban driving under multi-modal sensory inputs. |
arXiv 2026 | |
| DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving | DetailsApplies safety-oriented Direct Preference Optimization to end-to-end driving policy learning, aligning planner outputs with safe trajectory preferences. |
NeurIPS 2025 | |
| RAD: Training an End-to-End Driving Policy via Large-Scale 3DGS-based Reinforcement Learning | DetailsTrains end-to-end driving policies in large-scale 3D Gaussian Splatting environments, using reinforcement learning to bridge realistic simulation and closed-loop driving. |
NeurIPS 2025 | Project / Code |
| Sim2Real-AD: A Modular Sim-to-Real Framework for Deploying VLM-Guided Reinforcement Learning in Real-World Autonomous Driving | DetailsBuilds a modular sim-to-real pipeline for VLM-guided reinforcement learning, targeting deployable autonomous-driving policies beyond simulation. |
arXiv 2026 | |
| DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving | DetailsAdapts latent world-model reinforcement learning to autonomous driving, reducing environment interaction cost through imagined rollouts. |
arXiv 2026 | |
| DriveVLM-RL: Neuroscience-Inspired Reinforcement Learning with Vision-Language Models for Safe and Deployable Autonomous Driving | DetailsUses VLM-derived semantic rewards during offline training through a dual-pathway design, then removes VLM inference at deployment for real-time driving. |
arXiv 2026 | Project |
| CorrectionPlanner: Self-Correction Planner with Reinforcement Learning in Autonomous Driving | DetailsModels planning as a propose-evaluate-correct loop over motion tokens, using a collision critic and model-based RL to recover from unsafe candidate actions. |
arXiv 2026 | |
| TADPO: Reinforcement Learning Goes Off-road | DetailsOff-road autonomous driving presents sparse-reward and high-risk exploration challenges; TADPO proposes an RL optimization scheme tailored for off-road robustness and safety trade-offs. |
arXiv 2026 | |
| ThinkDrive: Chain-of-Thought Guided Progressive Reinforcement Learning Fine-Tuning for Autonomous Driving | DetailsCombines chain-of-thought supervision with progressive reinforcement fine-tuning to align driving reasoning, intent, and trajectory planning. |
arXiv 2026 | |
| Effective Learning Mechanism Based on Reward-Oriented Hierarchies for Sim-to-Real Adaption in Autonomous Driving Systems | Details |
TITS 2025 | |
| Safe-state enhancement method for autonomous driving via direct hierarchical reinforcement learning | Details |
TITS 2023 | |
| Imitation is not enough: Robustifying imitation with reinforcement learning for challenging driving scenarios | Details |
IROS 2023 | |
| Safe reinforcement learning for autonomous vehicle using monte carlo tree search | Details |
TITS 2021 | |
| Uncertainty-aware model-based reinforcement learning: Methodology and application in autonomous driving | Details |
TIV 2022 | |
| Safety-aware causal representation for trustworthy offline reinforcement learning in autonomous driving | Details |
RAL 2024 | |
| Improving generalization of transfer learning across domains using spatio-temporal features in autonomous driving | Details |
arXiv 2021 | |
| Towards socially responsive autonomous vehicles: A reinforcement learning framework with driving priors and coordination awareness | Details |
TIV 2023 | |
| Uncertainty-aware model-based offline reinforcement learning for automated driving | Details |
RAL 2023 | |
| Towards safe and robust autonomous vehicle platooning: A self-organizing cooperative control framework | Details |
arXiv 2024 | Project |
| Reinforced Refinement with Self-Aware Expansion for End-to-End Autonomous Driving | Details |
TPAMI 2026 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| DriveSafer: End-to-End Autonomous Driving with Safety Guidance | DetailsAdds explicit safety guidance to end-to-end driving, steering trajectory generation toward safer behavior in complex traffic scenes. |
arXiv 2026 | |
| MISTY: High-Throughput Motion Planning via Mixer-based Single-step Drifting | DetailsUses mixer-based single-step trajectory generation to accelerate motion planning while preserving multi-modal planning quality. |
arXiv 2026 | |
| FeaXDrive: Feasibility-aware Trajectory-Centric Diffusion Planning for End-to-End Autonomous Driving | DetailsIntroduces feasibility-aware diffusion planning centered on trajectory generation, improving the physical and driving-rule validity of predicted plans. |
arXiv 2026 | |
| RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework | DetailsScales reinforcement learning with a generator-discriminator framework, targeting stronger policy optimization for autonomous driving. |
arXiv 2026 | |
| HAD: Combining Hierarchical Diffusion with Metric-Decoupled RL for End-to-End Driving | DetailsCombines hierarchical diffusion planning with metric-decoupled reinforcement learning to improve safety, comfort, and task progress separately. |
arXiv 2026 | |
| Temporally Decoupled Diffusion Planning for Autonomous Driving | DetailsDecouples temporal components in diffusion-based planning, enabling more flexible long-horizon trajectory generation for autonomous driving. |
arXiv 2026 | |
| DiffRefiner: Coarse to Fine Trajectory Planning via Diffusion Refinement with Semantic Interaction for End to End Autonomous Driving | DetailsRefines coarse trajectory proposals through diffusion with semantic interaction modeling, improving fine-grained planning in end-to-end driving. |
AAAI 2026 | |
| Driving with Advice: Large Model as Motion Advisor for Joint Planning | DetailsUses a large model as a motion advisor to guide joint planning, injecting high-level semantic advice into trajectory generation. |
AAAI 2026 | |
| DiffE2E: Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy | DetailsCombines action diffusion with supervised policy learning, balancing generative trajectory diversity with stable end-to-end driving behavior. |
NeurIPS 2025 | Project |
| WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving | DetailsGenerates future waypoints through parallel coarse-to-fine discrete flow matching and further optimizes closed-loop behavior with simulator-guided rewards. |
CVPR 2026 | Code |
| Dichotomous Diffusion Policy Optimization | DetailsProposes DIPOLE, a stable RL method for diffusion policies that decomposes policy improvement into reward-maximizing and reward-minimizing branches, enabling controllable inference and VLA driving experiments on NAVSIM. |
ICLR 2026 | Project / Code |
| Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving | Details |
CVPR 2025 | Code |
| Diffvla: Vision-language guided diffusion planning for autonomous driving | Details |
arXiv 2025 | |
| DiffAD: A Unified Diffusion Modeling Approach for Autonomous Driving | Details |
arXiv 2025 | |
| A Knowledge-Driven Diffusion Policy for End-to-End Autonomous Driving Based on Expert Routing | Details |
arXiv 2025 | Code / Project |
| Diffusion-based planning for autonomous driving with flexible guidance | Details |
arXiv 2025 | |
| Diffusion-ES: Gradient-free planning with diffusion for autonomous and instruction-guided driving | Details |
CVPR 2024 | |
| Uncertainty-Based Alternative Diffusion Policy for Safe Autonomous Driving | Details |
TITS 2025 | |
| Recogdrive: A reinforced cognitive framework for end-to-end autonomous driving | Details |
arXiv 2025 | |
| FlowDrive: Energy Flow Field for End-to-End Autonomous Driving | Details |
arXiv 2025 | Project |
| Diffusion-based planning for autonomous driving with flexible guidance | Details |
arXiv 2025 |
| Title | Abstract | Year | Project |
|---|---|---|---|
| Lost in Fog: Sensor Perturbations Expose Reasoning Fragility in Driving VLAs | DetailsEvaluates driving VLA robustness under sensor perturbations such as fog, showing how perception degradation can expose fragile reasoning and planning behavior. |
arXiv 2026 | |
| SafeAlign-VLA: A Negative-Enhanced Safe Alignment Framework for Risk-Aware Autonomous Driving | DetailsAligns VLA driving behavior with safety preferences by emphasizing negative examples and risk-aware supervision during training. |
arXiv 2026 | |
| CLAP: Contrastive Latent-space Prompt Optimization for End-to-end Autonomous Driving | DetailsOptimizes prompts in latent space with contrastive objectives, improving end-to-end driving policy adaptation without heavy model retraining. |
arXiv 2026 | |
| MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving | DetailsIntroduces a unified streaming VLA architecture for autonomous driving, integrating perception, language understanding, and action generation in an online setting. |
arXiv 2026 | |
| OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation | DetailsPerforms one-step latent reasoning for planning while producing vision-language explanations, reducing multi-step reasoning overhead in VLA driving. |
arXiv 2026 | Project |
| OneDrive: Unified Multi-Paradigm Driving with Vision-Language-Action Models | DetailsUnifies multiple driving paradigms in a VLA framework, connecting perception, reasoning, and action generation across different supervision modes. |
arXiv 2026 | Code |
| SpanVLA: Efficient Action Bridging and Learning from Negative-Recovery Samples for Vision-Language-Action Model | DetailsImproves VLA driving through efficient action bridging and negative-recovery samples, helping policies learn to recover from unsafe or suboptimal actions. |
arXiv 2026 | |
| UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving | DetailsUnifies scene understanding, perception, and action planning in a VLA architecture for end-to-end autonomous driving. |
arXiv 2026 | Code |
| ICR-Drive: Instruction Counterfactual Robustness for End-to-End Language-Driven Autonomous Driving | DetailsEvaluates and improves instruction-conditioned driving robustness with counterfactual language and scene perturbations. |
arXiv 2026 | |
| Vega: Learning to Drive with Natural Language Instructions | DetailsTrains driving policies conditioned on natural language instructions, connecting high-level command following with trajectory-level planning. |
arXiv 2026 | |
| NoRD: A Data-Efficient Vision-Language-Action Model that Drives without Reasoning | DetailsShows that a data-efficient VLA policy can drive effectively without explicit chain-of-thought reasoning, reducing inference cost for deployment. |
CVPR 2026 | Project |
| VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving | DetailsAdds cross-view geometric grounding to VLMs, improving spatial understanding and planning reliability in multi-view autonomous driving. |
CVPR 2026 | Project / Code |
| FastDriveVLA: Efficient End-to-End Driving via Plug-and-Play Reconstruction-based Token Pruning | DetailsProposes ReconPruner, a plug-and-play MAE-style visual token pruner that preserves foreground driving information; introduces nuScenes-FG with 241K image-mask pairs and improves nuScenes open-loop planning across pruning ratios. |
AAAI 2026 | |
| VLM-AutoDrive: Post-Training Vision-Language Models for Safety-Critical Autonomous Driving Events | DetailsAdapts pretrained VLMs to safety-critical dashcam events with metadata captions, LLM descriptions, VQA pairs, and CoT supervision, improving collision and near-collision detection with interpretable reasoning traces. |
arXiv 2026 | |
| Senna-2: Aligning VLM and End-to-End Driving Policy for Consistent Decision Making and Planning | DetailsAligns VLM reasoning with end-to-end policy learning to reduce decision/planning inconsistency and improve reliability in complex driving scenes. |
arXiv 2026 | |
| SAMoE-VLA: A Scene Adaptive Mixture-of-Experts Vision-Language-Action Model for Autonomous Driving | DetailsProposes a scene-adaptive MoE VLA architecture that routes computation by scene context for stronger robustness and efficiency in end-to-end driving. |
arXiv 2026 | |
| LaST-VLA: Thinking in Latent Spatio-Temporal Space for Vision-Language-Action in Autonomous Driving | DetailsIntroduces latent spatio-temporal reasoning for driving VLA models to improve long-horizon planning quality and robustness under complex scene dynamics. |
arXiv 2026 | |
| Devil is in Narrow Policy: Unleashing Exploration in Driving VLA Models | DetailsAnalyzes narrow-policy collapse in driving VLAs and proposes exploration-centric training to improve robustness in long-tail scenarios. |
arXiv 2026 | |
| Modular Autonomy with Conversational Interaction: An LLM-driven Framework for Decision Making in Autonomous Driving | DetailsConnects an LLM-based conversational interface to modular autonomy software, translating passenger commands into validated driving-system actions. |
IV 2026 | |
| SGDrive: Scene-to-Goal Hierarchical World Cognition for Autonomous Driving | DetailsStructures VLM driving cognition into scene, agent, and goal levels, producing compact representations for trajectory planning. |
arXiv 2026 | |
| LatentVLA: Efficient Vision-Language Models for Autonomous Driving via Latent Action Prediction | DetailsUses self-supervised latent action prediction and distillation to build efficient VLA driving policies with reduced language-annotation dependence. |
arXiv 2026 | |
| A Vision-Language-Action Model with Visual Prompt for OFF-Road Autonomous Driving | DetailsIntroduces OFF-EMMA, an off-road VLA driving model that uses visual prompts and self-consistent reasoning to improve trajectory planning on rough terrain. |
arXiv 2026 | |
| FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder | Details |
WACV 2026 Workshop | |
| Counterfactual VLA: Self-Reflective Vision-Language-Action Model with Adaptive Reasoning | DetailsAdds counterfactual self-reflection to VLA driving, allowing the model to revise planned actions before trajectory generation in challenging scenes. |
arXiv 2025 | |
| KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System | DetailsCombines driving knowledge retrieval with a value model to guide interpretable, value-aligned trajectory assessment and planning. |
arXiv 2025 | |
| FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving | Details |
arXiv 2025 | Code |
| ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving | Details |
arXiv 2025 | Project / Code |
| DriveMoE: Mixture-of-Experts for Vision-Language-Action Model in End-to-End Autonomous Driving | Details |
arXiv 2025 | Project |
| DSDrive: Distilling Large Language Model for Lightweight End-to-End Autonomous Driving with Unified Reasoning and Planning | Details |
arXiv 2025 | |
| OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model | DetailsDevelops a large vision-language-action model for end-to-end autonomous driving, connecting multi-view perception, language reasoning, and trajectory output. |
AAAI 2026 | Code |
| Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models | DetailsReleases open weights and open data for driving VLA models, supporting reproducible research on language-conditioned autonomous driving. |
arXiv 2025 | Code / Project |
| Towards Human-Centric Autonomous Driving: A Fast-Slow Architecture Integrating Large Language Model Guidance with Reinforcement Learning | DetailsCombines fast low-level control with slower LLM-guided reasoning and reinforcement learning to improve human-centric driving decisions. |
ITSC 2025 | Project |
| DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving | DetailsDevelops a vision-language reasoning model for cross-task autonomous driving, connecting perception, reasoning, and planning tasks. |
arXiv 2025 | Project |
| DiffVLA: Vision-Language Guided Diffusion Planning for Autonomous Driving | DetailsUses vision-language guidance to condition diffusion-based trajectory planning, combining semantic reasoning with generative action prediction. |
arXiv 2025 | |
| AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning | DetailsCombines adaptive reasoning with reinforcement fine-tuning in a VLA driving model to improve planning under complex scene context. |
NeurIPS 2025 | Code / Project |
| Extending Large Vision-Language Model for Diverse Interactive Tasks in Autonomous Driving | DetailsExtends large vision-language models to interactive autonomous-driving tasks such as question answering, decision support, and scene-grounded reasoning. |
arXiv 2025 | |
| X-Driver: Explainable Autonomous Driving with Vision-Language Models | DetailsUses vision-language models to produce explainable driving decisions, connecting visual evidence with action-level reasoning. |
arXiv 2025 | |
| AlphaDrive: Unleashing the Power of VLMs in Autonomous Driving via Reinforcement Learning and Reasoning | DetailsCombines VLM reasoning with reinforcement learning to improve autonomous-driving decisions under complex scene context. |
arXiv 2025 | Code |
| Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning | DetailsBuilds a generalized MLLM framework for translating scene understanding into driving decisions and trajectories. |
arXiv 2025 | |
| VLM-MPC: Vision Language Foundation Model (VLM)-Guided Model Predictive Controller (MPC) for Autonomous Driving | DetailsGuides model predictive control with VLM-derived scene understanding and driving intent for autonomous driving. |
ICML 2025 | |
| VLM-E2E: Enhancing End-to-End Autonomous Driving with Multi-modal Driver Attention Fusion | DetailsFuses driver attention with multi-modal VLM features to enhance end-to-end autonomous-driving prediction and planning. |
arXiv 2025 | |
| VLM-Assisted Continual learning for Visual Question Answering in Self-Driving | DetailsUses VLM assistance for continual learning in self-driving visual question answering, reducing forgetting across driving domains. |
arXiv 2025 | |
| WiseAD: Knowledge Augmented End-to-End Autonomous Driving with Vision-Language Model | Details |
arXiv 2024 | Code / Project |
| CALMM-Drive: Confidence-Aware Autonomous Driving with Large Multimodal Model | DetailsUses confidence-aware multimodal reasoning to improve reliability in end-to-end autonomous-driving decisions. |
arXiv 2024 | |
| OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving | Details |
WACV 2025 | Code |
| VLM-AD: End-to-End Autonomous Driving through Vision-Language Model Supervision | Details |
arXiv 2024 | |
| TS-VLM: Text-Guided SoftSort Pooling for Vision-Language Models in Multi-View Driving Reasoning | DetailsIntroduces text-guided SoftSort pooling to improve multi-view driving reasoning in VLMs. |
arXiv 2025 | |
| LightEMMA: Lightweight End-to-End Multimodal Model for Autonomous Driving | DetailsBuilds a lightweight multimodal end-to-end driving model for efficient planning and reasoning. |
arXiv 2025 | Code |
| DriveGPT4: Interpretable End-to-end Autonomous Driving via Large Language Model | Details |
RAL 2024 | Project |
| ADAPT: Action-aware Driving Caption Transformer | Details |
ICRA 2023 | Code |
| Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving | Details |
NeurIPS 2024 | Code |
| DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models | Details |
arXiv 2024 | Project |
| LingoQA: Visual Question Answering for Autonomous Driving | Details |
ECCV 2024 | Code |
| Training-Free Open-Ended Object Detection and Segmentation via Attention as Prompts | Details |
NeurIPS 2024 | |
| ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation | DetailsGenerates driving actions from vision-language instructions in a holistic end-to-end autonomous-driving framework. |
arXiv 2025 | Code |
| Generative Planning with 3D-vision Language Pre-training for End-to-End Autonomous Driving | DetailsUses 3D vision-language pre-training to support generative trajectory planning for end-to-end autonomous driving. |
arXiv 2025 | |
| FutureSightDrive: Visualizing Trajectory Planning with Spatio-Temporal CoT for Autonomous Driving | DetailsUses spatio-temporal chain-of-thought visualization to make trajectory planning more interpretable and reasoning-aware. |
arXiv 2025 | Code |
| Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning | DetailsUses reinforcement learning to bridge VLM reasoning and trajectory planning for autonomous driving. |
arXiv 2025 | |
| ReCogDrive: A Reinforced Cognitive Framework for End-to-End Autonomous Driving | DetailsIntroduces a reinforced cognitive framework that aligns perception, reasoning, and planning in end-to-end driving. |
arXiv 2025 | Project / Code |
| ColaVLA: Leveraging Cognitive Latent Reasoning for Hierarchical Parallel Trajectory Planning in Autonomous Driving | DetailsUses cognitive latent reasoning for hierarchical parallel trajectory planning in driving VLA models. |
arXiv 2025 | Project |
| Large Multimodal Models for Embodied Intelligent Driving: The Next Frontier in Self-Driving? | DetailsDiscusses embodied intelligent driving with large multimodal models, combining semantic understanding with policy optimization for continuous decision learning. |
arXiv 2026 |
🔗Refer to Link
| Title | Abstract | Year | Project |
|---|---|---|---|
| HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models | DetailsUses trajectory-guided world modeling to connect heterogeneous sensor inputs with end-to-end planning for autonomous driving. |
arXiv 2026 | |
| Xiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving | DetailsIntegrates scene reconstruction and future generation in a joint driving world model, supporting both representation learning and planning-oriented simulation. |
arXiv 2026 | |
| EponaV2: Driving World Model with Comprehensive Future Reasoning | DetailsExtends driving world modeling with comprehensive future reasoning, improving long-horizon scene prediction and planning awareness. |
arXiv 2026 | |
| The DAWN of World-Action Interactive Models | DetailsStudies world-action interaction models that jointly learn how actions affect future scene evolution, bridging prediction and control. |
arXiv 2026 | |
| DriveDreamer-Policy: A Geometry-Grounded World-Action Model for Unified Generation and Planning | DetailsUnifies video generation and driving planning with a geometry-grounded world-action model for action-conditioned future simulation. |
arXiv 2026 | |
| Latent-WAM: Latent World Action Modeling for End-to-End Autonomous Driving | DetailsTrains a latent world-action model that predicts action-conditioned future dynamics for end-to-end autonomous driving. |
arXiv 2026 | |
| Bridging Scene Generation and Planning: Driving with World Model via Unifying Vision and Motion Representation | DetailsUnifies visual scene generation and motion planning representations so generated futures can directly support driving decisions. |
arXiv 2026 | |
| Risk-Aware World Model Predictive Control for Generalizable End-to-End Autonomous Driving | DetailsCombines world-model prediction with risk-aware MPC to improve generalization and safety in end-to-end driving. |
arXiv 2026 | |
| ResWorld: Temporal Residual World Model for End-to-End Autonomous Driving | DetailsUses temporal residual world modeling to improve future scene prediction and planning-relevant representation learning. |
ICLR 2026 | Code |
| Learning Vision-Language-Action World Models for Autonomous Driving | DetailsBuilds a VLA world model that learns action-conditioned future prediction and planning-relevant representations from driving data. |
CVPR 2026 Findings | Project |
| LMGenDrive: Bridging Multimodal Understanding and Generative World Modeling for End-to-End Driving | DetailsConnects multimodal scene understanding with generative world modeling, using future generation to support end-to-end driving. |
arXiv 2026 | |
| ExploreVLA: Dense World Modeling and Exploration for End-to-End Autonomous Driving | DetailsCombines dense world modeling with exploration-oriented training to improve VLA driving performance in diverse scenarios. |
arXiv 2026 | |
| Uni-World VLA: Interleaved World Modeling and Planning for Autonomous Driving | DetailsInterleaves world modeling and planning in a unified VLA framework so imagined futures and actions can refine each other. |
arXiv 2026 | |
| OccSim: Multi-kilometer Simulation with Long-horizon Occupancy World Models | DetailsUses long-horizon occupancy world models to simulate multi-kilometer driving scenes for scalable evaluation and training. |
arXiv 2026 | |
| AutoWorld: Scaling Multi-Agent Traffic Simulation with Self-Supervised World Models | DetailsScales multi-agent traffic simulation with self-supervised world models, enabling realistic interaction modeling for driving policy training. |
arXiv 2026 | |
| DLWM: Dual Latent World Models enable Holistic Gaussian-centric Pre-training in Autonomous Driving | DetailsUses dual latent world models for Gaussian-centric pre-training, aligning perception, reconstruction, and future prediction in autonomous driving. |
arXiv 2026 | |
| Infrastructure-Centric World Models: Bridging Temporal Depth and Spatial Breadth for Roadside Perception | DetailsBuilds roadside infrastructure-centric world models that combine long temporal context with broad spatial coverage for cooperative perception. |
arXiv 2026 | |
| DriveVA: Video Action Models are Zero-Shot Drivers | DetailsExplores whether video action models can act as zero-shot drivers by mapping visual context directly to driving actions. |
arXiv 2026 | |
| Latent Chain-of-Thought World Modeling for End-to-End Driving | DetailsIntroduces latent chain-of-thought reasoning inside driving world models to improve future prediction and downstream planning. |
arXiv 2025 | |
| GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation | DetailsUses 4D occupancy guidance for physics-aware driving video generation, improving spatial and temporal consistency in world-model rollouts. |
arXiv 2025 | |
| X-World: Controllable Ego-Centric Multi-Camera World Models for Scalable End-to-End Driving | DetailsBuilds a controllable ego-centric multi-camera world model to scale closed-loop evaluation and training for end-to-end driving policies. |
arXiv 2026 | |
| DynFlowDrive: Flow-Based Dynamic World Modeling for Autonomous Driving | DetailsIntroduces a flow-based dynamic world model to better capture multi-modal future scene evolution for robust planning. |
arXiv 2026 | Code |
| Latent World Models for Automated Driving: A Unified Taxonomy, Evaluation Framework, and Open Challenges | DetailsProvides a structured taxonomy and evaluation framework for latent driving world models, highlighting open challenges for reliable deployment. |
arXiv 2026 | |
| Kinematics-Aware Latent World Models for Data-Efficient Autonomous Driving | DetailsIntroduces kinematics-aware latent world modeling for improved data efficiency and physically consistent planning in autonomous driving. |
arXiv 2026 | |
| MAD: Motion Appearance Decoupling for efficient Driving World Models | DetailsDecouples structured motion learning from appearance synthesis, adapting video diffusion models into controllable driving world models more efficiently. |
arXiv 2026 | Project |
| WorldRFT: Latent World Model Planning with Reinforcement Fine-Tuning for Autonomous Driving | DetailsAligns latent world-model representation learning with planning through hierarchical decomposition and reinforcement fine-tuning for safer end-to-end driving. |
AAAI 2026 | |
| Other Vehicle Trajectories Are Also Needed: A Driving World Model Unifies Ego-Other Vehicle Trajectories in Video Latent Space | DetailsUnifies ego and surrounding-vehicle trajectory modeling in video latent space, improving interaction-aware driving world prediction. |
AAAI 2026 | |
| InDRiVE: Reward-Free World-Model Pretraining for Autonomous Driving via Latent Disagreement | DetailsUses latent ensemble disagreement as intrinsic motivation for reward-free world-model pretraining, enabling reusable exploration policies for driving. |
arXiv 2025 | |
| DriveLaW:Unifying Planning and Video Generation in a Latent Driving World | DetailsUnifies video generation and motion planning by sharing latent world representations between future prediction and trajectory generation. |
arXiv 2025 | |
| 3D-VLA: A 3D Vision-Language-Action Generative World Model | Details |
ICML 2024 | Code |
| CarDreamer: Open-source learning platform for world-model-based autonomous driving | DetailsProvides an open-source learning platform for autonomous-driving research with world-model-based simulation and policy training. |
IOTJ 2025 | Code |
| VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving | Details |
arXiv 2025 | Project |
| Dual-Mind World Models: A General Framework for Learning in Dynamic Wireless Networks | Details |
arXiv 2025 | |
| Addressing Corner Cases in Autonomous Driving: A World Model-based Approach with Mixture of Experts and LLMs | Details |
arXiv 2025 | |
| From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction | DetailsPredicts collaborative future states and actions with a policy world model, linking multi-agent forecasting to planning. |
NeurIPS 2025 | Code |
| Rethinking Driving World Model as Synthetic Data Generator for Perception Tasks | Details |
arXiv 2025 | Project |
| SparseWorld: A Flexible, Adaptive, and Efficient 4D Occupancy World Model Powered by Sparse and Dynamic Queries | Details |
arXiv 2025 | Code |
| OmniNWM: Omniscient Driving Navigation World Models | Details |
arXiv 2025 | Project |
| DriveVLA-W0: World Models Amplify Data Scaling Law in Autonomous Driving | Details |
arXiv 2025 | Code |
| IRL-VLA: Training a Vision-Language-Action Policy via Reward World Model | DetailsTrains a VLA policy with a reward world model, using learned future feedback to guide action optimization. |
arXiv 2025 | Project / Code |
| TeraSim-World: Worldwide Safety-Critical Data Synthesis for End-to-End Autonomous Driving | Details |
arXiv 2025 | Project |
| World4Drive: End-to-end autonomous driving via intention-aware physical latent world model | DetailsUses an intention-aware physical latent world model to connect future dynamics prediction with end-to-end trajectory planning. |
ICCV 2025 | Code |
| World model-based end-to-end scene generation for accident anticipation in autonomous driving | Details |
arXiv 2025 | |
| DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video Generation | Details |
AAAI 2025 | Project |
| GAIA-1: A Generative World Model for Autonomous Driving | Details |
arXiv 2023 | |
| Driving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous Driving | Details |
CVPR 2024 | Code / Project. |
| TrafficBots: Towards World Models for Autonomous Driving Simulation and Motion Prediction | Details |
ICRA 2023 | Code |
| MaskGWM: A Generalizable Driving World Model with Video Mask Reconstruction | DetailsLearns a generalizable driving world model through video mask reconstruction, improving future generation and representation transfer. |
CVPR 2025 | Project / Code |
| DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene Representation | DetailsUses world models as 4D data machines to generate temporally consistent driving scene representations for downstream perception tasks. |
CVPR 2025 | Project |
| X-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible Controllability | DetailsGenerates large-scale driving scenes with high fidelity and flexible controls, supporting simulation, data synthesis, and policy evaluation. |
NeurIPS 2025 | Project |
| Epona: Autoregressive Diffusion World Model for Autonomous Driving | DetailsBuilds an autoregressive diffusion world model for autonomous driving, generating future scene rollouts conditioned on prior context. |
ICCV 2025 | Project / Code |
| SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model | DetailsUses a generative world model for city-scale traffic simulation, enabling controllable multi-agent scenario synthesis. |
CVPR 2025 | |
| ReSim: Reliable World Simulation for Autonomous Driving | DetailsProvides a reliable world simulation framework for autonomous driving that emphasizes realistic closed-loop behavior and policy evaluation. |
NeurIPS 2025 | Project / Code |
| End-to-end driving with online trajectory evaluation via BEV world model | DetailsUses a BEV world model to evaluate trajectories online, improving end-to-end driving decisions through future-scene assessment. |
ICCV 2025 | Code |
| Raw2Drive: Reinforcement Learning with Aligned World Models for End-to-End Autonomous Driving (in CARLA v2) | DetailsAligns world models with reinforcement learning in CARLA v2, training end-to-end policies from raw observations through imagined and closed-loop feedback. |
NeurIPS 2025 | |
| Semi-supervised vision-centric 3d occupancy world model for autonomous driving | Details |
ICLR 2025 | |
| VL-SAFE: Vision-Language Guided Safety-Aware Reinforcement Learning with World Models for Autonomous Driving | Details |
arXiv 2025 | Project |
| AdaWM: Adaptive World-Model-Based Planning for Autonomous Driving | DetailsUses adaptive world-model-based planning to improve future prediction and trajectory selection in autonomous driving. |
ICLR 2025 | |
| Genad: Generative end-to-end autonomous driving | Details |
ECCV 2024 | Code |
| COME: Adding Scene-Centric Forecasting Control to Occupancy World Model | DetailsAdds scene-centric forecasting control to occupancy world models, improving controllable future prediction for autonomous driving. |
NeurIPS 2025 | Code |
| UniWorld: Autonomous Driving Pre-training via World Models | Details |
arXiv 2023 | |
| Occ-LLM: Enhancing Autonomous Driving with Occupancy-Based Large Language Models | Details |
ICRA 2025 | |
| UniDrive-WM: Unified Understanding, Planning and Generation World Model For Autonomous Driving | DetailsUnifies VLM-based scene understanding, trajectory planning, and trajectory-conditioned future image generation within one driving world model. |
arXiv 2026 | Project |
- nuScenes : A large-scale multimodal dataset widely used for various AD tasks, including 3D object detection, tracking, and prediction. It features data from cameras, LiDAR, and radar, along with full sensor suites and map information. Several works like LightEMMA , OpenDriveVLA , and GPT-Driver utilize nuScenes for evaluation or data generation. It is also used for tasks like BEV retrieval and dense captioning.
- 4DLidarOpen : An open multi-modal autonomous driving dataset centered on 4D FMCW LiDAR, providing point-wise radial velocity, multiple LiDAR types, surround-view cameras, ego poses, 3D boxes, track IDs, and benchmarks for detection, BEV segmentation, flow prediction, motion forecasting, and planning.
- XWOD : Extreme Weather Object Detection benchmark with real-world traffic images across rain, snow, fog, haze/sand/dust, flooding, tornado, and wildfire conditions for studying weather-robust autonomous driving perception.
- ScenePilot-Bench : A first-person driving VLM benchmark built on ScenePilot-4K, evaluating scene understanding, spatial perception, motion planning, and safety-aware reasoning across large-scale driving videos.
- Bench2Drive-Robust : A closed-loop robustness benchmark for E2E-AD under deployment perturbations such as camera-stream failures, ego-state errors, and compute-induced control delay.
- MDrive : A closed-loop cooperative driving benchmark for end-to-end multi-agent systems, covering V2X perception sharing, negotiation, Real2Sim conversion, and human-in-the-loop simulation.
- HiDrive : A closed-loop benchmark emphasizing high-level and long-tail driving capabilities such as rule compliance, ethical reasoning, emergency response, and rare-object interaction.
- EgoDyn-Bench : A diagnostic benchmark for evaluating ego-motion understanding and physical grounding in VLMs, MLLMs, and driving VLA models.
- WorldLens : A full-spectrum benchmark for driving world models, evaluating generation, reconstruction, action-following, downstream-task usefulness, and human preference with WorldLens-26K annotations.
- BDD-X (Berkeley DeepDrive eXplanation) : This dataset provides textual explanations for driving actions, making it particularly relevant for training and evaluating interpretable AD models. DriveGPT4 and ADAPT are evaluated on BDD-X. It contains video sequences with corresponding control signals and natural language narrations/reasoning.
- Waymo Open Dataset (WOMD) : A large and diverse dataset with high-resolution sensor data, including LiDAR and camera imagery. Used in works like OmniDrive for Q&A data generation and by LLMs Powered Context-aware Motion Prediction. Also used for scene simulation in ChatSim. WOMD-Reasoning is a language dataset built upon WOMD focusing on interaction descriptions and driving intentions.
- DriveLM : A benchmark and dataset focusing on driving with graph visual question answering. TS-VLM is evaluated on DriveLM. It aims to assess perception, prediction, and planning reasoning through QA pairs in a directed graph, with versions for CARLA and nuScenes.
- LingoQA : A benchmark and dataset specifically designed for video question answering in autonomous driving. It contains over 419k QA pairs from 28k unique video scenarios, covering driving reasoning, object recognition, action justification, and scene description. It also proposes the Lingo-Judge evaluation metric.
- DriveAction : DriveAction leverages real-world driving data actively collected by users of production-level autonomous vehicles to ensure broad and representative scenario coverage, provides high-level discrete behavior labels collected directly from users' actual driving operations, and implements a behavior-based tree-structured evaluation framework that explicitly links vision, language, and behavioral tasks to support comprehensive and task-specific evaluation.
- CARLA Simulator & Datasets : While a simulator, CARLA is extensively used to generate data and evaluate AD models in closed-loop settings. Works like LeapAD , LMDrive , and LangProp use CARLA for experiments and data collection. DriveLM-Carla is a specific dataset generated using CARLA.
- Argoverse : A dataset suite with a focus on motion forecasting, 3D tracking, and HD maps. Argoverse 2 is used in challenges like 3D Occupancy Forecasting.
- KITTI : One of the pioneering datasets for autonomous driving, still used for tasks like 3D object detection and tracking.
- Cityscapes : Focuses on semantic understanding of urban street scenes, primarily for semantic segmentation.
- UCU Dataset (In-Cabin User Command Understanding) : Part of the LLVM-AD Workshop, this dataset contains 1,099 labeled user commands for autonomous vehicles, designed for training models to understand human instructions within the vehicle.
- MAPLM (Large-Scale Vision-Language Dataset for Map and Traffic Scene Understanding) : Also from the LLVM-AD Workshop, MAPLM combines point cloud BEV and panoramic images for rich road scenario images and multi-level scene description data, used for QA tasks.
- NuPrompt : A large-scale language prompt set based on nuScenes for driving scenes, consisting of 3D object-text pairs, used in Prompt4Driving.
- nuDesign : A large-scale dataset (2300k sentences) constructed upon nuScenes via a rule-based auto-labeling methodology for 3D dense captioning.
- LaMPilot : An interactive environment and dataset designed for evaluating LLM-based agents in a driving context, containing scenes for command tracking tasks.
- DRAMA (Joint Risk Localization and Captioning in Driving) : Provides linguistic descriptions (with a focus on reasons) of driving risks associated with important objects.
- Rank2Tell : A multimodal ego-centric dataset for ranking importance levels of objects/events and generating textual reasons for the importance.
- HighwayEnv : A collection of environments for autonomous driving and tactical decision-making research, often used for RL-based approaches and LLM decision-making evaluations (e.g., by DiLu, MTD-GPT).
These repositories offer broader collections of resources that may overlap with or complement the focus of this list.
- Awesome-LLM4AD (LLM for Autonomous Driving):
- GitHub: https://github.com/Thinklab-SJTU/Awesome-LLM4AD
- Alternative Link:http://codesandbox.io/p/github/sorokinvld/Awesome-LLM4AD
- Description: A curated list of research papers about LLM-for-Autonomous-Driving, categorized by planning, perception, question answering, and generation. Continuously updated.
- Awesome-VLLMs (Vision Large Language Models):
- GitHub: JackYFL/awesome-VLLMs
- Description: Collects papers on Visual Large Language Models, with a dedicated section for "Vision-to-action" including "Autonomous driving" (Perception, Planning, Prediction).
- Awesome-Data-Centric-Autonomous-Driving:
- GitHub: https://github.com/LincanLi98/Awesome-Data-Centric-Autonomous-Driving
- Description: Focuses on data-driven AD solutions, including datasets, data mining, and closed-loop technologies. Mentions the role of LLMs/VLMs in scene understanding and decision-making.
- Awesome-World-Model (for Autonomous Driving and Robotics):
- GitHub: https://github.com/LMD0311/Awesome-World-Model
- Description: Records, tracks, and benchmarks recent World Models for AD or Robotics, supplementary to a survey paper. Includes many VLA-related and generative model papers.
- Awesome-Multimodal-LLM-Autonomous-Driving:
- GitHub: https://github.com/IrohXu/Awesome-Multimodal-LLM-Autonomous-Driving
- Description: A systematic investigation of Multimodal LLMs in autonomous driving, covering background, tools, frameworks, datasets, and future directions.
- Awesome VLM Architectures:
- GitHub: gokayfem/awesome-vlm-architectures
- Description: Contains information on famous Vision Language Models (VLMs), including details about their architectures, training procedures, and datasets.
If you find this repository helpful, a citation to our paper would be greatly appreciated:
@ARTICLE{xu2026survey,
author={Xu, Chengkai and Cui, Yiming and Liu, Jiaqi and Guo, Yicheng and Qin, Cheng and Zhang, Geyuan and Dong, Xinwei and Fang, Shiyu and Hang, Peng and Sun, Jian},
journal={IEEE Transactions on Intelligent Transportation Systems},
title={A Survey on End-to-End Autonomous Driving Training From the Perspectives of Data, Strategy, and Platform},
year={2026},
volume={},
number={},
pages={1-20},
keywords={Modeling;Training;Optimization;Autonomous driving;Safety;Surveys;Vehicles;Testing;Learning (artificial intelligence);Reinforcement learning;Autonomous vehicle;end-to-end;artificial intelligence;intelligent transportation system},
doi={10.1109/TITS.2026.3695999}
}