-
Shanghai AI Lab
- Shanghai
- https://tai-wang.github.io/
- @wangtai97
Stars
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
Elemental Diagnosis of Generalist Mobile Manipulation Policies
InternVLA-A1: Unifying Understanding, Generation, and Action for Robotic Manipulation
MMSI-Video-Bench: A Holistic Benchmark for Video-Based Spatial Intelligence
[CVPR 2026] G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
InternVLA-M1: A Spatially Guided Vision-Language-Action Framework for Generalist Robot Policy
[ICLR 2026] The offical Implementation of "Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model"
An All-in-one robot manipulation learning suite for policy models training and evaluation on various datasets and benchmarks.
InternRobotics' open-source toolbox for vision-based embodied spatial intelligence.
A versatile, all-in-one toolbox for whole-body humanoid robot control.
InternRobotics' open platform for building generalized navigation foundation models.
[NeurIPS 2025] InternScenes: A Large-scale Interactive Indoor Scene Dataset with Realistic Layouts.
[ICRA 2026] Official implementation of the paper: "StreamVLN: Streaming Vision-and-Language Navigation via SlowFast Context Modeling"
[ICLR 2026] MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
[ICRA 2026] NavDP: Learning Sim-to-Real Navigation Diffusion Policy with Privileged Information Guidance
[NeurIPS 2025] OST-Bench: Evaluating the Capabilities of MLLMs in Online Spatio-temporal Scene Understanding
[ICCV 2025] GLEAM: Learning Generalizable Exploration Policy for Active Mapping in Complex 3D Indoor Scene
A Next-Generation Training Engine Built for Ultra-Large MoE Models
Official Implementation of paper "MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion"
[ICCV 2025] A Simple yet Effective Pathway to Empowering LLaVA to Understand and Interact with 3D World
A simulation platform for versatile Embodied AI research and developments.
Code&Data for Grounded 3D-LLM with Referent Tokens
Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
[CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
Official implementation of the paper "PACER+: On-Demand Pedestrian Animation Controller in Driving Scenarios" (CVPR 2024).
[NeurIPS 2023] Official Code for "SMPLer-X: Scaling Up Expressive Human Pose and Shape Estimation"
[CVPR 2024 & NeurIPS 2024] EmbodiedScan: A Holistic Multi-Modal 3D Perception Suite Towards Embodied AI
Learning-based locomotion control from OpenRobotLab, including Hybrid Internal Model & H-Infinity Locomotion Control
[NeurIPS 2023] OV-PARTS: Towards Open-Vocabulary Part Segmentation



