Curated papers, datasets, systems, and benchmarks for robot data engines across robot-centric, UMI, human/egocentric, and simulation data.
-
Updated
Oct 6, 2026 - CSS
Curated papers, datasets, systems, and benchmarks for robot data engines across robot-centric, UMI, human/egocentric, and simulation data.
Scale AI — independent third-party profile of a public API surface, by API Evangelist. Scale AI is the data engine for AI. The company turns raw data into training data by combining ML-powered pre-labeling with multi-tier human review, and ships an extensive REST API and SDKs for managing labeling, evaluation, and generative-AI data pipelines.
Eventual — independent third-party profile of a public API surface, by API Evangelist. Eventual is a data infrastructure company building Daft, an open-source, high-performance data engine for AI and multimodal workloads. Written in Rust with Python and SQL interfaces, Daft lets teams query and process images, audio, video, documents, embeddings, a
LOGOS — an event-driven data engine that transforms raw inputs into structured, append-only data within the AIOS system.
Local-first Rust data engine: SQL, native structures, lexical/vector search, WAL, MVCC, recovery, and verifiable proofs in one binary.
Build and deploy AI agents for automated labeling, quality assurance, and workflow automation in the Encord ecosystem.
An India-native, self-improving data engine for autonomous driving. Auto-labeling with a calibrated confidence gate, active learning, HD maps, per-object dynamics, and a governed closed loop.
Library for building offline-first browser-based applications :: платформа автономных веб-приложений
An immutable database that follows Snowflake architecture, designed for scalable, replayable systems beyond analytics.
Automated visual data harvesting and VLM grounding engine for indoor door detection, video frame filtering, pseudo-labeling, and leakage-free YOLO dataset synthesis.
Mcity Data Engine
特斯拉自动驾驶架构思想向自驱动实验室 SDL 迁移—— AI for Science 基础设施视角 | Transferring Tesla's autonomous-driving architecture ideas to Self-Driving Labs — an AI for Science infrastructure perspective
End-to-end pipeline ingesting product docs from GitHub, refining through bronze/silver/gold on AWS, and serving to an AI agent
Auto-labeling and VLM scene understanding pipeline for driving data — open-vocabulary detection (Grounding DINO), SAM segmentation, traffic-rule-aware scene tagging (Qwen2.5-VL), pseudo-label mAP evaluation and long-tail mining. Designed to fit in 6 GB VRAM.
Zyquo Database (ZDB) — a native Apple Silicon data engine in Swift: columnar tables, vectors, documents, mini-SQL, compression, mmap, HTTP dashboard. Zero dependencies.
Zero-copy Rust data engine for Python. Memory-safe, blazingly fast, seamless Python integration via Arrow PyCapsule Interface.
A robust In-Memory Database Engine implemented in Java, focusing on professional software architecture, Clean Code, and the application of 7 Design Patterns.
Low-latency, multi-threaded data processing system designed to interface directly with low-level hardware
[NeurIPS 2025] Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection
A unified hierarchical engine for managing parameters, state, workflows, and knowledge with structure, meaning, and AI-powered reasoning.
To associate your repository with the data-engine topic, visit your repo's landing page and select "manage topics."