VideoDB’s cover photo
VideoDB

VideoDB

Technology, Information and Internet

San Francisco, California 3,262 followers

An applied research lab for visual intelligence. Data infrastructure for video, built to train and deploy visual AI.

About us

VideoDB is an applied research lab for visual intelligence. We build the data infrastructure for video that trains and deploys visual AI. We turn video into data you can use. We investigate where your model fails, find the data that fixes it, and bring scalable infrastructure and the expertise to go with it, so your team builds intelligence from video faster. Video primitives: in-house transcoding, frame tiling, a frame-addressable streaming engine, real-time ingestion and high-throughput pipelines. 100,000 hours turned into scene-level samples in two weeks, files left in place. Machine annotation: any model as an analyzer over every hour you have, at archive scale, with human QC through our partners. Episode retrieval: world-leading visual search over every episode. Moments and frames in 500 ms, deep search agents on top. Realtime ingestion: connect 1000s of cameras, streams and screens, analyze in real time, alerts with the clip attached. Connected to your robot data: MCAP, LeRobot and RLDS in. Manifests and loaders out. We publish benchmarks, build and test systems, and ship what holds up. Run in our cloud or yours. 10k+ developers on the platform. 250+ TB processed.

Website
https://videodb.io/
Industry
Technology, Information and Internet
Company size
11-50 employees
Headquarters
San Francisco, California
Type
Privately Held
Founded
2024
Specialties
visual intelligence, physical AI, video infrastructure, data infrastructure, machine annotation, video retrieval, visual search, realtime video, computer vision, multimodal AI, AI agents, model training, robot learning data, MCAP, LeRobot, RLDS, benchmarks, and research

Locations

Employees at VideoDB

Updates

  • A language model can tell you what usually comes next. A world model should tell you what happens if you act. That is the core of JEPA (Joint-Embedding Predictive Architecture), Yann LeCun's bet on how AI systems should learn. Our essay follows the idea from language models to vision-language-action models and world models for robots. Next-token prediction took us far, but its loss only checks the next visible token. A model can land on the right word while its hidden state drifts away from a coherent future. JEPA changes what gets predicted. It encodes a context view and predicts the embedding of a target view, such as a masked region or a future state, with no pixel reconstruction and no token decoding. Add actions, and the model can predict the consequence of an action before taking it. The early results are already sharp. In one hierarchical world-model paper the essay discusses, a flat planner (VJEPA2-AC) gets 0% success on a Franka pick-and-place task, while the hierarchical version reaches 70% on the cup task. The essay is careful to say these numbers hold under the evaluated setup. It also covers the main risk: everything happens in latent space, and a latent space can collapse or warp while the math still looks elegant. The stack the essay expects keeps language as the interface for instructions and explanations, with a world model holding and predicting state underneath. The full essay is in the first comment. We would love to hear how your team is thinking about world models.

  • How much of video search comes down to the model, and how much to the system around it? Our team answered that in a technical report, Search over the Visual World (arXiv 2608.08075), with the ideas built and tested in VideoDB. The starting point: video keeps arriving from archives, recordings and live streams, and an answer only helps if someone can press play on it and check it. So the report treats search as a set of system decisions, starting with the analyzer: a model run over the video with its own settings: how it cuts the timeline into scenes, what it is asked, and the shape of its answer. VideoDB keeps analyzer outputs aligned on source time, so a new index can be built without running the model again. Every result stays playable at the original source. Then we tested it. Our pipeline uses general-purpose components, none trained end-to-end for video retrieval. We ran it against a commercial video-native retrieval engine, with both systems indexing the same videos and receiving the same queries: +7.3 points on top-1 accuracy (R@1) across 9,800+ queries on four open datasets It also leads at R@3 and R@10. The practical point for anyone building on video: today, video search is a systems problem. Swapping an analyzer or an index changes what you can find, and a video-native model can join as one more analyzer. The full paper is in the first comment. We would love to hear how you search your own video.

  • Can a VLM annotate a robot run well enough to train on? We wanted a real answer, so we took WGO-Bench, MacroData's public benchmark of 100 human and robot manipulation videos, and ran our annotation pipeline against a public baseline. Same model on both sides, Gemini 3.7 Flash. Here is where we landed: VideoDB: 28.93% semantic F1 Public baseline: 19.92% Before scoring anything, we wanted a reference we could trust. So our team went through all 100 videos by hand and rebuilt the labels around one rule: one completed manipulation per segment, so a reach stays with the grasp it leads to. That took the reference from 743 segments to 534, with changes in 90 of the videos. Then we scored both pipelines against it. A segment only counted if its timing matched (IoU 0.5) and a judge accepted its label. Most of our lead came from the labels themselves: 149 accepted, compared with 100 for the baseline. This matters because robot policies such as π0.5 learn from labels like these. A clean boundary and the right object name give them a better training signal. We are now taking this further: tightening segment timing on HomER's egocentric video, sharpening labels on DROID, and flagging the annotations that deserve a human review. The full benchmark is in the first comment. We would love to hear how you are annotating your own robot data.

  • VideoDB reposted this

    We’re making the egocentric internet searchable. 12M+ task episodes are now searchable with VideoDB. Type “a person throwing a ball” and get thousands of matching episodes back in seconds. Export to LeRobot, MCAP or RLDS. Connect the MCP server and your agents can run the same search. Have an egocentric dataset? Submit it so researchers can find it too. Watch it in action below. Link in the comments.

  • Most robotics labs keep more video than anyone can look at. Teleoperation sessions, test runs on the robot and rollouts from a world model are all recorded, and they all go into a bucket. When a researcher wants to know how often the gripper slips on one part, someone writes a pipeline. The usual approach cuts the archive into clips up front, before anyone knows the next question. Each new question means a new pipeline and the whole archive processed again, so compute is spent per experiment rather than per hour of video. Test runs are worse: they land in a separate bucket and never join the training data. We built a query engine that keeps every episode whole. Any model can run as an analyzer over the archive and write down what happened and when. The index is versioned, so a better model means re-reading the archive instead of re-collecting it. A question in plain English decides the window: a slip needs about five seconds of context, a regrasp ninety, a whole shift twenty minutes. Six question types cover most of what researchers ask: find, count, filter, sample, look closer, watch. The hits from one question become a cohort, the cohort becomes a training manifest, and the next model's runs land in the same archive, searchable the moment they arrive. Full write-up in the first comment.

    • No alternative text description for this image
  • VideoDB reposted this

    Dyna Robotics trained a robot model on a million hours of human video. Figure's app receives thirty minutes of video every second. Let's take this scenario: what happens to the video after the run? In most labs it is processed once and stored forever. Every new question a researcher has means a new pipeline over the whole archive. And the test runs, the hours where the model actually failed, mostly never get indexed at all. The scaling law is proven. The recording is funded. The missing layer is the one that lets a lab ask its archive a question in plain language and get the exact moments back in under a second. I wrote down the questions I think every lab should be able to ask, and the loop that makes every run improve the next one. [Slides below] Where we come in: we start with three questions. Why can the model not do this task? Under what conditions does it fail? Is there enough variety in the data to fix it? We build the infrastructure that makes those questions cheap to ask. Your agents live inside your archive, watch every training run and test run as it lands, and come back with the failure modes and the evidence, not a summary. Then we fix the mix: the episodes you already have, and more sources, open or proprietary, when you run out. It sits on top of your bucket, reads MP4, MCAP, LeRobot, RLDS and live streams, and your files never move. We are starting pilots with labs on about a hundred hours of their own runs. If your test runs sit in a bucket nobody can ask, I would like to compare notes. I will be at Long Horizon in San Francisco on 20-21 October. The next models will not be trained on more video. They will be trained on the right moments.

  • Multimodal models have a problem that has nothing to do with perception. A model can identify every person in a crowded photograph and still count the group wrong. Not because it missed anyone - because it lost track of who it already counted mid-reasoning. DeepSeek's recent paper calls this the Reference Gap. Their proposed solution is direct: let the model emit bounding boxes and point sequences while it reasons, interleaved with language. A box becomes working memory for an object. A point sequence becomes working memory for a path. The idea matters more than any individual benchmark score. Multimodal reasoning needs a representation for spatial state that survives a long reasoning trace. Language can describe spatial state, but descriptions become ambiguous in crowded scenes, overlapping layouts, and multi-step deductions. Our team wrote a detailed breakdown - covering the math, the five-stage training pipeline, the reward design for counting and maze navigation, and what the paper establishes versus what it leaves open. Runnable Python is included throughout. If you work with vision-language models, this is worth reading carefully. Link in comments.

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • VideoDB reposted this

    We beat a video-native embedding model at video search. We didn't train a model. We built a system. 73.1% Recall@1 versus 65.8%, across 9,834 queries and four public datasets — using off-the-shelf analyzers you can swap any time. Here's the reasoning behind it. One hour of video is about a million tokens. Ten novels, every hour. That's the reason video infrastructure sits a decade behind text. Models can already read a frame. Hand GPT or Gemini an image and you get everything in it. But video isn't a series of frames. It's speech, motion, faces, emotion and sound over time and what comes before and after changes what a moment means. So we stopped treating retrieval as a model problem and started treating it as a system. Ingest the video, split it into scenes, run any model you like, and index every output: speech, objects, on-screen text, embeddings, your own domain signals. At query time the system resolves the boundary from your question, not from wherever the file happened to be cut. Then it returns a playable stream of exactly those seconds, in under a second. Paper link in the comments.

  • VideoDB reposted this

    We beat a video-native embedding model at video search. We didn't train a model. We built a system. 73.1% Recall@1 versus 65.8%, across 9,834 queries and four public datasets — using off-the-shelf analyzers you can swap any time. Here's the reasoning behind it. One hour of video is about a million tokens. Ten novels, every hour. That's the reason video infrastructure sits a decade behind text. Models can already read a frame. Hand GPT or Gemini an image and you get everything in it. But video isn't a series of frames. It's speech, motion, faces, emotion and sound over time and what comes before and after changes what a moment means. So we stopped treating retrieval as a model problem and started treating it as a system. Ingest the video, split it into scenes, run any model you like, and index every output: speech, objects, on-screen text, embeddings, your own domain signals. At query time the system resolves the boundary from your question, not from wherever the file happened to be cut. Then it returns a playable stream of exactly those seconds, in under a second. Paper link in the comments.

  • VideoDB reposted this

    I thought I was walking into another AI meetup. Instead within half an hour everyone had their laptops open and was building. Recently I attended Morning Sessions with Builders by VideoDB After a quick introduction to VideoDB and how it helps AI understand and retrieve information from videos... there weren't any long presentations or endless slides. The room shifted straight into building mode. I started working on Rehearse that is an AI interview analyzer built using VideoDB. It's not finished. It wasn't ready for the demo session. And that's okay. Watching people around me ship ideas reminded me that building isn't about getting everything right on the first try. It's about taking an idea from your head and turning it into something real… even if it's still rough around the edges. One thing that stood out was VideoDB itself. Instead of treating videos as files to store… it treats them as data that AI can understand… search and reason over. That opens up a lot of interesting possibilities beyond simple transcription. I'm leaving with: a project that I genuinely want to finish, a better understanding of AI infrastructure for video, and a reminder that the best way to learn is still to build. Looking forward to shipping Rehearse soon. Had a great time with Omkar, PRATHAMESH, Pranay, Siddhi! Thanks to the VideoDB team for putting together a session that was less about talking and more about building. Om Gate Almas S. #BuildInPublic #AI #VideoDB #Builders #GenAI #Engineering #Technology #studentbuilder

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
      +3

Similar pages

Browse jobs