NarrativeFlow

Flow-Based Vision-Language-Action Model Using Robot Velocity Fields

Shota Kobayashi, Koki Seno, Daichi Yashima, Komei Sugiura
Keio University

{shotakoba10267, koki.seno, ydaichi1207, komei.sugiura}@keio.jp

ACCV 2026

NarrativeFlow overview figure

Overview of our approach to language-conditioned flow-based manipulation. We formulate robot flows as continuous velocity fields via flow matching. For training, we use cross-embodiment data. At test time, given a language instruction and an initial image, our model generates a robot flow. Based on the robot flow, the robot then executes the object manipulation task.

Abstract

We focus on language-conditioned flow-based manipulation, where robot flows (robot velocity fields) serve as embodiment-agnostic, motion-centric representations for leveraging cross-embodiment data. This task is crucial because language-conditioned manipulation is essential for practical robotic systems, yet scaling robot foundation models remains bottlenecked by the labor-intensive collection of embodiment-specific data.

Existing methods either coarsely approximate robot flows with sparse keypoint displacements, or cannot handle language-conditioned manipulation. To address this limitation, we propose NarrativeFlow, which models robot flows as continuous velocity fields using a flow-matching formulation conditioned on language.

Moreover, we introduce an auxiliary objective that aligns the model's intermediate representations with task-relevant scene changes. Accordingly, NarrativeFlow captures task-relevant semantics, and generates robot flows that are physically consistent with real-world manipulation.

To validate NarrativeFlow, we have conducted experiments on standard datasets for language-conditioned manipulation. The experimental results show that NarrativeFlow outperforms representative baseline methods on standard evaluation metrics. Furthermore, through real-world experiments, we show that NarrativeFlow achieves higher success rates than baseline methods across multiple manipulation tasks.

Methodology

We propose NarrativeFlow, a method for language-conditioned robot flow generation inspired by flow-based manipulation policies. The novelties of NarrativeFlow are as follows.

  1. NarrativeFlow formulates robot flows as probability velocity fields within a flow-matching framework, enabling high-quality robot flow generation.
  2. NarrativeFlow introduces the Narrative Delta loss, which aligns the model's intermediate representations with multiple narrative representations of task-relevant scene changes between an initial image and a goal image.
  3. NarrativeFlow uses two subtask tokens that produce separate joint representations for conditioning flow generation and computing the Narrative Delta loss, mitigating interference between the learning signals from the two downstream paths.
NarrativeFlow model architecture

Architecture of NarrativeFlow. The vision-language fusion encoder integrates the embeddings of the language instruction and the initial image into the two subtask tokens: one for robot flow generation, and the other for the Narrative Delta loss. During training, the Narrative Delta loss aligns the corresponding subtask token with the narrative representations of task-relevant scene changes.

Experiments

Experiment rollout 1
Experiment rollout 2
Experiment rollout 3
Qualitative flow prediction results

Qualitative comparison between NarrativeFlow and a baseline method (Im2Flow2Act). In each subfigure, the first row presents the language instruction, the initial image, and the ground-truth robot flow. The second and third rows show the predicted robot flows of the baseline method and the proposed method, respectively.

Quantitative results on Fractal and Bridge V2

Quantitative comparison between the proposed method and baseline methods. The best scores for each metric are shown in bold.

Ablation study for subtask tokens and Narrative Delta loss

Ablation study for subtask tokens and Narrative Delta loss. ✓ indicates the use of the Narrative Delta loss. The best scores for each metric are shown in bold.

Real-World Experiments

Real-world mobile manipulation rollouts

Qualitative results of real-world experiments. The panels show successful executions of three manipulation tasks: (i) mobile close drawer, (ii) mobile bin picking, and (iii) mobile stack cup. In each panel, the top row shows the language instruction, the initial image, and the generated robot flow, while the bottom row shows the downstream manipulation conditioned on the robot flow.

1x
4x
2x
Real-world mobile manipulation results

Quantitative results of real-world experiments. For each method, we report success rates across 20 trials per task, along with the average rate over the three tasks. For the oracle method, the policy was conditioned on the corresponding ground-truth robot flows. The best scores among the proposed method and the baseline methods are shown in bold.

BibTeX

Coming soon.