Skip to content

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

RoboHarn-Evo

Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Turn physical experience into reusable knowledge.

Paper Project Page Demos License

Full affiliations & contributions

Shifeng Bao1,3*, Fanding Huang2*, Yihan Lin1,3, Youhe Feng1,3, Guanlin Li1,3, Chen Zhao1,3, Yang Li1,3, Jiawei He5, Cheng Chi1,4‡, Jing Zhang1,4†

1 School of Information, Renmin University of China
2 Tsinghua University
3 Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China
4 Engineering Research Center of Database and Business Intelligence, Beijing, China
5 XYZ Embodied AI, Beijing, China

* Equal contribution. † Corresponding author. ‡ Project leader.

Watch RoboHarn-Evo on a real robot: uncover blocks, count them, and press the matching buttons
Watch the real-robot demonstrations →

🔥 News

  • 2026.10.05: Our project page is live, with three real-robot demonstrations, method figures, and interactive results.

🧭 Overview

Physical experience can improve what a robot knows. RoboHarn-Evo develops Hierarchical Physical Knowledge (HPK) through interaction, with the vision-language model and low-level executor held fixed.

  • Task Knowledge guides subtask selection, ordering, goals, and completion conditions.
  • Action Knowledge guides object-relative geometry, interaction strategies, and physical effects.

An inner execution loop retrieves both levels of knowledge and grounds actions in the current scene. An outer knowledge-update loop uses recorded physical outcomes to revise, consolidate, and organize knowledge for later episodes. Action effects and subtask completion are assessed separately.

RoboHarn-Evo architecture: the inner loop uses Task and Action Knowledge during execution; the outer loop updates knowledge from physical evidence

📖 Abstract

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.

🧩 Method

  1. Retrieve knowledge at each decision level. Skill summaries route queries to related entries; Task Knowledge guides planning, while Action Knowledge supports action grounding.
  2. Ground actions in the current scene. Object-relative strategies and the task goal constrain motion targets and expected effects.
  3. Check physical outcomes. Execution evidence distinguishes the effect of an action from completion of the broader subtask.
  4. Update knowledge across episodes. Reflect on trajectories, consolidate related entries, and revise their content and applicability using supporting, opposing, or unresolved evidence.
Skill-routed retrieval and evidence-driven maintenance

Skill summaries route knowledge retrieval, while evidence from trajectories supports incremental knowledge maintenance

📊 Results

RMBench: six-task mean success. Each model is compared with and without HPK using the same executor. The paper reports 20 held-out episodes per task.

Model Without HPK Full HPK Gain
Qwen3.8-27B 19.2% 41.7% +22.5 pp
GPT-5.5 58.3% 82.5% +24.2 pp
GPT-6 72.5% 91.7% +19.2 pp

Continued interaction. In the separate learning experiment, knowledge evolves over 80 source rollouts. Held-out evaluation uses frozen knowledge snapshots and contributes no feedback to updates.

Model Before interaction After 80 rollouts
GPT-5.5 48.3% 75.0%
GPT-6 70.0% 88.3%

Values are mean held-out success over three learning histories. After 80 rollouts, 83.3% / 91.7% of initial knowledge errors are repaired or deactivated for GPT-5.5 / GPT-6, while both retain 95.8% of initially correct entries.

Transfer to physical robots. Across three tasks with eight scenes per task, mean success rises from 41.7% to 62.5% after real-world knowledge updates; mean normalized task progress rises from 70.5% to 84.2%. These paper statistics are separate from the illustrative recordings below. Frozen knowledge also transfers from RMBench to RoboDojo, improving zero-shot success by 35.0 pp for GPT-5.5 and 25.0 pp for GPT-6.

See the paper and interactive results for task-level comparisons and evaluation details.

🎬 Real-Robot Demonstrations

Task Demonstration Video
Cover blocks Apply simulation-derived knowledge to real-world covering Watch MP4
Press by number Use accumulated experience for ordered button presses Watch MP4
Uncover, count & press Reuse knowledge in a composed manipulation task Watch MP4

Clips are edited excerpts at 3× recorded speed, with waiting intervals omitted. They illustrate knowledge use rather than a controlled success-rate comparison. The project page provides chapters, knowledge summaries, and recorded outcome qualifications, including unresolved automatic verification of the blue-button effect in the third task.

🛠️ Installation

Use Python 3.10 or newer in the environment for your selected benchmark:

git clone https://github.com/RUCKBReasoning/RoboHarn-Evo.git
cd RoboHarn-Evo
python -m pip install -e '.[api,expert]'

The distribution name is roboharn-evo; the import namespace is roboharn_evo. Simulator dependencies, model checkpoints, SAM3 source and weights, and licensed simulation assets are installed separately. Benchmark applications run from the source checkout.

Consult the benchmark overview, RMBench setup, and LIBERO-PRO setup for their respective runtime requirements.

🧪 Evaluation and Configuration

Discover tasks and inspect a run

These commands do not start a simulator:

python scripts/run_rmbench.py --list-tasks
python scripts/run_libero_pro.py --list-suites
python scripts/run_rmbench.py --dry-run --task cover_blocks --seed 0

After configuring the required services, model resources, and assets, use scripts/run_rmbench.py --run with explicit task, seed, deployment configuration, asset root, and output arguments. Runtime outputs are written under eval_result/.

Model and segmentation services

Inspect provider, model, checkpoint, and endpoint arguments with:

python scripts/serve_rmbench_agent_api.py --help
python scripts/serve_rmbench_pi05.py --help
python scripts/serve_rmbench_sam3.py --help

Set RMBENCH_ASSETS_ROOT to the licensed asset directory. Credentials belong in environment variables or external credential files; the key-pool example shows multiple independently configured keys. Loopback addresses refer to local services; replace /path/to/ placeholders with installed resources.

Cover Blocks service orchestration additionally uses RMBENCH_PYTHON, SAM3_PYTHON, SAM3_REPO, SAM3_CHECKPOINT, SAM3_BPE_PATH, ROBOHARN_EVO_PROVIDER_AUTH_FILE, and ROBOHARN_EVO_PROVIDER_CONFIG_FILE.

Knowledge use and experiment protocols

Setting Purpose
agent.hpk_v3.mode Select HPK usage mode
knowledge_updates_enabled Control persistent knowledge updates
agent.recovery.enable_execution_evidence Enable execution evidence; disabling it also disables online knowledge updates and reflection

The supplied deployment configuration currently disables execution evidence. Select explicit experiment settings when evaluating verified maintenance or continued learning.

RMBench experiment definitions specify source/evaluation seeds and Off, flat reflection, Task-only, Action-only, and Full conditions. Definitions retain their own budgets, instruction conditions, and split sizes; use the intended protocol rather than treating every YAML as the paper's main evaluation.

🗂️ Repository Structure

roboharn_evo/
  agent/hpk/          Knowledge extraction, storage, retrieval and maintenance
  agent/recovery/     Tool dispatch, action-effect verification and recovery
  agent/vla/          Low-level executor interfaces
  models/            Planner and model interfaces
  services/          Model and segmentation service entry points
  benchmark_adapters/ Shared observation, action and feedback interfaces
  resources/         Agent configuration and planning/execution prompts
benchmarks/          RMBench and LIBERO-PRO applications and task definitions
configs/             Service and worker configuration examples
scripts/             Runners, services, knowledge tools and evaluations
Implementation map
  • roboharn_evo/agent/runtime.py constructs the agent; agent/core/img_agent.py implements the interaction and control loops.
  • Under agent/hpk/, hierarchical_knowledge.py defines trajectory packages and Task/Action atomization; vlm_hierarchical_reflector.py interprets trajectories with the model.
  • hierarchical_store.py, semantic_consolidator.py, and family_store.py organize persistent knowledge and its Skill index; incremental_maintainer.py and rgb_maintenance.py perform evidence-aware maintenance.
  • hierarchical_retriever.py, family_router.py, and rgb_retrieval.py retrieve knowledge; goal_consistency.py connects subtask goals to action realizations.
  • roboharn_evo/agent/grasp_attachment_contract.py handles grasp evidence. Prompts for planning, perception, memory, and execution are in roboharn_evo/resources/skills/.
  • HPK configuration uses the agent.hpk_v3 namespace.

📝 Citation

If you use RoboHarn-Evo in your research, please cite:

@misc{bao2026roboharnevo,
  title={RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation},
  author={Bao, Shifeng and Huang, Fanding and Lin, Yihan and Feng, Youhe and Li, Guanlin and Zhao, Chen and Li, Yang and He, Jiawei and Chi, Cheng and Zhang, Jing},
  year={2026},
  eprint={2609.37583},
  archivePrefix={arXiv},
  primaryClass={cs.RO},
  url={https://arxiv.org/abs/2609.37583}
}

📄 License and Third-Party Attribution

RoboHarn-Evo original code is licensed under Apache-2.0. Copyright 2026 Shifeng Bao.

Third-party code retains its original copyright notices and licenses; see the benchmark-specific LICENSE and NOTICE files. OpenPI source under benchmarks/rmbench/policy/pi05/ retains its Apache-2.0 license. Model weights, simulation assets, and external datasets have their own license requirements.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages