A lightweight, general-purpose harness for tool-using LLM agents, fair benchmark evaluation, harness baselines, and personal assistant workflows.
-
Updated
Jul 16, 2026 - Python
A lightweight, general-purpose harness for tool-using LLM agents, fair benchmark evaluation, harness baselines, and personal assistant workflows.
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
Runtime AI Scientist (formerly RAC AI Scientist): official code for "Can AI Scientists Coordinate at Runtime?" (arXiv 2610.00980). Runtime Agent Coordination (RAC) for multi-agent AI scientists such as ARK, Agent Laboratory, and EvoScientist, evaluated on ResearchClawBench.
To associate your repository with the researchclawbench topic, visit your repo's landing page and select "manage topics."