A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.
-
Updated
Sep 21, 2026
A curated list of autonomous improvement loops, research agents, and autoresearch-style systems inspired by Karpathy's autoresearch.
Group Evolving Agents: Open-Ended Self-Improvement via Experience Sharing
Agentic quant research workbench with CSV data inspection, local Codex and compatible APIs, predictive models, independent reviews, and reproducible backtests.
Learning to Evaluate Before Improving: Automatic Rubric Induction for Automatic Research Agents
LitReview Skill is an installable agent skill for end-to-end literature review generation. It helps agents conduct literature reviews with a well-designed and widely used review framework so the search process is broad, iterative, and less likely to miss relevant articles.
SutroYaro — Sutro Group research workspace for energy-efficient AI training. Point any coding agent at the repo and it becomes a research agent. 34 experiments, eval environment, weekly catch-ups, multi-researcher workflow.
🤖 CodeForge AI: An autonomous multi-agent coding system powered by LangGraph for agentic software development and automated workflows. SOTA custom agentic GraphRag, shared-state memory, auto-model routing for cost optimization, and a range of custom tooling.
🔍 DISCOVER — Auditable knowledge graph of claims and contributions from primary literature. See the state of a field, contradictions, and gaps.
Faraday: An Autonomous Web Research Agent (LangGraph/Streamlit). 🕵️♀️ Investigates queries using dynamic tools (Tavily, Google, NewsAPI, etc.), gathers multi-source info, and synthesizes structured reports in a Streamlit UI. Features agentic workflow & source tracking.
Open-source Bittensor subnet for a global AI agent competition in sales intelligence
Curated paper-related AI skills and GitHub repositories for idea discovery, literature search, experiments, writing, citations, LaTeX/DOCX, review, and submission.
📝 REVIEW — 10-stage pipeline turning a paper PDF into claim/evidence map, novelty and rigor checks, and a self-critique pass. Built to help authors improve.
Evaluation code, tasks, prompts, and scientific tools for EarthVerse.
Six MCP servers that automate the full academic research pipeline — from refining a vague research question to generating a publication-ready report. Each server handles a distinct stage of the workflow: question development, data processing, code generation, script execut
Agent-assisted open-source toolkit for detecting and reviewing suspicious regions in reconstructed papyrus surfaces
Faraday: An Autonomous Web Research Agent (LangGraph/Streamlit). 🕵️♀️ Investigates queries using dynamic tools (Tavily, Google, NewsAPI, etc.), gathers multi-source info, and synthesizes structured reports in a Streamlit UI. Features agentic workflow & source tracking.
Research workrooms for AI agents: evidence retrieval, citation audit, adversarial review, peer review gates, and clean final reports.
Open agent network for reproducible research: AI agents test hypotheses through code, falsification, review, and scientific memory.
AgentX — community registry of research AI agents. Public hub for issue reports, agent suggestions and data corrections; the registry itself is served via the website API.
Lightweight Python CLI for the Exa API (Search, Contents, Find Similar, Answer, Research, Context) with JSON-first output, SSE streaming, and model-aware polling. LLM‑agnostic: integrate with OpenAI Agents SDK/Codex CLI or Claude tool use by invoking CLI commands, no MCP server required.
To associate your repository with the research-agents topic, visit your repo's landing page and select "manage topics."