Skip to content

Repository files navigation

@aldine/confucius-agent

"Files are state, memory is cache"

Confucius Agent - A feature-complete implementation of the Confucius Code Agent architecture for autonomous AI agent development.

npm version License: MIT

πŸ“š Based on Research

This SDK implements the architecture described in:

Confucius: Iterative Tool Learning from Introspection Feedback by Easy-to-Difficult Curriculum
Shen et al., 2024
arXiv:2512.10398v5
https://arxiv.org/abs/2512.10398

The paper introduces a scalable agent scaffold designed for real-world codebases with:

  • Orchestrator Loop (Algorithm 1) - Core execution cycle
  • Extension System - Pluggable tool architecture
  • Hierarchical Working Memory - Session/Entry/Runnable scopes
  • Sub-Agents - Architect (compression), NoteTaker (sessions), Meta-Agent (learning)

πŸ†• What's New in v2.0.0

🧠 Confucius SDK (NEW)

Full implementation of the Confucius Code Agent paper:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚                      CONFUCIUS AGENT                            β”‚
β”œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”‚
β”‚  β”‚   Session   β”‚    β”‚    Entry    β”‚    β”‚  Runnable   β”‚         β”‚
β”‚  β”‚   Scope     │───▢│   Scope     │───▢│   Scope     β”‚         β”‚
β”‚  β”‚ (immutable) β”‚    β”‚ (task)      β”‚    β”‚ (trace)     β”‚         β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜         β”‚
β”‚         β”‚                                      β”‚                β”‚
β”‚         β–Ό                                      β–Ό                β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚               ORCHESTRATOR (Algorithm 1)                 β”‚   β”‚
β”‚  β”‚  while iteration < max_iters:                           β”‚   β”‚
β”‚  β”‚    1. Invoke LLM with memory                            β”‚   β”‚
β”‚  β”‚    2. Parse actions from response                       β”‚   β”‚
β”‚  β”‚    3. Route to extensions β†’ Execute β†’ Update memory     β”‚   β”‚
β”‚  β”‚    4. Check completion/compression                      β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β”‚         β”‚                                      β”‚                β”‚
β”‚         β–Ό                                      β–Ό                β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”         β”‚
β”‚  β”‚  Architect  β”‚    β”‚  NoteTaker  β”‚    β”‚ Meta-Agent  β”‚         β”‚
β”‚  β”‚ (compress)  β”‚    β”‚ (sessions)  β”‚    β”‚ (learning)  β”‚         β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜         β”‚
β”‚         β”‚                  β”‚                   β”‚                β”‚
β”‚         β–Ό                  β–Ό                   β–Ό                β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   β”‚
β”‚  β”‚              .ralph/ (Persistent Storage)               β”‚   β”‚
β”‚  β”‚  sessions/session-*.md  β”‚  knowledge.md (learned rules) β”‚   β”‚
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Key Features:

  • Hierarchical Memory: Session (system prompt), Entry (task), Runnable (execution trace)
  • Context Compression: Architect agent summarizes runnable scope when tokens exceed threshold
  • Session Notes: NoteTaker generates structured Markdown summaries after each run
  • Self-Improvement: Meta-Agent extracts lessons and injects them into future sessions
  • Built-in Extensions: bash, file_edit, think, finish
  • Multi-Provider LLM: OpenRouter (default), OpenAI, Anthropic

πŸ“¦ What's Included

This package unifies four complementary systems:

🧠 Confucius SDK (NEW in v2.0.0)

Implementation of the Confucius Code Agent paper:

  • Orchestrator Loop: Algorithm 1 from the paper
  • Extension System: Pluggable tools (bash, file_edit, think, finish)
  • Hierarchical Memory: Three-scope architecture (Session/Entry/Runnable)
  • Sub-Agents: Architect, NoteTaker, Meta-Agent
  • Knowledge Base: Persistent learning across sessions

πŸ”„ Ralph Protocol v3

A text-based operating system for LLM agents with:

  • 3-Strike Failure Policy: Auto-reset after 3 failed attempts
  • Token Rot Prevention: Hard cap on state files, automatic archiving
  • Loop Scripts: Bash/PowerShell scripts for continuous agent execution
  • Subagent Spawning: Fresh context generation for hard resets
  • Supervised Recursion: Quality gates with signed trace execution

🌐 Confucius Browser MCP

MCP server for browser automation and testing:

  • Visual QA: Screenshot capture and comparison
  • Accessibility Testing: WCAG contrast auditing
  • Console Monitoring: Error and warning collection
  • Chrome DevTools Protocol: Direct browser control

βœ… AI Agent Focusing System (NEW in v1.0.2)

Built-in verification to prevent typos and linting errors:

  • ESLint Strict Mode: Catches 7+ error types automatically
  • TypeScript Strict Mode: Already enabled, now with explicit verification
  • Verification Commands: npm run verify before commits
  • Zero Tolerance: No any types, no var, explicit return types required

πŸ“‹ Test Prompts & Verification

This package includes comprehensive test prompts to verify functionality:

  • PROMPT_2_RALPH_SANITY: Quick sanity check (5 min)

    • Tests Ralph Protocol CLI commands
    • Verifies state file management
    • Validates strike system
  • PROMPT_1_RALPH_AUDIT: Full accessibility audit (15 min)

    • Ralph + Browser MCP integration
    • Computed WCAG contrast ratios
    • Multi-phase workflow with verification
  • BUNDLE_PLAN: Complete ecosystem architecture

    • Integration guide
    • Package boundaries
    • Relationship to Python framework

These prompts force real computed values and verifiable state changes, catching LLM hallucinations.

πŸš€ Quick Start

Installation

# Global installation (recommended for CLI tools)
npm install -g @aldine/confucius-agent

# Local installation
npm install @aldine/confucius-agent

Confucius SDK - Run an Autonomous Task

# Set your API key (OpenRouter by default)
export OPENROUTER_API_KEY=sk-or-v1-...

# Or pass it directly
confucius run "Create a hello.txt file with 'Hello World'" --api-key sk-or-v1-...

# Use different providers
confucius run "List files in current directory" --provider openai --model gpt-4o
confucius run "Create a test file" --provider anthropic --model claude-3-5-sonnet-20241022

# Verbose mode shows all internal operations
confucius run "Check if README.md exists" --verbose

What happens during a run:

  1. Session Scope initialized with system prompt + learned rules from .ralph/knowledge.md
  2. Entry Scope set with your task
  3. Orchestrator Loop executes until task complete or max iterations
  4. NoteTaker generates session summary β†’ .ralph/sessions/session-*.md
  5. Meta-Agent extracts lesson β†’ appends to .ralph/knowledge.md

Ralph Protocol - Initialize a Project

cd your-project
ralph init --name "My AI Project" --vision "Build something amazing"

This creates the Ralph scaffold:

your-project/
β”œβ”€β”€ IDEA.md          # Problem, User, Outcome
β”œβ”€β”€ PRD.md           # Scope, Non-goals, User Stories, Metrics
β”œβ”€β”€ tasks.md         # Task management
β”œβ”€β”€ progress.txt     # Append-only progress log
β”œβ”€β”€ confucius.md     # State document (<200 lines)
β”œβ”€β”€ PROMPT.md        # Agent operating instructions
└── .ralph/
    β”œβ”€β”€ archive/     # Archived state files
    β”œβ”€β”€ logs/        # Command execution logs
    └── strikes.json # Strike counter state

Browser MCP - Initialize for VS Code Copilot

confucius-browser init --host vscode

Or for Claude Code:

confucius-browser init --host claude

πŸ“– Ralph Protocol Commands

Core Commands

ralph init                    # Initialize Ralph Protocol
ralph status                  # Show status, strikes, token health
ralph context                 # Display full agent context
ralph task "Build API"        # Set the current task
ralph progress "Fixed bug"    # Append to progress.txt

Strike System

ralph strike "Agent looping"  # Record a strike (max 3)
ralph unstrike                # Clear all strikes on success
ralph reset --error "..."     # Hard reset: new run ID, fresh context
ralph history                 # View strike history
ralph check "agent output"    # Check output for failure patterns

3-Strike Rule:

  • Strike 1: Ask for single smallest change
  • Strike 2: Force diagnosis (reproduce β†’ isolate β†’ test)
  • Strike 3: Hard reset - spawn fresh subagent

Loop Execution

ralph loop --powershell       # Generate PowerShell loop script
ralph loop --bash             # Generate Bash loop script
ralph loop --agent "claude"   # Specify agent CLI command

Subagent Spawning

ralph subagent                # Generate minimal context for fresh agent
ralph subagent --error "..."  # Include the failing error
ralph subagent --output ctx.md # Write to file

🌐 Browser MCP Tools

When configured with VS Code Copilot or Claude Code, these MCP tools become available:

open_url

Navigate to URLs with configurable wait conditions.

"Navigate to http://localhost:3000 and wait for the page to load"

screenshot

Capture full page or viewport screenshots.

"Take a screenshot of the login page"

console_errors

Collect console messages, errors, and warnings.

"Check for console errors on the dashboard"

contrast_audit

WCAG contrast checking for accessibility compliance.

"Audit the page for WCAG contrast violations"

πŸ”§ Prerequisites

  • Node.js: 18.18.0+
  • Google Chrome: Latest stable (for Browser MCP)
  • Chrome Remote Debugging: Launch with --remote-debugging-port=9222

πŸ›‘οΈ Security

The Browser MCP implements defense-in-depth security:

  • Localhost-only by default: Only allows http://localhost and http://127.0.0.1
  • Approval tokens: External URLs require explicit approval via CONFUCIUS_APPROVAL_TOKEN
  • Secrets redaction: Automatically redacts sensitive data from logs
  • Chrome binding: Requires Chrome to bind to 127.0.0.1 only

See SECURITY.md for the full security policy.

πŸ—οΈ Architecture

@aldine/confucius-agent/
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ index.ts              # Main entry point
β”‚   β”œβ”€β”€ ralph/                # Ralph Protocol v3
β”‚   β”‚   └── index.ts          # CLI & core logic
β”‚   └── browser/              # Browser MCP Server
β”‚       β”œβ”€β”€ index.ts          # CLI entry point
β”‚       β”œβ”€β”€ public.ts         # Public API exports
β”‚       β”œβ”€β”€ mcp/              # MCP protocol implementation
β”‚       β”‚   β”œβ”€β”€ server.ts     # Stdio server
β”‚       β”‚   └── logging.ts    # Structured logging
β”‚       β”œβ”€β”€ runtime/          # Chrome DevTools integration
β”‚       β”‚   β”œβ”€β”€ cdp_client.ts # CDP client
β”‚       β”‚   β”œβ”€β”€ browser_session.ts
β”‚       β”‚   └── allowlist.ts  # URL security
β”‚       └── cli/              # Config writers
β”œβ”€β”€ ralph-loop.ps1            # PowerShell loop script
β”œβ”€β”€ package.json
β”œβ”€β”€ tsconfig.json
β”œβ”€β”€ SECURITY.md
└── README.md

🎯 Use Cases

Autonomous Agent Development

# Initialize project with Ralph scaffold
ralph init --name "AutoCoder" --vision "Self-improving code agent"

# Start development loop
ralph loop --powershell

# Agent works autonomously with automatic strike tracking

Visual QA & Accessibility

# Set up Browser MCP
confucius-browser init --host vscode

# In Copilot/Claude:
"Navigate to localhost:3000 and check for WCAG contrast violations"
"Take screenshots of all form states"
"Check for console errors during checkout flow"

Combined Workflow

Use Ralph Protocol to manage agent state while Browser MCP provides visual verification:

  1. Agent reads PRD.md and tasks.md
  2. Makes code changes
  3. Browser MCP verifies UI changes
  4. Ralph tracks progress and handles failures

οΏ½ Code Quality & AI Agent Focusing

Built-in Verification System

The package includes ESLint + TypeScript strict mode to prevent common AI agent mistakes:

# Run type checking and linting together
npm run verify

# Type check only (strict mode enabled)
npm run typecheck

# Lint only
npm run lint

# Auto-fix linting issues
npm run lint:fix

# Test the focusing system
npm run test:focus       # Basic verification
npm run test:verify      # Comprehensive test

What Gets Caught

  • ❌ Typos in variable/interface names (UserProflie β†’ UserProfile)
  • ❌ Missing return type annotations on functions
  • ❌ var usage instead of const/let
  • ❌ any type usage - enforces proper typing
  • ❌ Unused variables - catches dead code
  • ⚠️ Console statements - warnings

For AI Agent Developers

Add these rules to your AI agent instructions:

**CRITICAL**: Run `npm run verify` before committing any code.
**CRITICAL**: Fix ALL linting errors and type errors. Zero tolerance.
**CRITICAL**: Use explicit return types on all functions.
**CRITICAL**: Never use `any` type - always provide proper types.

See .eslintrc.json and test files for configuration details.

οΏ½πŸ“„ Migration from Previous Packages

If you were using the separate packages:

# Old packages (deprecated)
npm uninstall @aldine/ralph-protocol @aldine/confucius-mcp-browser

# New unified package
npm install -g @aldine/confucius-agent

The CLI commands remain the same:

  • ralph - Ralph Protocol commands
  • confucius-browser - Browser MCP commands

πŸ“ License

MIT License - see LICENSE for details.

πŸ”— Links

πŸ™ Acknowledgments

Research Attribution

The Confucius SDK implementation is based on the architecture described in:

Confucius: Iterative Tool Learning from Introspection Feedback by Easy-to-Difficult Curriculum
Shen et al., 2024
arXiv: 2512.10398v5

Key concepts implemented from the paper:

  • Algorithm 1: Orchestrator execution loop
  • Section 2.2: Extension system architecture
  • Section 2.3.1: Hierarchical working memory (Session/Entry/Runnable scopes)
  • Section 2.3.2: Note-taking agent for session summarization
  • Section 2.3.3: Meta-agent self-improvement loop

Built With

Built with ❀️ for autonomous AI agent development, visual QA, and accessibility testing.


Replaces:

  • @aldine/ralph-protocol (v0.2.0)
  • @aldine/confucius-mcp-browser (v0.1.0)

About

No description, website, or topics provided.

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages