An in-the-wild benchmark for AI agents in the production harness.
-
Updated
Sep 18, 2026 - Python
An in-the-wild benchmark for AI agents in the production harness.
Gen AI Evaluation Toolkit on AWS is a flexible, cloud-native accelerator built on AWS serverless architecture that enables comprehensive evaluation of generative AI applications.
DataClawEval: A Benchmark for Engineering Data Agents in Real Industrial Harness
A SnitchBench-style benchmark inverted for the dark-forest problem: does a listener AI alert humans about an alien signal when alerting may doom humanity?
🚀 AI Evolution Factory - From evaluation tool to continuous AI self-improvement platform. Agentic evaluation, auto-finetuning, global P2P testing, and hardware telemetry for local LLMs
Open-source verification for AI systems. Declare expected behavior, check real runs, and keep the evidence behind each result
A small benchmark for agent skills, verification artifacts, and fresh-session resumability.
Autonomous Multi-Turn Agentic Metamorphic Fuzzer & Swarm Consensus Stress-Testing Engine. Formulates Martingale context drift estimators, metamorphic execution DAGs, and Sybil groupthink inoculation filters.
Proof-bound evaluator stress testing with oracle-witnessed reward-hacking exploits and replayable evidence.
Paired rollouts for group-relative RL of LLM agents in stochastic environments
Two-tier honest-evaluation harness for agentic RTL design (research, WIP)
Runnable lab measuring invalidation/staleness in agent memory — paper: Are We Ready For An Agent-Native Memory System? (arXiv:2606.24775)
Pipeline to investigate structured reasoning and instruction adherence in multimodal LLMs
The specification behind ESAC, a program of compact benchmarks that measure difficult capabilities without measuring budget. ESAC-GI is released and runnable; ESAC-AG has a harness and one of its seven categories. Items are generators rather than stored questions, and every score carries the version tag of the instrument that produced it.
Scorekeeper is a server that benchmarks AI agents and assistant platforms
Deterministic synthetic fixtures and hard-gate scoring for agentic biosafety safeguard routing.
A Harbor-format RL environment for pass@k estimation, with a verifier built to resist reward hacking. 17 tests, 10 adversarial solutions rejected, amd64 CI.
Measure AI agents’ performance with standardized tests across 314 tasks, 33 domains, and 4 difficulty levels for clear, reproducible comparison.
A GPU-backed Harbor/Terminal-Bench RL environment for data-poisoning defense: containerized task, oracle solution, and a 25-check verifier that rejects 17 reward-hacking solutions.
To associate your repository with the agentic-evaluation topic, visit your repo's landing page and select "manage topics."