Symbolic Reasoning & Hermeneutic Benchmark for AI Agents
-
Updated
Oct 7, 2026 - Python
Symbolic Reasoning & Hermeneutic Benchmark for AI Agents
Can AI agents finish real engineering work? 17 long, held-out coding tasks graded by hidden tests.
aicharts plots published AI benchmark scores against cost and tokens per task, and a local collector measures your own agents' token use.
Track 20,710+ AI benchmark, eval, dataset, and data-quality records from 39 public sources, with linked evidence and daily updates.
같은 프롬프트로 다섯 개의 AI 모델이 만든 브라우저 전용 3D 수족관을 나란히 실행하고 비교하는 벤치마크 갤러리
📊 Daily auto-updated snapshots of all Arena AI (LMSYS Chatbot Arena) leaderboards — LLM, Vision, Code, Video, Image & more. Structured JSON with historical tracking.
A curated list of evaluation tools, benchmark datasets, leaderboards, frameworks, and resources for assessing model performance.
Фрактальный литературно-документальный корпус и AI-бенчмарк для длинноконтекстных мультимодальных рассуждений. Двуязычный (RU/EN). Включает RAG-движок SuperCrichton (FAISS). CC-BY 4.0.
PlayBench is a platform that evaluates AI models by having them compete in various games and creative tasks. Unlike traditional benchmarks that focus on text generation quality or factual knowledge, PlayBench tests models on skills like strategic thinking, pattern recognition, and creative problem-solving.
🔍 Evaluate AI models' ability to detect ambiguity and manage uncertainty with the ERR-EVAL benchmark for reliable epistemic reasoning.
Free Windows PC benchmark with 63 scored results across CPU, GPU, memory, storage, AI compute and real-world workloads, plus local reports and optional online comparisons.
Benchmark comparing LLM models on one-shot multi-page Flask movie quizzes, with blind reviews, browser evidence, scorecards, and results.
Daily AI model performance comparison across code generation, translation, and long-context summarization. Updated automatically at 03:00 UTC+8.
24 个普通人的 46 件人生大事,实测 12 款开源命理工具:八字 BaZi、紫微斗数 Zi Wei Dou Shu、印度占星 Vedic Astrology、奇门遁甲 Qi Men、六爻 Liu Yao、塔罗 Tarot。工具组第一名没超过「四句好话」对照,原始回答与评分可离线复核。
Research Software Engineering para Ciencias da Terra: pipeline climatico/geoespacial clima x incendio (ERA5-Land/FIRMS-shaped), QC cientifico, proveniencia auditavel e benchmark auto-avaliavel para agentes de IA.
Benchmark for AI-generated parametric CAD, scored on whether the part can actually be printed
Same task, different models and harnesses. Inspect the work.
MindTrial: Evaluate and compare AI language models (LLMs) on text-based tasks with optional file/image attachments and tool use. Supports multiple providers (OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, OpenRouter), custom tasks in YAML, and HTML/CSV/JSON reports.
Race the same agentic coding task across Gemini & Claude on Vertex AI and compare correctness, speed, tokens & cost — live. Node/Express app, deploys to Google Cloud Run with keyless CI.
To associate your repository with the ai-benchmark topic, visit your repo's landing page and select "manage topics."