Skip to content
#

ai-benchmark

Here are 158 public repositories matching this topic...

aicharts

aicharts plots published AI benchmark scores against cost and tokens per task, and a local collector measures your own agents' token use.

  • Updated Oct 7, 2026
  • TypeScript
benchmark-radar

Track 20,710+ AI benchmark, eval, dataset, and data-quality records from 39 public sources, with linked evidence and daily updates.

  • Updated Oct 7, 2026
  • Python

Фрактальный литературно-документальный корпус и AI-бенчмарк для длинноконтекстных мультимодальных рассуждений. Двуязычный (RU/EN). Включает RAG-движок SuperCrichton (FAISS). CC-BY 4.0.

  • Updated Oct 6, 2026
  • HTML

24 个普通人的 46 件人生大事,实测 12 款开源命理工具:八字 BaZi、紫微斗数 Zi Wei Dou Shu、印度占星 Vedic Astrology、奇门遁甲 Qi Men、六爻 Liu Yao、塔罗 Tarot。工具组第一名没超过「四句好话」对照,原始回答与评分可离线复核。

  • Updated Oct 3, 2026
  • Python

MindTrial: Evaluate and compare AI language models (LLMs) on text-based tasks with optional file/image attachments and tool use. Supports multiple providers (OpenAI, Google, Anthropic, DeepSeek, Mistral AI, xAI, Alibaba, Moonshot AI, OpenRouter), custom tasks in YAML, and HTML/CSV/JSON reports.

  • Updated Oct 7, 2026
  • Go

Add this topic to your repo

To associate your repository with the ai-benchmark topic, visit your repo's landing page and select "manage topics."

Learn more