Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
-
Updated
Oct 3, 2026 - Python
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
🔍 Minimal examples of machine learning tests for implementation, behaviour, and performance.
Community registry of LLM serving-path traps that produce confidently wrong measurements: templates, tool parsers, reasoning fields, quant kernel paths, CUDA toolchains, KV allocation, eval harnesses, versioning. Symptom-first, with the check that catches each.
Evaluate visual models on your own images, JSON Schema, and production constraints with LangGraph, Pareto analysis, and a no-key Replay demo.
Measure and visualize machine learning model performance without the usual boilerplate.
A High-level Scorecard Modeling API | 评分卡建模尽在于此
A 124.6M LLM trained from scratch on 13B tokens on a single RTX 4090 — tokenizer, pretraining, post-training, evaluation, and reproducible inference.
Ollama Model Test - Figure out the best model for the task
PolyCouncil is an open-source multi-model deliberation engine for LM Studio. It runs multiple LLMs in parallel, gathers their answers, scores each response using a shared rubric, and produces a final, consensus-driven result. Designed for testing, comparing, and orchestrating local models with ease.
Valor is a lightweight, numpy-based library designed for fast and seamless evaluation of machine learning models.
Evaluate the performance of computer vision models and prompts for zero-shot models (Grounding DINO, CLIP, BLIP, DINOv2, ImageBind, models hosted on Roboflow)
🎓 2020 Undergraduate Graduation Project in Jiangnan University ALL codes including Data-convert, keras-Train, model-Evaluate and Web-App
Python tools for climate and air quality model evaluation
Signal diagnostics, statistical validation, and backtest evaluation for quantitative trading workflows.
Local-first macOS app for evaluating and choosing AI coding model configurations.
Detect and classify fraudulent transactions using SQL and Python. Generate behavioral features with SQLite, train a Logistic Regression model, and evaluate performance with AUC, precision, recall, and ROC analysis. A complete supervised fraud detection workflow.
A hands-on TensorFlow image recognition project teaching a computer to identify 10 everyday objects, originally for a linear algebra class, with tools to train a CNN, auto-tune settings, and test accuracy on random internet images.
Rapid Evaluation Framework for climate data
A comprehensive resource for machine learning interview preparation, featuring coding challenges, algorithm explanations, and practical Python examples. Covers supervised and unsupervised learning, model evaluation, and data preprocessing for technical interviews.
Catches fake hi-fi: music sold as lossless that was secretly compressed. It only says so when the proof is overwhelming.
To associate your repository with the model-evaluation topic, visit your repo's landing page and select "manage topics."