Trace, evaluate, guard, and improve AI agents and coding agents with OpenTelemetry. Agent observability Β· Evals Β· Guardrails Β· Prompt management Β· Cost & GPU monitoring Β· Self-host free (Apache 2.0)
β StarΒ Β Β Β π QuickstartΒ Β Β Β π Docs
Website Β· Documentation Β· Quickstart Β· Compare Β· Join Slack
OpenLIT is an open-source agent harness engineering platform. It gives teams OpenTelemetry-native tracing, evaluations, guardrails, prompt and context management, and cost and GPU monitoring for the harness around their AI agents and coding agents, so every agent failure can be traced, scored, and turned into a harness fix. OpenLIT is free to self-host under Apache 2.0 and works with any model, framework, or harness: Claude Code, Codex, Cursor, OpenAI Agents SDK, LangGraph, CrewAI, and 70+ more integrations.
An AI agent is a model plus a harness. The harness is everything except the model: tools, context, prompts, memory, hooks, guardrails, and feedback loops. Agent harness engineering is the discipline of designing, measuring, and improving that harness so agents are reliable in production. The core loop is run β observe β evaluate β fix the harness β verify, and OpenLIT gives you each step:
| Harness engineering step | OpenLIT feature |
|---|---|
| Observe every LLM call, tool call, MCP request, retrieval and agent step | Agent observability & OpenTelemetry LLM tracing |
| Evaluate quality, safety and cost on real traces | Evals (LLM-as-a-judge, programmatic, human feedback) + CI gates |
| Guard the agent at runtime | Guardrails (prompt injection, sensitive topics, topic restriction) |
| Fix the harness without redeploying | Prompt Hub, Context, Rule Engine, Vault (secrets) |
| Compare models and prompts before shipping | OpenGround |
| Account for cost, latency and hardware | Cost tracking, custom pricing, GPU monitoring (NVIDIA, AMD, Intel) |
A production agent can involve:
flowchart TD
U([User]) --> A[AI Agent]
A --> L[LLM calls]
A --> T[Tool calls]
A --> R[Retrieval]
A --> M[Memory]
A --> S[Sub-agents]
A --> P[Prompts]
A --> C[Code changes]
L & T & R & M & S & P & C --> E{{Evaluation}}
E --> O[["Cost / Quality / Errors"]]
style U fill:#F97316,stroke:#7C2D12,color:#fff
style A fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style E fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style O fill:#F97316,stroke:#7C2D12,color:#fff
git clone https://github.com/openlit/openlit.git
cd openlit
docker compose up -dOpen:
http://127.0.0.1:3000
Python:
pip install openlitTypeScript:
npm install openlitPython:
import openlit
openlit.init()That's it.
OpenLIT automatically instruments supported LLM providers, frameworks, vector databases, and other AI infrastructure and exports OpenTelemetry traces and metrics.
By default, configure the OTLP endpoint:
export OTEL_EXPORTER_OTLP_ENDPOINT="http://127.0.0.1:4318"Or:
import openlit
openlit.init(
otlp_endpoint="http://127.0.0.1:4318"
)Open your dashboard and start exploring your AI application's traces, metrics, costs, and performance.
AI coding agents are powerful β but understanding what they actually did can be difficult.
OpenLIT gives you an OpenTelemetry-native view of coding-agent sessions.
Install the CLI:
curl -fsSL https://raw.githubusercontent.com/openlit/openlit/main/cli/scripts/install.sh | shiwr -useb https://raw.githubusercontent.com/openlit/openlit/main/cli/scripts/install.ps1 | iexConfigure OpenLIT:
openlit configure --endpoint http://127.0.0.1:4318Install coding-agent instrumentation:
openlit coding install --vendor=allOr install individual integrations:
openlit coding install --vendor=cursor
openlit coding install --vendor=claude-code
openlit coding install --vendor=codexCheck your installation:
openlit doctorNow OpenLIT can capture:
flowchart LR
S([Coding Agent Session]) --> P[User prompt]
S --> L[LLM calls]
S --> T[Tool calls]
T --> T1[File reads]
T --> T2[File edits]
T --> T3[Shell commands]
T --> T4[Search]
S --> SA[Sub-agent activity]
S --> TU[Token usage]
S --> CO[Cost]
S --> CI[Code impact]
style S fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style T fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
Explore the resulting sessions in the Coding Agents dashboard.
Understand exactly what happened during an AI request.
All represented using OpenTelemetry.
Track the cost of your AI applications across:
Support custom pricing for custom and fine-tuned models.
Automatically evaluate LLM and agent outputs using LLM-as-a-Judge evaluations.
Built-in evaluation types include:
Use evaluations to move from:
"The agent produced an answer."
to:
"The agent produced a good answer."
Find the requests that matter.
Investigate:
Go from:
Something went wrong.
to a fully traced root cause:
flowchart TD
A[Agent] --> P[Prompt] --> L1[LLM] --> T[Tool call] --> R[Retrieval] --> L2[LLM] --> E([Error])
style E fill:#DC2626,stroke:#7F1D1D,color:#fff
style A fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
Use Prompt Hub to:
Example:
prompt = openlit.prompts.get(
"customer-support"
)Keep prompt management separate from application code while maintaining version control and observability.
Use Context to store reusable RAG content once and have the Rule Engine return the right piece at runtime. See the Context docs.
Define runtime rules based on trace attributes.
Use rules to dynamically control:
Example:
IF
environment = production
AND
model = expensive-model
THEN
run cost evaluation
+ retrieve production prompt
Use OpenLIT SDK guardrails to detect and block risky prompts at runtime:
- Prompt injection / jailbreak attempts
- Sensitive topics
- Topic restriction
Guardrail checks are traced with OpenTelemetry so you can see when and why a request was blocked. See the guardrails docs and quickstart.
Store LLM API keys and other secrets in Vault, then retrieve them at runtime via the SDK or API β without hard-coding credentials in application code. See the Vault docs.
Run the same prompt across multiple providers in OpenGround and compare cost, latency, and output quality before you ship. See the OpenGround docs.
Monitor NVIDIA, AMD, and Intel GPUs used for LLM inference with the OpenTelemetry GPU collector: utilization, memory, power, and temperature, correlated with your traces. See the GPU collector docs.
OpenLIT is built around OpenTelemetry, rather than creating a proprietary telemetry format. Traces and metrics follow OpenTelemetry GenAI semantic conventions (gen_ai.*).
Your telemetry can flow through the OpenTelemetry ecosystem:
flowchart TD
A["AI App / AI Agent"] -->|OTLP| R[OpenLIT OTLP receiver]
A -->|optional sidecar| C[Your OpenTelemetry Collector]
C --> R
C --> O["Other OTel backends<br/>(Datadog, Grafana, Honeycomb, ...)"]
R --> B[ClickHouse]
B --> D[OpenLIT Dashboard]
style A fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style R fill:#F97316,stroke:#7C2D12,color:#fff
style C fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style B fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style D fill:#F97316,stroke:#7C2D12,color:#fff
OpenLIT listens for OTLP on :4317 (gRPC) and :4318 (HTTP). A bundled Collector is not required. You can still put your own Collector in front for fan-out or processing.
This means you can integrate OpenLIT into an existing OpenTelemetry architecture instead of replacing it.
OpenLIT auto-instruments a growing ecosystem of AI providers, frameworks, vector databases, and GPU infrastructure with a single line of code. Click any badge to view its integration guide.
LLM Providers
Vector & Data Stores
AI Frameworks & Agents
Governance & Protocols
GPU Monitoring
See the complete integration list in the documentation.
OpenLIT provides OpenTelemetry-native SDKs for:
pip install openlitnpm install openlitOpenLIT is designed to run in your infrastructure.
A typical deployment looks like:
flowchart TD
subgraph App["Your application"]
direction LR
Agent --> LLM --> Tools --> RAG --> DB
end
App -->|OTLP| Receiver[OpenLIT OTLP receiver]
Receiver --> CH[(ClickHouse)]
CH --> Dash[OpenLIT Dashboard]
style App fill:#1F2937,stroke:#F97316,stroke-width:2px,color:#fff
style Receiver fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style CH fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style Dash fill:#F97316,stroke:#7C2D12,color:#fff
Run OpenLIT inside your own infrastructure using Docker or Kubernetes.
Your telemetry stays under your control.
docker compose up -dFor Kubernetes, see the installation documentation.
Observability is only the beginning.
OpenLIT is designed around a continuous AI engineering loop:
flowchart LR
T[Trace] --> E[Evaluate] --> A[Analyze] --> O[Optimize] --> M[Manage] -.-> T
style T fill:#F97316,stroke:#7C2D12,color:#fff
style E fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style A fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style O fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
style M fill:#111827,stroke:#F97316,stroke-width:2px,color:#fff
The goal is simple:
Make AI systems observable, measurable, debuggable, and continuously improvable.
What is OpenLIT? OpenLIT is an open-source agent harness engineering platform. It provides OpenTelemetry-native tracing, evaluations, guardrails, prompt management, and cost and GPU monitoring for AI agents and coding agents, and it is free to self-host under Apache 2.0.
What is an agent harness? An agent harness is everything in an AI agent except the model: the tools, context, prompts, memory, hooks, guardrails, and feedback loops that turn a model into a working agent. Claude Code, Codex, and frameworks like LangGraph or CrewAI are harnesses.
What is agent harness engineering? Agent harness engineering is the discipline of designing, measuring, and improving the harness around a model so agents are reliable in production. Teams run agents on real tasks, observe failures in traces, evaluate them, fix the harness (prompts, tools, rules, guardrails), and verify the fix with regression evals.
Is OpenLIT an agent framework or a harness runtime? No. OpenLIT doesn't run your agent loop. It instruments and improves whatever harness you already use, through OpenTelemetry, so you can switch models, frameworks, or harnesses without losing your traces, evals, or prompts.
How is OpenLIT different from Langfuse, LangSmith, or Arize Phoenix? OpenLIT is Apache-2.0 and OpenTelemetry-native end to end, adds GPU monitoring, guardrails, Vault, and coding-agent observability in the same self-hosted platform, and exports to any OTLP backend. See https://openlit.io/compare for feature-by-feature comparisons.
Does OpenLIT work with Claude Code, Codex, and Cursor?
Yes. The openlit CLI ingests each coding agent's hook events and maps them to OpenTelemetry gen_ai.* conventions, with no SDK and no code changes.
Can I run agent evals in CI? Yes. Run LLM-as-a-judge and programmatic evals online on production traces, or offline through the SDK as CI/CD gates.
Is OpenLIT free? Yes. Self-hosted OpenLIT is free under Apache 2.0, with no license key and no per-trace fee.
Does OpenLIT add latency? No proxy is required. SDKs instrument in-process and export telemetry asynchronously over OTLP.
Can I send data to Grafana, Datadog, or my existing OpenTelemetry Collector? Yes. OpenLIT emits standard OTLP traces and metrics, so they can go to OpenLIT, Grafana, Datadog, New Relic, SigNoz, or any OTLP backend.
Feature-by-feature comparisons on openlit.io/compare:
- OpenLIT vs Langfuse
- OpenLIT vs LangSmith
- OpenLIT vs Arize Phoenix
- OpenLIT vs Braintrust
- OpenLIT vs Helicone
- OpenLIT vs Datadog
- OpenLIT vs Comet Opik
OpenLIT is open source and built with the AI engineering community.
Join us:
If OpenLIT is useful to you, please consider giving the repository a β.
It helps other AI engineers discover the project.
Contributions are welcome.
You can contribute by:
- fixing bugs
- adding integrations
- improving documentation
- creating examples
- improving SDKs
- adding evaluations
- building dashboards
- reporting issues
- sharing OpenLIT with other developers
Check the repository's issues for opportunities to contribute.
OpenLIT is licensed under the Apache License 2.0.
See LICENSE for details.
Names for the same contributors are on openlit.io/about-us.





