Single-binary OpenAI-compatible inference proxy: API key GUI, exact usage metering, full request logging, OTel GenAI export
-
Updated
Oct 7, 2026 - Rust
Single-binary OpenAI-compatible inference proxy: API key GUI, exact usage metering, full request logging, OTel GenAI export
Self-hosted multi-model gateway for Claude Code: cut your Claude bill by fanning cheap-model swarms, with session-pinned prompt-cache locality.
Self-hosted AI gateway in Go: pool Claude/ChatGPT/Gemini subscription accounts behind one OpenAI-compatible endpoint, with virtual keys, budgets, built-in token savers and an MCP gateway. A single-binary LiteLLM alternative.
LLM gateway comparison: self-hosted and managed AI API gateways compared on protocol compatibility, billing units, timeout and retry semantics, idempotency, quotas and observability.
Open WebUI starter for connecting to a unified OpenAI-compatible model gateway.
Sub-ms Bun AI gateway with Directive Key routing, 70% reasoning token stripping, zero-Docker key pooling, and a built-in 60s agentic eval gauntlet.,
Local OpenAI-compatible proxy that stitches together multiple provider free-tier quotas, with automatic failover when one runs dry
Self-hosted AI gateway for small teams: a key per person, budgets that fall back to local models, PII masking and free SSO.
Local-first AI gateway with bidirectional protocol conversion (Anthropic ↔ OpenAI ↔ Responses). Route Claude Code / Codex CLI to any provider by quality tag, with failover. Rust + Tauri.
A lightweight, highly secure AI API Gateway/Proxy written in Go. Acts as transparent middleware between local AI coding clients (OpenCode/Pi/Cursor) and upstream LLM providers (Gemini, DeepSeek, Zhipu z.ai).
Command-line tool / Terminal User Interface to install and manage Nenya AI Gateway.
LiteLLM Universal Token Connection Sync
Sandhi — the metering layer for AI agents. An open-source AI usage gateway: meter, attribute, and budget every model call across shared keys.
Lightweight, provider-agnostic Python LLM library — one API for OpenAI, Gemini, Anthropic, Groq, Mistral, Cohere, Azure, Bedrock & Ollama. Vision, streaming, tools, batch.
Thin, memory-bounded OpenAI/Anthropic-compatible LLM reverse proxy in Rust — transparent streaming passthrough with bounded usage/cost/rate-limit observability, opt-out optimizations, and a multi-instance dashboard.
Connect OpenCode to local LLM servers like Ollama, vLLM, and LM Studio using one provider with automatic runtime model detection.
Lightweight reverse proxy for OpenAI API with per-key token budgets, rate limiting, cost dashboards, model downgrade, and caching — deploy in 5 minutes with Docker
The visibility gap LiteLLM doesn't close: attribute LLM spend per RAG pipeline stage (retrieval / reranking / generation / evaluation). OpenAI-compatible, Ollama-backed, <8ms overhead.
Unlock AI Coding with OpenCode Local Provider 2026 - Auto-Detect Ollama LM Studio
To associate your repository with the litellm-alternative topic, visit your repo's landing page and select "manage topics."