One API endpoint for your LLM services.
GPROXY is a self-hosted LLM API gateway written in Rust. Add your upstream accounts, configure model routes in the console, and call different services through the same endpoint and gateway API key. Provider changes, credential updates, usage, and costs are managed in one place.
v4 runs as a desktop or mobile application with a setup wizard, a CLI service, a container, or on Cloudflare, Netlify, Vercel, and Deno. Use the latest stable release for regular deployments; development snapshots are available under nightly.
|
Accepts OpenAI Chat Completions, Responses, Claude Messages, and Gemini GenerateContent, with protocol conversion and streaming. Client guides cover Codex CLI, Claude Code, and other tools. |
Use API keys, OAuth, or Cookies as supported by each channel. Manage refresh, health, and quotas, with configurable failover to other credentials or providers. |
|
Give clients a fixed model name such as |
Configure system text, cache breakpoints, JSON edits, regular expressions, and header rules for each provider to accommodate different clients. |
|
Manage users, organizations, teams, API keys, permissions, and spending budgets. Inspect usage, upstream and downstream request logs, and audit records in the console. |
Use the application on a personal device, run the CLI or container on a server, or deploy to Cloudflare, Netlify, Vercel, and Deno. Rust projects can embed the gateway through its SDK. |
Built-in channels include OpenAI, Claude API / Code / Web, Codex, Gemini CLI, Google AI Studio, Copilot, DeepSeek, Kimi, OpenRouter, AWS Bedrock, Azure, Vertex, and xAI. Other services that speak OpenAI, Claude, or Gemini protocols can use the custom channel.
Support for images, audio, files, WebSocket / Realtime, and other operations depends on the channel and upstream. See Providers and credentials and the client guides.
Screenshots of the v4 console, using a separate demo instance with example data.
Providers and models — Manage upstream accounts and model catalogs in one place.
View model routes and request rules
Model routes — Configure providers, weights, and fallback tiers for a fixed model name.
Request rules — Configure system text and cache breakpoints for a provider.
- Download Application (
gproxy-tauri-*) for your platform from Releases. - Open it and follow the wizard to configure the instance, administrator account, and optional configuration import.
- Save the service URL and gateway API key. Open the console and add a provider and upstream credential.
- Test the credential, create a model route named
fast, and add the working model as a member.
Application uses its in-app console; its HTTP port serves gateway API calls.
Download CLI (gproxy-*), extract it, and run:
chmod +x ./gproxy
./gproxy serve --data-dir ./data --port 8787On Windows, use gproxy.exe. Save the administrator password and API key printed on first startup, then open http://127.0.0.1:8787/console/. Add a provider, test its credential, and create a model route as above.
The CLI listens on localhost by default, with SQLite at ./data/gproxy.db. Configure and save GPROXY_MASTER_KEY before adding credentials; without it, upstream secrets are stored unencrypted. Keep the data directory and master key when updating. Configure the listening address and HTTPS for remote access. See the configuration reference.
Replace the key below with your gateway API key. fast is the route created above:
export GPROXY_KEY='your-gproxy-api-key'
curl -sS http://127.0.0.1:8787/v1/chat/completions \
-H "Authorization: Bearer $GPROXY_KEY" \
-H 'Content-Type: application/json' \
-d '{"model":"fast","messages":[{"role":"user","content":"Hello"}],"stream":true}'Templates for Cloudflare, Netlify, Vercel, and Deno use prebuilt v4.0.3 bundles without compiling Rust. Cloudflare supports D1 or libSQL/Turso; the other three use PostgreSQL. Configure the database, administrator password, and master key, then deploy and sign in at /console/.
See the hosted deployment guide for deploy buttons, platform or external databases, and differences such as WebSocket support.
| Option | Use case | Guide |
|---|---|---|
| Application | Desktop or mobile, managed in the app | Platform installation |
| CLI / container | Persistent server with a browser console | Installation and containers |
| Cloudflare / Netlify / Vercel / Deno | Hosted service without your own server | Deployment guide |
| Rust SDK | Embed in your own program | gproxy-sdk |
Windows MSIX files are unsigned Store submission packages; use ZIP for a regular installation. macOS apps are not notarized. The HarmonyOS HAP is experimental and requires your own signature. See the installation guide for platform requirements.
Configuration is loaded into in-memory snapshots, and outbound clients are reused by connection configuration. Log capture can be switched off, with separate settings for body logging.
These measurements use the official v4.0.0 Linux x86_64 release on a Ryzen 7 8745H (8 cores, 16 threads) with SQLite on NVMe. The load generator, gateway, and mock upstream share one machine. Non-streaming Chat Completions requests go through authentication, model routing, usage extraction, pricing, and persistence. The mock reports 25 input tokens and 18 output tokens with no artificial delay.
Each round warms up each logging configuration for 2 seconds, then measures each connection count for 10 seconds with oha 1.16.0. There are three rounds. Each metric below is the median of those three runs. Body logging is off; usage recording and cost settlement are enabled in both logging configurations.
| Request logs | Connections | Requests/s | p50 | p99 |
|---|---|---|---|---|
| Off | 1 | 2,882 | 0.34 ms | 0.50 ms |
| Off | 32 | 28,303 | 0.76 ms | 2.71 ms |
| Off | 64 | 27,263 | 1.47 ms | 41.91 ms |
| Upstream and downstream metadata | 64 | 7,765 | 2.81 ms | 45.16 ms |
With one connection, the gateway adds about 0.30 ms to median latency compared with calling the mock directly. All 1,989,395 measured requests returned HTTP 200. After writes completed, persisted usage counts matched request counts, and every row passed token and cost checks. Input and output were each priced at $1 per million tokens, producing a recorded cost of $0.000043 per request.
At 64 connections, throughput stops increasing and p99 rises to about 42 ms, or 45 ms with metadata logging. This test covers short responses, same-protocol forwarding, and local SQLite. Long streams, protocol conversion, remote databases, and real upstreams need measurements with the intended workload.
Stop v3, back up the database, master key, and startup configuration, then start v4 with the same configuration. Supported v3 SQLite, PostgreSQL, MySQL, and D1 databases migrate automatically, preserving accounts, passwords, API keys, and historical usage. The migration report lists configuration that could not be mapped. Do not run v3 and v4 against the same database at once.
See Migrating v3 to v4 for migration scope, backups, and failure handling.
Native builds require stable Rust, Go, and Clang. The console and docs require Node.js 22.12+ (24 LTS recommended) and pnpm. Linux desktop builds also require the WebKitGTK 4.1, GTK 3, and libsoup 3 development packages.
pnpm --dir console install --frozen-lockfile
pnpm --dir console build
cargo run -p gproxy -- serveThe console build synchronizes embedded assets. A fresh checkout built with only cargo build has no console bundle.
cargo fmt --all --check
cargo clippy --all-targets -- -D warnings
cargo test
pnpm --dir console lint
pnpm --dir console test
pnpm --dir docs install --frozen-lockfile
pnpm --dir docs check
pnpm --dir docs buildSee Building from source for desktop and WASM checks, Architecture for the workspace layout, and Adding a channel for extensions.
Report bugs through Issues. Report vulnerabilities privately through Security.
The gateway application is AGPL-3.0-or-later; see LICENSE. gproxy-protocol, gproxy-protocol-macros, gproxy-client, gproxy-cache, gproxy-file, gproxy-seaorm, and gproxy-tokenizer are MIT, with a license in each directory.


