Thank you for your interest in contributing! This guide covers how to build the project, run tests, understand the architecture, and open pull requests.
# Clone with submodules (llama.cpp is a Git submodule — REQUIRED)
git clone --recurse-submodules https://github.com/ferrumox/fox
cd fox
# Standard build — GPU backend (CUDA / ROCm / Vulkan / Metal) detected at runtime.
# build.rs compiles llama.cpp via CMake, so this needs the submodule + a C/C++ toolchain.
cargo build --release
# Stub build — sets cfg(fox_stub) and swaps in a no-op model, so the crate builds and
# tests run without llama.cpp, the submodule, or GPU drivers. CI runs entirely like this.
FOX_SKIP_LLAMA=1 cargo build
# Fast type-check (no codegen)
cargo checkFOX_SKIP_LLAMA=1 is the key escape hatch for any work that doesn't touch real inference.
build.rs picks the GPU backend at build time from the toolchains it finds on the
host (CUDA via nvcc, ROCm via hipcc, Vulkan via glslc, Metal on macOS), and the
backend is then dlopen-ed at runtime — one binary runs on GPU or CPU, falling back to
CPU automatically when no GPU is present.
Key split: building a GPU backend needs its toolchain; running it only needs the GPU driver. So you can build in a container and run the binary on the host.
Vulkan (AMD / Intel iGPUs and any Vulkan GPU — validated on an AMD Radeon 890M
gfx1150). Building the Vulkan backend needs, on top of the standard cmake clang libclang-dev ninja-build:
glslc glslang-tools libvulkan-dev spirv-headers
(llama.cpp's ggml-vulkan CMake wants both the Vulkan/glslc tooling and
SPIRV-Headers.) These are all packaged on Ubuntu 24.04; 22.04 does not ship glslc
easily. build.rs checks for all three before switching the backend on, and builds
CPU-only with an explanatory warning when any is missing — enabling GGML_VULKAN
without them is a fatal CMake error, not a fallback. Two escape hatches:
| Variable | Effect |
|---|---|
FOX_NO_VULKAN=1 |
never build the Vulkan backend, however complete the toolchain looks |
FOX_FORCE_VULKAN=1 |
build it even if the check says a piece is missing |
Note that installing a package does not make cargo re-run build.rs, so after
apt install spirv-headers a plain rebuild reuses the previous answer. Rebuilding with
FOX_FORCE_VULKAN=1 changes the environment and re-runs the check.
On Windows the same rule applies to the LunarG SDK: 1.3.246 and older ship glslc and
the loader but no SPIRV-HeadersConfig.cmake, so they need a newer SDK. The reproducible way to build + run is Dockerfile.vulkan:
# Build the image (installs the toolchain above and compiles the Vulkan backend)
docker build -f Dockerfile.vulkan -t fox:vulkan .
# Run with the host GPU passed through (the image ships the Mesa driver)
docker run --rm --device /dev/dri --group-add video \
-p 8080:8080 -v ~/.cache/ferrumox/models:/root/.cache/ferrumox/models \
fox:vulkan serve
# …or extract the self-contained bundle and run it natively on a host that already
# has a Vulkan driver (Mesa/RADV, etc.):
id=$(docker create fox:vulkan) \
&& docker cp "$id:/usr/local/lib/fox" ./fox-vulkan && docker rm "$id" \
&& ./fox-vulkan/fox serve --model-path <model.gguf>make vulkan wraps the build-and-extract step: it produces ./fox-vulkan/ ready to
run (./fox-vulkan/fox serve --model-path <model.gguf>).
make ci runs exactly what CI runs — do this before pushing:
FOX_SKIP_LLAMA=1 cargo fmt --all -- --check
FOX_SKIP_LLAMA=1 cargo clippy --all-targets --features test-helpers -- -D warnings
FOX_SKIP_LLAMA=1 cargo test --all --features test-helpers- Format: run
cargo fmtbefore committing — CI rejects unformatted code. - Lints: the project must compile clean under
clippy -- -D warnings. make setupinstalls a pre-push git hook that runs these automatically.
The pre-push hook also rejects a commit whose message states a throughput or
latency number without saying how it was measured. This is not pedantry: on
2026-08-02 five separate performance claims had to be retracted, none of them
because fox's code was wrong — the measurements were (two servers running at
once, before/after pairs taken hours apart, a rebuild that silently never
applied, a micro-benchmark on synthetic data, a "winner" declared from
overlapping ranges). The full account is in
docs/design/rocm-benchmarking-2026-08.md.
Use scripts/ab_bench.sh for any before/after comparison — it runs one server
at a time, alternates the two arms, reports what each actually loaded, and
returns INCONCLUSIVE rather than a winner when the ranges overlap:
./scripts/ab_bench.sh \
--a-label before --a-cmd './target/release/fox serve --model-path M --port 8097' \
--b-label after --b-cmd './target/release/fox serve --model-path M --port 8097' \
--prep-b 'cargo build --release' \
--url http://localhost:8097 --model llama-3.2-1b-instruct-q8_0 --rounds 3Then quote its output, or name whatever harness/instrumentation you used. If
the number is illustrative rather than a claim, FOX_SKIP_PERF_CHECK=1 git push.
Understanding how the pieces fit together makes it much easier to know where to make changes.
CLI (src/cli/)
│
├─ fox serve → builds AppState (ModelRegistry + config)
│ starts Axum HTTP server
│
└─ fox run → loads one model directly, runs inference loop
API layer (src/api/)
├─ router.rs Axum router + AppState; routes.rs lists the routes
├─ v1/ OpenAI-compatible handlers (chat, completions, embeddings, models)
├─ ollama/ Ollama-compatible handlers (chat, generate, embed, management)
├─ shared/ Inference + streaming helpers reused by both API families
├─ types/ Request/response types split by surface (v1, ollama, embeddings, …)
├─ auth.rs Optional FOX_API_KEY bearer-token middleware
└─ pull_handler.rs POST /api/pull — SSE download from HuggingFace
Model registry (src/model_registry/)
ModelRegistry — DashMap<String, EngineEntry> + LRU eviction
get_or_load() — loads a model on first request, starts its engine loop
EngineEntry — holds Arc<InferenceEngine>, aborted on Drop (eviction)
loader.rs — resolves model names/aliases to GGUF files
Engine (src/engine/)
InferenceEngine — receives InferenceRequest, runs the decode loop
model/llama_cpp/ — wraps the llama.cpp C bindings (FFI in engine/ffi.rs)
model/stub.rs — the cfg(fox_stub) no-op model for FOX_SKIP_LLAMA builds
Scheduler (src/scheduler/)
Scheduler — priority queue, assigns KV cache blocks to requests
prefix_cache.rs — reuses already-processed shared prefixes (continuous batching)
KV cache (src/kv_cache/)
KVCacheManager — manages free/used blocks, implements PagedAttention-style allocation
Request lifecycle (HTTP → token stream):
POST /v1/chat/completionsarrives at the router (src/api/router.rs) → handlerv1/chat.rs- Handler resolves model name → calls
registry.get_or_load(model_name) - Registry returns
Arc<InferenceEngine>(loading model if not cached) - Handler formats prompt, builds
InferenceRequest, submits to engine - Engine's run loop picks up the request, allocates KV blocks via scheduler
- Tokens are sent back through an
mpsctoken channel - Handler streams tokens as SSE (OpenAI) or NDJSON (Ollama) to the client
- When the client disconnects,
send().is_err()signals the engine to preempt
- Add the request/response types under
src/api/types/ - Write the handler in the matching
src/api/v1/(OpenAI) orsrc/api/ollama/module, reusingsrc/api/shared/helpers - Register the route in
src/api/router.rs/src/api/routes.rs - Add documentation under
docs/api/
- Create
src/cli/<command>.rswith a<Command>Argsstruct (clapParser) and arun_<command>()async function - Add the variant to the
Commandenum insrc/cli/mod.rsand wire its match arm inrun() - Add documentation to
docs/cli/<command>.md - Place the page under
docs/(readable directly on GitHub)
Integration tests live in tests/. The HTTP-layer tests need the test-helpers feature
(which exposes StubModel, EngineEntry::for_test, ModelRegistry::preload_for_test, and
src/api/test_helpers.rs); the live tests in tests/integration.rs need a running server:
# Full suite as CI runs it (stub model, no server needed)
FOX_SKIP_LLAMA=1 cargo test --all --features test-helpers
# Live integration tests — start a server first (in a separate terminal)
fox serve --model-path models/<model>.gguf --port 8081
cargo test --test integration -- --test-threads=1Unit tests inside src/ (marked #[cfg(test)]) can run without a server:
cargo test --libThe built-in model registry lives in registry.json at the repo root. Each entry maps a short
alias to a HuggingFace repo and the preferred GGUF filename pattern:
{
"llama3.2": {
"repo": "bartowski/Llama-3.2-3B-Instruct-GGUF",
"pattern": "Q4_K_M"
}
}To add a new model:
- Add an entry to
registry.jsonfollowing the format above. - Verify
fox modelsdisplays it andfox pull <alias>resolves correctly. - Include the change in your PR with a brief description of the model.
- Rust — memory safety without GC, zero-cost async, small binary. The inference hot path runs with minimal allocations.
- llama.cpp as a submodule — avoids a separate install step and guarantees version parity between the Rust bindings and the C library. This is why
--recurse-submodulesis required when cloning. - DashMap for the model registry — lock-free concurrent reads; models are loaded infrequently but looked up on every request.
- PagedAttention-style KV cache — prevents memory fragmentation under concurrent load. Enables accurate eviction without OOM.
- mpsc channel for token streaming — decouples the engine loop from the HTTP handler. The handler detects client disconnects via
send().is_err()without polling.
Decide the bump before writing the entry: COMPATIBILITY.md says which surfaces are a promise and which are not, and therefore whether a change is a patch or a minor. A Tier 1 change in a patch release is the mistake that policy exists to catch.
The version bump, the CHANGELOG entry and make e2e come first. The tag comes last,
immediately before pushing — never earlier.
This is a rule because breaking it costs more than it looks. On 2026-08-03 three tags
were created the moment their release commit landed, and every one of them had to be
deleted and recreated when a later measurement showed the CHANGELOG claiming something
untrue: a default that did not do what the entry said, and an "advantage narrows" line
that a proper experiment reversed. Nothing was pushed, so nothing broke — but a tag that
moves is not a marker, and main/develop currently sit far behind, which makes tags
the only durable record these releases have.
The release: X.Y.Z commit is the marker while work continues. Tag that commit when you
push it, not when you write it. If a marker is genuinely needed sooner, use an annotated
-rc1 tag: the churn then shows up in history instead of being hidden behind a
git tag -d.
Write the CHANGELOG entry first — including what the release does not do and any measured limits — then let the tooling do the rest:
make release VERSION=0.21.0 # checks, bump, `release: X.Y.Z` commit
# merge to develop, snapshot to main (below)
make publish VERSION=0.21.0 # ONE tag, then verifies a Release run startedmake release refuses to continue on a dirty tree, from main, when the tag already
exists locally or on the remote, when the CHANGELOG has no entry for the version (or an
empty one), when make ci or make e2e fail, or when the bumped version does not end
up matching in Cargo.toml, Cargo.lock and the README badge. That last check exists
because 0.14 through 0.18 shipped with Cargo.toml still reading 0.11.0 — six
releases whose binary reported the wrong version.
make publish pushes one tag and then asks the API whether a Release run actually
started. Never git push --tags: GitHub fires no workflow when more than three tags
arrive in a single push, which is how ten tags were published and nothing was built.
release.yml also has workflow_dispatch now, so a missed trigger is re-run rather
than fixed by deleting a tag.
They are not two views of one history — git merge-base main develop is empty. Treat
them as two mechanisms:
-
developtakes ordinary merges, one per feature branch, keeping the merge commit:git merge --no-ff feature/0.X -m "Merge branch 'feature/0.X' into develop". -
maintakes a snapshot per released version, in order, one commit each (release: vX.Y.Z), and the tag lives there.Historically
mainwas an orphan branch: its own root commit (470e139 release: v0.1.0) with no ancestor in common withdevelop, sogit merge-base main developwas empty and every snapshot was a single-parent commit built by copying a tree. That is a defect, not a design — it mademainunmergeable, hid which develop commit each release actually came from, and is why an earlier version of this section said merging would "splice unrelated histories".Fixed in 0.22.1 by grafting rather than rewriting: the snapshot is now a two-parent commit. The first parent is
main's tip, sogit log --first-parent mainstill shows nothing but the clean release chain; the second is therelease: X.Y.Zcommit on the working branch, which is what gives the two branches a common ancestor. Every published tag keeps its SHA — nothing was rewritten.REL=$(git rev-parse release/X.Y.Z) # the `release: X.Y.Z` commit git checkout main NEW=$(git commit-tree "$REL^{tree}" -p main -p "$REL" -m "release: vX.Y.Z") git reset --hard "$NEW"
The tree still comes wholesale from the release commit, so
maincontinues to hold release snapshots and nothing else. Verify both properties before tagging:test "$(git rev-parse main^{tree})" = "$(git rev-parse $REL^{tree})" # same tree git merge-base main develop # must not be empty
Version by version, tag by tag. Never one snapshot covering several releases: main is
the only record of what each version actually contained, and collapsing two of them
throws that away.
develop is the active trunk; main is the release branch. Branch from and target develop.
- Fork the repository and create a feature branch from
develop:git checkout develop git checkout -b feat/my-feature
- Make your changes. Keep commits focused — one logical change per commit.
- Run
make cilocally (fmt + clippy + tests) before pushing. - Open a PR targeting
develop. Fill in the PR description explaining what changed and why. - A maintainer will review and may request changes. Once approved, the PR is squash-merged.
Open a GitHub issue with:
- fox version (
fox --version) - OS and architecture
- Steps to reproduce
- Expected vs actual behaviour