Skip to content

Latest commit

 

History

History
340 lines (262 loc) · 15.6 KB

File metadata and controls

340 lines (262 loc) · 15.6 KB

Contributing to fox

Thank you for your interest in contributing! This guide covers how to build the project, run tests, understand the architecture, and open pull requests.

Building and running tests

# Clone with submodules (llama.cpp is a Git submodule — REQUIRED)
git clone --recurse-submodules https://github.com/ferrumox/fox
cd fox

# Standard build — GPU backend (CUDA / ROCm / Vulkan / Metal) detected at runtime.
# build.rs compiles llama.cpp via CMake, so this needs the submodule + a C/C++ toolchain.
cargo build --release

# Stub build — sets cfg(fox_stub) and swaps in a no-op model, so the crate builds and
# tests run without llama.cpp, the submodule, or GPU drivers. CI runs entirely like this.
FOX_SKIP_LLAMA=1 cargo build

# Fast type-check (no codegen)
cargo check

FOX_SKIP_LLAMA=1 is the key escape hatch for any work that doesn't touch real inference.

GPU backends

build.rs picks the GPU backend at build time from the toolchains it finds on the host (CUDA via nvcc, ROCm via hipcc, Vulkan via glslc, Metal on macOS), and the backend is then dlopen-ed at runtime — one binary runs on GPU or CPU, falling back to CPU automatically when no GPU is present.

Key split: building a GPU backend needs its toolchain; running it only needs the GPU driver. So you can build in a container and run the binary on the host.

Vulkan (AMD / Intel iGPUs and any Vulkan GPU — validated on an AMD Radeon 890M gfx1150). Building the Vulkan backend needs, on top of the standard cmake clang libclang-dev ninja-build:

glslc  glslang-tools  libvulkan-dev  spirv-headers

(llama.cpp's ggml-vulkan CMake wants both the Vulkan/glslc tooling and SPIRV-Headers.) These are all packaged on Ubuntu 24.04; 22.04 does not ship glslc easily. build.rs checks for all three before switching the backend on, and builds CPU-only with an explanatory warning when any is missing — enabling GGML_VULKAN without them is a fatal CMake error, not a fallback. Two escape hatches:

Variable Effect
FOX_NO_VULKAN=1 never build the Vulkan backend, however complete the toolchain looks
FOX_FORCE_VULKAN=1 build it even if the check says a piece is missing

Note that installing a package does not make cargo re-run build.rs, so after apt install spirv-headers a plain rebuild reuses the previous answer. Rebuilding with FOX_FORCE_VULKAN=1 changes the environment and re-runs the check.

On Windows the same rule applies to the LunarG SDK: 1.3.246 and older ship glslc and the loader but no SPIRV-HeadersConfig.cmake, so they need a newer SDK. The reproducible way to build + run is Dockerfile.vulkan:

# Build the image (installs the toolchain above and compiles the Vulkan backend)
docker build -f Dockerfile.vulkan -t fox:vulkan .

# Run with the host GPU passed through (the image ships the Mesa driver)
docker run --rm --device /dev/dri --group-add video \
  -p 8080:8080 -v ~/.cache/ferrumox/models:/root/.cache/ferrumox/models \
  fox:vulkan serve

# …or extract the self-contained bundle and run it natively on a host that already
# has a Vulkan driver (Mesa/RADV, etc.):
id=$(docker create fox:vulkan) \
  && docker cp "$id:/usr/local/lib/fox" ./fox-vulkan && docker rm "$id" \
  && ./fox-vulkan/fox serve --model-path <model.gguf>

make vulkan wraps the build-and-extract step: it produces ./fox-vulkan/ ready to run (./fox-vulkan/fox serve --model-path <model.gguf>).

Code style and CI

make ci runs exactly what CI runs — do this before pushing:

FOX_SKIP_LLAMA=1 cargo fmt --all -- --check
FOX_SKIP_LLAMA=1 cargo clippy --all-targets --features test-helpers -- -D warnings
FOX_SKIP_LLAMA=1 cargo test --all --features test-helpers
  • Format: run cargo fmt before committing — CI rejects unformatted code.
  • Lints: the project must compile clean under clippy -- -D warnings.
  • make setup installs a pre-push git hook that runs these automatically.

Performance claims

The pre-push hook also rejects a commit whose message states a throughput or latency number without saying how it was measured. This is not pedantry: on 2026-08-02 five separate performance claims had to be retracted, none of them because fox's code was wrong — the measurements were (two servers running at once, before/after pairs taken hours apart, a rebuild that silently never applied, a micro-benchmark on synthetic data, a "winner" declared from overlapping ranges). The full account is in docs/design/rocm-benchmarking-2026-08.md.

Use scripts/ab_bench.sh for any before/after comparison — it runs one server at a time, alternates the two arms, reports what each actually loaded, and returns INCONCLUSIVE rather than a winner when the ranges overlap:

./scripts/ab_bench.sh \
  --a-label before --a-cmd './target/release/fox serve --model-path M --port 8097' \
  --b-label after  --b-cmd './target/release/fox serve --model-path M --port 8097' \
  --prep-b 'cargo build --release' \
  --url http://localhost:8097 --model llama-3.2-1b-instruct-q8_0 --rounds 3

Then quote its output, or name whatever harness/instrumentation you used. If the number is illustrative rather than a claim, FOX_SKIP_PERF_CHECK=1 git push.

Architecture overview

Understanding how the pieces fit together makes it much easier to know where to make changes.

CLI (src/cli/)
    │
    ├─ fox serve  →  builds AppState (ModelRegistry + config)
    │               starts Axum HTTP server
    │
    └─ fox run    →  loads one model directly, runs inference loop

API layer (src/api/)
    ├─ router.rs       Axum router + AppState; routes.rs lists the routes
    ├─ v1/             OpenAI-compatible handlers (chat, completions, embeddings, models)
    ├─ ollama/         Ollama-compatible handlers (chat, generate, embed, management)
    ├─ shared/         Inference + streaming helpers reused by both API families
    ├─ types/          Request/response types split by surface (v1, ollama, embeddings, …)
    ├─ auth.rs         Optional FOX_API_KEY bearer-token middleware
    └─ pull_handler.rs POST /api/pull — SSE download from HuggingFace

Model registry (src/model_registry/)
    ModelRegistry — DashMap<String, EngineEntry> + LRU eviction
    get_or_load()  — loads a model on first request, starts its engine loop
    EngineEntry    — holds Arc<InferenceEngine>, aborted on Drop (eviction)
    loader.rs      — resolves model names/aliases to GGUF files

Engine (src/engine/)
    InferenceEngine        — receives InferenceRequest, runs the decode loop
    model/llama_cpp/       — wraps the llama.cpp C bindings (FFI in engine/ffi.rs)
    model/stub.rs          — the cfg(fox_stub) no-op model for FOX_SKIP_LLAMA builds

Scheduler (src/scheduler/)
    Scheduler       — priority queue, assigns KV cache blocks to requests
    prefix_cache.rs — reuses already-processed shared prefixes (continuous batching)

KV cache (src/kv_cache/)
    KVCacheManager  — manages free/used blocks, implements PagedAttention-style allocation

Request lifecycle (HTTP → token stream):

  1. POST /v1/chat/completions arrives at the router (src/api/router.rs) → handler v1/chat.rs
  2. Handler resolves model name → calls registry.get_or_load(model_name)
  3. Registry returns Arc<InferenceEngine> (loading model if not cached)
  4. Handler formats prompt, builds InferenceRequest, submits to engine
  5. Engine's run loop picks up the request, allocates KV blocks via scheduler
  6. Tokens are sent back through an mpsc token channel
  7. Handler streams tokens as SSE (OpenAI) or NDJSON (Ollama) to the client
  8. When the client disconnects, send().is_err() signals the engine to preempt

Adding a new API endpoint

  1. Add the request/response types under src/api/types/
  2. Write the handler in the matching src/api/v1/ (OpenAI) or src/api/ollama/ module, reusing src/api/shared/ helpers
  3. Register the route in src/api/router.rs / src/api/routes.rs
  4. Add documentation under docs/api/

Adding a new CLI command

  1. Create src/cli/<command>.rs with a <Command>Args struct (clap Parser) and a run_<command>() async function
  2. Add the variant to the Command enum in src/cli/mod.rs and wire its match arm in run()
  3. Add documentation to docs/cli/<command>.md
  4. Place the page under docs/ (readable directly on GitHub)

Integration tests

Integration tests live in tests/. The HTTP-layer tests need the test-helpers feature (which exposes StubModel, EngineEntry::for_test, ModelRegistry::preload_for_test, and src/api/test_helpers.rs); the live tests in tests/integration.rs need a running server:

# Full suite as CI runs it (stub model, no server needed)
FOX_SKIP_LLAMA=1 cargo test --all --features test-helpers

# Live integration tests — start a server first (in a separate terminal)
fox serve --model-path models/<model>.gguf --port 8081
cargo test --test integration -- --test-threads=1

Unit tests inside src/ (marked #[cfg(test)]) can run without a server:

cargo test --lib

Adding models to the registry

The built-in model registry lives in registry.json at the repo root. Each entry maps a short alias to a HuggingFace repo and the preferred GGUF filename pattern:

{
  "llama3.2": {
    "repo": "bartowski/Llama-3.2-3B-Instruct-GGUF",
    "pattern": "Q4_K_M"
  }
}

To add a new model:

  1. Add an entry to registry.json following the format above.
  2. Verify fox models displays it and fox pull <alias> resolves correctly.
  3. Include the change in your PR with a brief description of the model.

Design decisions

  • Rust — memory safety without GC, zero-cost async, small binary. The inference hot path runs with minimal allocations.
  • llama.cpp as a submodule — avoids a separate install step and guarantees version parity between the Rust bindings and the C library. This is why --recurse-submodules is required when cloning.
  • DashMap for the model registry — lock-free concurrent reads; models are loaded infrequently but looked up on every request.
  • PagedAttention-style KV cache — prevents memory fragmentation under concurrent load. Enables accurate eviction without OOM.
  • mpsc channel for token streaming — decouples the engine loop from the HTTP handler. The handler detects client disconnects via send().is_err() without polling.

Cutting a release

Decide the bump before writing the entry: COMPATIBILITY.md says which surfaces are a promise and which are not, and therefore whether a change is a patch or a minor. A Tier 1 change in a patch release is the mistake that policy exists to catch.

The version bump, the CHANGELOG entry and make e2e come first. The tag comes last, immediately before pushing — never earlier.

This is a rule because breaking it costs more than it looks. On 2026-08-03 three tags were created the moment their release commit landed, and every one of them had to be deleted and recreated when a later measurement showed the CHANGELOG claiming something untrue: a default that did not do what the entry said, and an "advantage narrows" line that a proper experiment reversed. Nothing was pushed, so nothing broke — but a tag that moves is not a marker, and main/develop currently sit far behind, which makes tags the only durable record these releases have.

The release: X.Y.Z commit is the marker while work continues. Tag that commit when you push it, not when you write it. If a marker is genuinely needed sooner, use an annotated -rc1 tag: the churn then shows up in history instead of being hidden behind a git tag -d.

Write the CHANGELOG entry first — including what the release does not do and any measured limits — then let the tooling do the rest:

make release VERSION=0.21.0    # checks, bump, `release: X.Y.Z` commit
# merge to develop, snapshot to main (below)
make publish VERSION=0.21.0    # ONE tag, then verifies a Release run started

make release refuses to continue on a dirty tree, from main, when the tag already exists locally or on the remote, when the CHANGELOG has no entry for the version (or an empty one), when make ci or make e2e fail, or when the bumped version does not end up matching in Cargo.toml, Cargo.lock and the README badge. That last check exists because 0.14 through 0.18 shipped with Cargo.toml still reading 0.11.0 — six releases whose binary reported the wrong version.

make publish pushes one tag and then asks the API whether a Release run actually started. Never git push --tags: GitHub fires no workflow when more than three tags arrive in a single push, which is how ten tags were published and nothing was built. release.yml also has workflow_dispatch now, so a missed trigger is re-run rather than fixed by deleting a tag.

How develop and main differ

They are not two views of one history — git merge-base main develop is empty. Treat them as two mechanisms:

  • develop takes ordinary merges, one per feature branch, keeping the merge commit: git merge --no-ff feature/0.X -m "Merge branch 'feature/0.X' into develop".

  • main takes a snapshot per released version, in order, one commit each (release: vX.Y.Z), and the tag lives there.

    Historically main was an orphan branch: its own root commit (470e139 release: v0.1.0) with no ancestor in common with develop, so git merge-base main develop was empty and every snapshot was a single-parent commit built by copying a tree. That is a defect, not a design — it made main unmergeable, hid which develop commit each release actually came from, and is why an earlier version of this section said merging would "splice unrelated histories".

    Fixed in 0.22.1 by grafting rather than rewriting: the snapshot is now a two-parent commit. The first parent is main's tip, so git log --first-parent main still shows nothing but the clean release chain; the second is the release: X.Y.Z commit on the working branch, which is what gives the two branches a common ancestor. Every published tag keeps its SHA — nothing was rewritten.

    REL=$(git rev-parse release/X.Y.Z)          # the `release: X.Y.Z` commit
    git checkout main
    NEW=$(git commit-tree "$REL^{tree}" -p main -p "$REL" -m "release: vX.Y.Z")
    git reset --hard "$NEW"

    The tree still comes wholesale from the release commit, so main continues to hold release snapshots and nothing else. Verify both properties before tagging:

    test "$(git rev-parse main^{tree})" = "$(git rev-parse $REL^{tree})"   # same tree
    git merge-base main develop                                            # must not be empty

Version by version, tag by tag. Never one snapshot covering several releases: main is the only record of what each version actually contained, and collapsing two of them throws that away.

Pull request process

develop is the active trunk; main is the release branch. Branch from and target develop.

  1. Fork the repository and create a feature branch from develop:
    git checkout develop
    git checkout -b feat/my-feature
  2. Make your changes. Keep commits focused — one logical change per commit.
  3. Run make ci locally (fmt + clippy + tests) before pushing.
  4. Open a PR targeting develop. Fill in the PR description explaining what changed and why.
  5. A maintainer will review and may request changes. Once approved, the PR is squash-merged.

Reporting bugs

Open a GitHub issue with:

  • fox version (fox --version)
  • OS and architecture
  • Steps to reproduce
  • Expected vs actual behaviour