Skip to content

Latest commit

 

History

History
286 lines (231 loc) · 14.9 KB

File metadata and controls

286 lines (231 loc) · 14.9 KB

Inference backends

llmman does not ship an inference engine. llmman serve picks one that already exists for the model format it finds, and runs it unmodified.

Model format Backend Where it comes from
GGUF llama-server in a container --runtime docker / podman (Linux only; what auto tries first): the ghcr.io/ggml-org/llama.cpp:server-<backend> image for your GPU
GGUF llama-server --runtime bin: a prebuilt upstream release matching your OS/arch/GPU, downloaded and cached; --runtime path: the one on your PATH
safetensors vllm in a container --runtime docker / podman (Linux only): the vllm/vllm-openai, rocm/vllm or vllm/vllm-openai-cpu image for your GPU and architecture
safetensors vllm Your PATH (every non-container runtime)
safetensors sglang in a container --runtime docker / podman (Linux only) with LLMMAN_SAFETENSORS_ENGINE=sglang: the lmsysorg/sglang image for your GPU
safetensors sglang Your PATH, with LLMMAN_SAFETENSORS_ENGINE=sglang
safetensors mlx_lm.server macOS: llmman's own uv-installed copy, installed when the daemon starts, or the one on your PATH — per --runtime, as for llama-server; preferred over vllm
GGUF diffusion (LTX-2) llmman itself, on ggml The libggml/libllama next to llama-server; see the blog post
Diffusers safetensors (Qwen-Image 2.1) llmman itself, on ggml The same libraries; see below
Diffusers safetensors vllm serve --omni Your PATH's vllm with the vllm-omni package installed
Diffusers safetensors vllm serve --omni in a container --runtime docker / podman (Linux only): the vllm/vllm-omni image (CUDA only)

Choosing a runtime

llmman serve --runtime <auto|docker|podman|bin|path> (or the LLMMAN_RUNTIME environment variable, which also reaches the daemon llmman run/launch start for you) says where the engine comes from:

Value llama-server Notes
docker, podman ghcr.io/ggml-org/llama.cpp:server-<backend>-<release> container Linux only; safetensors models get the vLLM images too
bin llmman's own download of llama.cpp's prebuilt release for this OS/arch/GPU Installed under ~/.local/share/llmman/llama.cpp/<release>/; nothing needed on PATH
path the llama-server already on PATH, as-is (and, on macOS, the mlx_lm.server on PATH) Never downloads anything; --llama-cpp-version does not apply
auto (default) the first of docker, podman, bin, path that works here Off Linux: bin, then path

Under auto, a container engine "works" when its CLI is on PATH, its daemon answers docker info/podman info, an NVIDIA host has the NVIDIA Container Toolkit, and the llama.cpp image pulls; bin works when the release downloads (or is already cached). Each step skipped is logged with the reason. Whatever is chosen is fetched before the listener binds, so the first request is never stuck behind a silent download. On macOS mlx_lm.server follows the same rule at the same point: bin installs llmman's own uv and mlx-lm, path uses PATH's, auto installs and falls back to PATH (see MLX).

Everything llmman installs for itself — llama.cpp, uv, mlx-lm — has one shape on disk: ~/.local/share/llmman/<tool>/<version>/, a .llmman-complete sentinel written last (without it the directory is rebuilt), one installer at a time across processes (<tool>/.lock), and earlier versions deleted once a new one completes. The pins are the repo-root LLAMA_CPP_RELEASE, UV_RELEASE, MLX_LM_RELEASE and PYTHON_RELEASE files, which CI tests against.

If that fetch fails, the daemon still starts: the failure is a warning, and the first model load that needs llama.cpp retries it and, failing again, reports that to the request. Nothing remembers the failure. A safetensors model under bin or path needs no llama.cpp and loads regardless. mlx-lm is the same: a failed install is a warning, retried by the first safetensors load. Only --pull-only treats a failed fetch as an error.

llmman serve --pull-only does exactly those fetches, in the foreground with their own progress, then exits — run it once before starting a detached daemon. With a container runtime and a safetensors MODEL argument it pulls the vLLM image as well.

llama.cpp

The release is pinned: unless --llama-cpp-version <tag> says otherwise, both bin and the container images use the llama.cpp build llmman's own CI tested (the LLAMA_CPP_RELEASE file in the repository), so what runs is what was tested. --llama-cpp-version latest takes upstream's floating latest instead.

For bin, llmman probes for CUDA, ROCm, OpenCL (Windows ARM64 Adreno only), Vulkan or Metal (in that order) and downloads the matching prebuilt asset from llama.cpp's GitHub releases. LLMMAN_LLM_LIBRARY overrides the probe; LLMMAN_DEBUG=1 shows what it found. llama.cpp publishes no prebuilt Linux CUDA binary, so an NVIDIA host on Linux gets the CPU build from bin — the container runtimes (which auto prefers for that reason) have CUDA images.

Context length, parallel slots, flash attention, KV-cache type and GPU split are environment variables; see configuration.md.

In a container

On Linux, --runtime docker (or podman) runs llama-server from the ghcr.io/ggml-org/llama.cpp image instead, picking the server-cuda/server-cuda13/server-rocm/server-vulkan/server tag for the host, suffixed with the pinned release. CUDA_VISIBLE_DEVICES and friends are forwarded into the container.

The container is a sibling of the daemon, not part of its cgroup, so a CPU limit on llmman serve is forwarded: when one binds (cgroup quota or affinity mask), the container gets --cpus <n>. llama-server always gets --threads <n>: half the CPUs the daemon can use above 8, else all of them (a quota alone would leave it autodetecting a thread per host core and throttling). No limit, no --cpus. An explicit LLAMA_ARG_THREADS still wins. vLLM containers get the same --cpus.

Each of those images is also published with llmman in it, as docker.io/ai/llmman:<tag> (latest is server) and <tag>-<llmman version>, built from packaging/Dockerfile on every release against the llama.cpp build CI tests. Same entrypoint and GPU flags as upstream, plus /usr/local/bin/llmman. LLMMAN_HOST is preset to 0.0.0.0:17434 (a loopback bind inside a container is unreachable even with -p), so the daemon requires LLMMAN_API_KEYS or LLMMAN_AUTH=off (configuration.md); publish the port on the host's loopback to keep it local. The store is /root/.local/share/llmman:

docker run -p 127.0.0.1:17434:17434 -e LLMMAN_API_KEYS=<key> \
  -v llmman:/root/.local/share/llmman --gpus all \
  --entrypoint llmman ai/llmman:server-cuda serve

vLLM

Safetensors models are served by a separately installed vllm unless LLMMAN_SAFETENSORS_ENGINE picks SGLang, or the host is a Mac (see MLX). Plain vllm is CPU-only on macOS unless vllm-metal is installed. LLMMAN_CONTEXT_LENGTH is forwarded as --max-model-len; LLMMAN_LOAD_TIMEOUT (default 10 minutes) bounds a stalled load. LLMMAN_VLLM_ARGS appends anything else (--dtype bfloat16 --tp 2) to the vllm serve command line, local or in a container.

In a container

On Linux, --runtime docker (or podman) runs vllm serve from a vLLM image for a safetensors model, picked by the same GPU probe plus the host architecture:

Host GPU x86_64 aarch64
NVIDIA (CUDA 13) vllm/vllm-openai:latest-x86_64 vllm/vllm-openai:latest-aarch64
NVIDIA (CUDA 12) vllm/vllm-openai:latest-x86_64-cu129 vllm/vllm-openai:latest-aarch64-cu129
AMD (ROCm) rocm/vllm:latest not published upstream (LLMMAN_LLM_LIBRARY=cpu for the CPU image)
Vulkan-only or none vllm/vllm-openai-cpu:latest-x86_64 vllm/vllm-openai-cpu:latest-arm64

--vllm-version <tag> pins the release (the arch suffix is added for the vllm/ images; for rocm/vllm it is the whole tag). llmman serve --runtime docker --pull-only <model> pulls the image an already-pulled model needs (without a model, the llama.cpp image). CUDA_VISIBLE_DEVICES and friends plus every VLLM_* variable are forwarded into the container.

vLLM-Omni (Diffusers pipelines)

A safetensors repository laid out as a Diffusers pipeline (a root model_index.json next to transformer/, vae/, ...) is served by vLLM-Omni: the same vllm launcher with --omni. Plain vllm serve cannot load one. The exception is a pipeline llmman runs itself (see Qwen-Image 2.1).

uv pip install vllm==0.28.0 vllm-omni     # into the environment `vllm` runs from
llmman run ORG/MODEL "A robot arm cleaning a plate in a kitchen"
llmman run ORG/MODEL --video --seconds 2 "A robot arm cleaning a plate"

If vllm is a #!/path/to/python script whose Python cannot import vllm_omni, the load fails up front with a message saying so; otherwise vLLM itself reports unrecognized arguments: --omni. The model answers /v1/images/generations and /v1/videos, in llama-server's dialect (width/height/steps/cfg_scale, streamed image_generation.* events, a video job with a content_url) and vLLM-Omni's own (size, num_inference_steps, guidance_scale, num_frames, extra_params); unsent fields are left to the model's defaults. There is no /v1/audio/speech for these models.

The server is started with --no-guardrails; LLMMAN_VLLM_OMNI_GUARDRAILS=1 leaves the pipeline's safety guardrails on (they may need extra packages and gated weights). LLMMAN_LOAD_TIMEOUT is also passed as --init-timeout.

With a container runtime, the image is vllm/vllm-omni:latest-x86_64 or latest-aarch64 (CUDA only; --vllm-version pins vLLM-Omni's release, e.g. v0.28.0). --pull-only <model> picks it for a pulled Diffusers model.

Qwen-Image 2.1

QwenImage21Pipeline repositories (docker.io/ai/qwen-image-2.1, Qwen/Qwen-Image-2.1) run in llmman on the same ggml libraries, straight from their bf16 safetensors:

llmman run qwen-image-2.1 "A manatee in a sunlit lagoon"

Text to image only: 1024x1024, 40 steps, RGBA PNG. The text encoder and the transformer load on first use and take turns when both do not fit.

vllm serve from llmman's store

The inverse: the vllm-llmman plugin lets vllm serve oci://<reference> pull a CNCF ModelPack image from any OCI registry, via llmman (LLMMAN_BIN if it is not on PATH). See vllm-plugin/README.md.

SGLang

LLMMAN_SAFETENSORS_ENGINE=sglang sends safetensors models to SGLang instead of vllm (uv pip install sglang puts its sglang console script on PATH):

LLMMAN_SAFETENSORS_ENGINE=sglang llmman serve
llmman run Qwen/Qwen3-0.6B "hello"

llmman runs sglang serve --model-path <dir> --served-model-name <ref> on the model's directory; LLMMAN_CONTEXT_LENGTH is forwarded as --context-length and LLMMAN_LOAD_TIMEOUT bounds the load as for vLLM. LLMMAN_SGLANG_ARGS appends anything else (--disable-cuda-graph --tp 2). Which device SGLang uses is its own autodetection and environment, which a local child inherits: SGLANG_USE_CPU_ENGINE=1 for its CPU engine (Linux; x86_64 needs AVX-512), SGLANG_USE_MLX=1 for its MLX runtime on Apple Silicon. llmman ps shows sglang (local).

LLMMAN_SAFETENSORS_ENGINE=vllm is the other explicit choice — vllm even where mlx_lm.server would otherwise be preferred; unset leaves the choice to llmman as before.

In a container

With --runtime docker (or podman) and LLMMAN_SAFETENSORS_ENGINE=sglang, python3 -m sglang.launch_server runs from the lmsysorg/sglang image (multi-arch, no arch suffix), picked by the same GPU probe:

Host GPU Image
NVIDIA (CUDA 13) lmsysorg/sglang:latest
NVIDIA (CUDA 12) lmsysorg/sglang:latest-cu129 (upstream's last CUDA 12 lane is v0.5.19)
AMD (ROCm) no floating tag upstream: --sglang-version must be the whole tag for your GPU family, e.g. v0.5.19-rocm700-mi30x (x86_64 only)
none not offered: the -xeon CPU images need Intel AMX and launch flags llmman does not pass

--sglang-version <tag> pins the release (-cu129 is added for CUDA 12). llmman serve --runtime docker --pull-only <model> pulls it for a pulled safetensors model when the variable is set. CUDA_VISIBLE_DEVICES and friends plus every SGLANG_* variable are forwarded into the container, which runs with --ipc=host like vLLM's.

MLX (macOS)

On macOS, mlx_lm.server is preferred over vllm for safetensors when LLMMAN_SAFETENSORS_ENGINE is unset: Metal-accelerated (every Mac is assumed to have Metal), no vLLM dependency, more model families than vllm-metal. LLMMAN_CONTEXT_LENGTH is not forwarded and /v1/embeddings is unsupported.

Nothing needs installing first. llmman serve installs its own mlx_lm.server at startup, before it starts listening — the same way it fetches llama-server — with its own uv (the repo-root UV_RELEASE pin, under ~/.local/share/llmman/uv/<version>/). uv installs the mlx-lm in MLX_LM_RELEASE on a uv-managed CPython at the PYTHON_RELEASE pin — never your own Python; a version mlx, vllm and sglang all publish wheels for — under ~/.local/share/llmman/mlx-lm/<mlx-lm>-py<python>/ (python/ and venv/). --runtime applies as it does to llama-server: bin runs only this install, path only the mlx_lm.server on PATH, auto (the default) this install with PATH as the fallback if it fails.

llmman serve --pull-only does the same install with its progress on your terminal, then exits. A llmman run or launch that starts the daemon shows what it is fetching while it waits. If the install fails, the daemon starts anyway and the first safetensors load retries it. LLMMAN_MLX_LM_VERSION=<version> overrides the mlx-lm pin. LLMMAN_SAFETENSORS_ENGINE=vllm skips MLX (and the install).

Registry transport

Registry access goes through a Go shim compiled into the binary, with two implementations behind Cargo features. Building needs Rust and Go 1.22+.

Docker (default), via containerd's resolver:

cargo build --release

Podman, via container-libs:

cargo build --release --no-default-features --features podman