llmman does not ship an inference engine. llmman serve picks one that
already exists for the model format it finds, and runs it unmodified.
| Model format | Backend | Where it comes from |
|---|---|---|
| GGUF | llama-server in a container |
--runtime docker / podman (Linux only; what auto tries first): the ghcr.io/ggml-org/llama.cpp:server-<backend> image for your GPU |
| GGUF | llama-server |
--runtime bin: a prebuilt upstream release matching your OS/arch/GPU, downloaded and cached; --runtime path: the one on your PATH |
| safetensors | vllm in a container |
--runtime docker / podman (Linux only): the vllm/vllm-openai, rocm/vllm or vllm/vllm-openai-cpu image for your GPU and architecture |
| safetensors | vllm |
Your PATH (every non-container runtime) |
| safetensors | sglang in a container |
--runtime docker / podman (Linux only) with LLMMAN_SAFETENSORS_ENGINE=sglang: the lmsysorg/sglang image for your GPU |
| safetensors | sglang |
Your PATH, with LLMMAN_SAFETENSORS_ENGINE=sglang |
| safetensors | mlx_lm.server |
macOS: llmman's own uv-installed copy, installed when the daemon starts, or the one on your PATH — per --runtime, as for llama-server; preferred over vllm |
| GGUF diffusion (LTX-2) | llmman itself, on ggml | The libggml/libllama next to llama-server; see the blog post |
| Diffusers safetensors (Qwen-Image 2.1) | llmman itself, on ggml | The same libraries; see below |
| Diffusers safetensors | vllm serve --omni |
Your PATH's vllm with the vllm-omni package installed |
| Diffusers safetensors | vllm serve --omni in a container |
--runtime docker / podman (Linux only): the vllm/vllm-omni image (CUDA only) |
llmman serve --runtime <auto|docker|podman|bin|path> (or the
LLMMAN_RUNTIME environment variable, which also reaches the daemon
llmman run/launch start for you) says where the engine comes from:
| Value | llama-server | Notes |
|---|---|---|
docker, podman |
ghcr.io/ggml-org/llama.cpp:server-<backend>-<release> container |
Linux only; safetensors models get the vLLM images too |
bin |
llmman's own download of llama.cpp's prebuilt release for this OS/arch/GPU | Installed under ~/.local/share/llmman/llama.cpp/<release>/; nothing needed on PATH |
path |
the llama-server already on PATH, as-is (and, on macOS, the mlx_lm.server on PATH) |
Never downloads anything; --llama-cpp-version does not apply |
auto (default) |
the first of docker, podman, bin, path that works here |
Off Linux: bin, then path |
Under auto, a container engine "works" when its CLI is on PATH, its
daemon answers docker info/podman info, an NVIDIA host has the NVIDIA
Container Toolkit, and the llama.cpp image pulls; bin works when the
release downloads (or is already cached). Each step skipped is logged
with the reason. Whatever is chosen is fetched before the listener binds,
so the first request is never stuck behind a silent download. On macOS
mlx_lm.server follows the same rule at the same point: bin installs
llmman's own uv and mlx-lm, path uses PATH's, auto installs
and falls back to PATH (see MLX).
Everything llmman installs for itself — llama.cpp, uv, mlx-lm — has
one shape on disk: ~/.local/share/llmman/<tool>/<version>/, a
.llmman-complete sentinel written last (without it the directory is
rebuilt), one installer at a time across processes (<tool>/.lock), and
earlier versions deleted once a new one completes. The pins are the
repo-root LLAMA_CPP_RELEASE, UV_RELEASE, MLX_LM_RELEASE and
PYTHON_RELEASE files, which CI tests against.
If that fetch fails, the daemon still starts: the failure is a warning,
and the first model load that needs llama.cpp retries it and, failing
again, reports that to the request. Nothing remembers the failure. A
safetensors model under bin or path needs no llama.cpp and loads
regardless. mlx-lm is the same: a failed install is a warning, retried
by the first safetensors load. Only --pull-only treats a failed fetch
as an error.
llmman serve --pull-only does exactly those fetches, in the foreground
with their own progress, then exits — run it once before starting a
detached daemon. With a container runtime and a safetensors MODEL
argument it pulls the vLLM image as well.
The release is pinned: unless --llama-cpp-version <tag> says otherwise,
both bin and the container images use the llama.cpp build llmman's own
CI tested (the LLAMA_CPP_RELEASE file in the repository), so what runs
is what was tested. --llama-cpp-version latest takes upstream's
floating latest instead.
For bin, llmman probes for CUDA, ROCm, OpenCL (Windows ARM64 Adreno only),
Vulkan or Metal (in that order) and downloads the matching prebuilt asset from
llama.cpp's GitHub releases. LLMMAN_LLM_LIBRARY overrides the probe; LLMMAN_DEBUG=1
shows what it found. llama.cpp publishes no prebuilt Linux CUDA binary,
so an NVIDIA host on Linux gets the CPU build from bin — the container
runtimes (which auto prefers for that reason) have CUDA images.
Context length, parallel slots, flash attention, KV-cache type and GPU split are environment variables; see configuration.md.
On Linux, --runtime docker (or podman) runs llama-server from the
ghcr.io/ggml-org/llama.cpp image instead, picking the
server-cuda/server-cuda13/server-rocm/server-vulkan/server tag
for the host, suffixed with the pinned release.
CUDA_VISIBLE_DEVICES and friends are forwarded into the container.
The container is a sibling of the daemon, not part of its cgroup, so a
CPU limit on llmman serve is forwarded: when one binds (cgroup quota or
affinity mask), the container gets --cpus <n>. llama-server always gets
--threads <n>: half the CPUs the daemon can use above 8, else all of them
(a quota alone would leave it autodetecting a thread per host core and
throttling). No limit, no --cpus. An explicit LLAMA_ARG_THREADS still
wins. vLLM containers get the same --cpus.
Each of those images is also published with llmman in it, as
docker.io/ai/llmman:<tag> (latest is server) and <tag>-<llmman version>, built from packaging/Dockerfile on
every release against the llama.cpp build CI tests. Same entrypoint and GPU
flags as upstream, plus /usr/local/bin/llmman. LLMMAN_HOST is preset to
0.0.0.0:17434 (a loopback bind inside a container is unreachable even with
-p), so the daemon requires LLMMAN_API_KEYS or LLMMAN_AUTH=off
(configuration.md); publish the port on
the host's loopback to keep it local. The store is /root/.local/share/llmman:
docker run -p 127.0.0.1:17434:17434 -e LLMMAN_API_KEYS=<key> \
-v llmman:/root/.local/share/llmman --gpus all \
--entrypoint llmman ai/llmman:server-cuda serveSafetensors models are served by a separately installed vllm unless
LLMMAN_SAFETENSORS_ENGINE picks SGLang, or the host is a
Mac (see MLX).
Plain vllm is CPU-only on macOS unless
vllm-metal is installed.
LLMMAN_CONTEXT_LENGTH is forwarded as --max-model-len;
LLMMAN_LOAD_TIMEOUT (default 10 minutes) bounds a stalled load.
LLMMAN_VLLM_ARGS appends anything else (--dtype bfloat16 --tp 2) to
the vllm serve command line, local or in a container.
On Linux, --runtime docker (or podman) runs vllm serve from a vLLM
image for a safetensors model, picked by the same GPU probe plus the
host architecture:
| Host GPU | x86_64 | aarch64 |
|---|---|---|
| NVIDIA (CUDA 13) | vllm/vllm-openai:latest-x86_64 |
vllm/vllm-openai:latest-aarch64 |
| NVIDIA (CUDA 12) | vllm/vllm-openai:latest-x86_64-cu129 |
vllm/vllm-openai:latest-aarch64-cu129 |
| AMD (ROCm) | rocm/vllm:latest |
not published upstream (LLMMAN_LLM_LIBRARY=cpu for the CPU image) |
| Vulkan-only or none | vllm/vllm-openai-cpu:latest-x86_64 |
vllm/vllm-openai-cpu:latest-arm64 |
--vllm-version <tag> pins the release (the arch suffix is added for the
vllm/ images; for rocm/vllm it is the whole tag).
llmman serve --runtime docker --pull-only <model> pulls the image an
already-pulled model needs (without a model, the llama.cpp image).
CUDA_VISIBLE_DEVICES and friends plus every VLLM_* variable are
forwarded into the container.
A safetensors repository laid out as a Diffusers pipeline (a root
model_index.json next to transformer/, vae/, ...) is
served by vLLM-Omni: the same
vllm launcher with --omni. Plain vllm serve cannot load one. The
exception is a pipeline llmman runs itself (see Qwen-Image 2.1).
uv pip install vllm==0.28.0 vllm-omni # into the environment `vllm` runs from
llmman run ORG/MODEL "A robot arm cleaning a plate in a kitchen"
llmman run ORG/MODEL --video --seconds 2 "A robot arm cleaning a plate"If vllm is a #!/path/to/python script whose Python cannot import
vllm_omni, the load fails up front with a message saying so; otherwise
vLLM itself reports unrecognized arguments: --omni. The model answers
/v1/images/generations and /v1/videos, in llama-server's dialect
(width/height/steps/cfg_scale, streamed image_generation.*
events, a video job with a content_url) and vLLM-Omni's own (size,
num_inference_steps, guidance_scale, num_frames, extra_params);
unsent fields are left to the model's defaults. There is no
/v1/audio/speech for these models.
The server is started with --no-guardrails; LLMMAN_VLLM_OMNI_GUARDRAILS=1
leaves the pipeline's safety guardrails on (they may need extra packages
and gated weights). LLMMAN_LOAD_TIMEOUT is also passed as --init-timeout.
With a container runtime, the image is vllm/vllm-omni:latest-x86_64 or
latest-aarch64 (CUDA only; --vllm-version pins vLLM-Omni's release,
e.g. v0.28.0). --pull-only <model> picks it for a pulled Diffusers model.
QwenImage21Pipeline repositories (docker.io/ai/qwen-image-2.1,
Qwen/Qwen-Image-2.1) run in llmman on the same ggml libraries, straight
from their bf16 safetensors:
llmman run qwen-image-2.1 "A manatee in a sunlit lagoon"Text to image only: 1024x1024, 40 steps, RGBA PNG. The text encoder and the transformer load on first use and take turns when both do not fit.
The inverse: the vllm-llmman
plugin lets vllm serve oci://<reference> pull a CNCF ModelPack image
from any OCI registry, via llmman (LLMMAN_BIN if it is not on
PATH). See vllm-plugin/README.md.
LLMMAN_SAFETENSORS_ENGINE=sglang sends safetensors models to
SGLang instead of vllm
(uv pip install sglang puts its sglang console script on PATH):
LLMMAN_SAFETENSORS_ENGINE=sglang llmman serve
llmman run Qwen/Qwen3-0.6B "hello"llmman runs sglang serve --model-path <dir> --served-model-name <ref>
on the model's directory; LLMMAN_CONTEXT_LENGTH is forwarded as
--context-length and LLMMAN_LOAD_TIMEOUT bounds the load as for
vLLM. LLMMAN_SGLANG_ARGS appends anything else (--disable-cuda-graph --tp 2). Which device SGLang uses is its own autodetection and
environment, which a local child inherits: SGLANG_USE_CPU_ENGINE=1 for
its CPU engine (Linux; x86_64 needs AVX-512), SGLANG_USE_MLX=1 for its
MLX runtime on Apple Silicon. llmman ps shows sglang (local).
LLMMAN_SAFETENSORS_ENGINE=vllm is the other explicit choice — vllm
even where mlx_lm.server would otherwise be preferred; unset leaves
the choice to llmman as before.
With --runtime docker (or podman) and LLMMAN_SAFETENSORS_ENGINE=sglang,
python3 -m sglang.launch_server runs from the lmsysorg/sglang image
(multi-arch, no arch suffix), picked by the same GPU probe:
| Host GPU | Image |
|---|---|
| NVIDIA (CUDA 13) | lmsysorg/sglang:latest |
| NVIDIA (CUDA 12) | lmsysorg/sglang:latest-cu129 (upstream's last CUDA 12 lane is v0.5.19) |
| AMD (ROCm) | no floating tag upstream: --sglang-version must be the whole tag for your GPU family, e.g. v0.5.19-rocm700-mi30x (x86_64 only) |
| none | not offered: the -xeon CPU images need Intel AMX and launch flags llmman does not pass |
--sglang-version <tag> pins the release (-cu129 is added for CUDA 12).
llmman serve --runtime docker --pull-only <model> pulls it for a pulled
safetensors model when the variable is set. CUDA_VISIBLE_DEVICES and
friends plus every SGLANG_* variable are forwarded into the container,
which runs with --ipc=host like vLLM's.
On macOS, mlx_lm.server is preferred over vllm for safetensors when
LLMMAN_SAFETENSORS_ENGINE is unset: Metal-accelerated (every Mac is
assumed to have Metal), no vLLM dependency, more model families than
vllm-metal. LLMMAN_CONTEXT_LENGTH is not forwarded and /v1/embeddings
is unsupported.
Nothing needs installing first. llmman serve installs its own
mlx_lm.server at startup, before it starts listening — the same way it
fetches llama-server — with its own uv
(the repo-root UV_RELEASE pin, under ~/.local/share/llmman/uv/<version>/).
uv installs the mlx-lm in MLX_LM_RELEASE on a uv-managed CPython at
the PYTHON_RELEASE pin — never your own Python; a version mlx, vllm
and sglang all publish wheels for — under
~/.local/share/llmman/mlx-lm/<mlx-lm>-py<python>/ (python/ and
venv/). --runtime applies as it does to llama-server: bin runs
only this install, path only the mlx_lm.server on PATH, auto
(the default) this install with PATH as the fallback if it fails.
llmman serve --pull-only does the same install with its progress on
your terminal, then exits. A llmman run or launch that starts the
daemon shows what it is fetching while it waits. If the install fails,
the daemon starts anyway and the first safetensors load retries it.
LLMMAN_MLX_LM_VERSION=<version> overrides the mlx-lm pin.
LLMMAN_SAFETENSORS_ENGINE=vllm skips MLX (and the install).
Registry access goes through a Go shim compiled into the binary, with two implementations behind Cargo features. Building needs Rust and Go 1.22+.
Docker (default), via containerd's resolver:
cargo build --release
Podman, via container-libs:
cargo build --release --no-default-features --features podman