KoboldCpp is free and open-source software for running GGUF large language models (LLMs) on your own computer. Chat with an AI assistant, write stories, roleplay, or connect other apps to a local API. KoboldCpp runs on CPU or GPU and also includes text, image, video, speech and music generation with compatible models, an integrated agent, a bundled KoboldAI Lite WebUI, and many additional powerful features.
Inspired by KoboldAI and built on llama.cpp
One executable file, no installation required. Ready-to-run downloads are available for Windows, Linux, and macOS.
Download KoboldCpp | Documentation and FAQ | API reference | Discord community
- Local text generation: Run any GGUF language model to generate text in chat, adventure, instruct, or story writing modes. Compatible vision models also support image understanding.
- Image generation and editing: Supports Stable Diffusion 1.5, SDXL, SD3, Flux, Qwen Image, Ideogram, Z-Image, Klein, Krea2 and more.
- Video generation: Supports WAN 2.2, LTX2.3, MiniMax H3 and more.
- Voice recognition: Speech-to-text with Whisper, and multimodal audio from Gemma4 E2B and E4B.
- Speech generation: Text to speech with Qwen3TTS, Kokoro, OuteTTS, Parler, and Dia.
- Music generation: ACE Step 1.5 and ACE Step XL.
- Tools and agents: MCP server support, tool calling, web search, retrieval-augmented generation (RAG) through TextDB, and an integrated KoboldCpp Agent for writing and editing code, running programs, and scheduling tasks.
- To start the agent, select Launch KoboldCpp Agent in the launcher's Admin tab, or add
--agentto your launch command.
- To start the agent, select Launch KoboldCpp Agent in the launcher's Admin tab, or add
- Writing and roleplay tools: Bundled KoboldAI Lite WebUI includes multiple UI themes, editing tools, memory, world info, author's notes, characters, scenarios, and persistent story saves. Import Tavern character cards and other supported formats from file or external sites. Also includes the classic llama.cpp WebUI.
- App integrations: Compatible endpoints for KoboldAI, OpenAI, Ollama, A1111/Forge, ComfyUI, Whisper transcription, XTTS, and OpenAI speech clients. See APIs and integrations.
- Fully Portable - Single standalone executable for Windows, Linux or macOS, with no installation required and no external dependencies. Runs on CPU or GPU, with full or partial offloading. Can also run on Colab, Docker, also supports other platforms if self-compiled (like Android via Termux and Raspberry PI).
- Phishing SCAM Warning:
koboldcpp.comis a malicious fake site and is not affiliated with this project. You should ONLY trust official downloads from the release binaries on the official github at https://github.com/LostRuins/koboldcpp/releases/latest
- Download KoboldCpp for your operating system from the latest KoboldCpp release. See the platform instructions below for help choosing a file.
- Download a GGUF text model. Models are separate from the software. If unsure, start with the example models, or open Get Help and pick from Newbie Templates in the launcher for an easy setup.
- Open KoboldCpp and select your model in the GGUF Text Model field. Choose hardware settings suited to your computer. Generally the defaults should work, see GPU and performance settings if needed.
- Need help launching? See the platform instructions.
- A dedicated GPU is optional; the model size and context length determine how much memory you need.
- Click Launch and wait for the model to load. Keep KoboldCpp running while you use it.
- Connect to the Web UI in your browser once ready at http://localhost:5001
You should choose the correct KoboldCpp executable from the Assets section of the latest KoboldCpp release. Here are direct links and a quick overview for each supported platform. Models must be obtained separately
- Download
koboldcpp.exeand double-click it to open the launcher. - If you do not need NVIDIA CUDA support,
koboldcpp-nocuda.exeis a smaller download with CPU and Vulkan support. - If you're using a PC with an older CPU or GPU, try
koboldcpp-oldpc.exeif you encounter compatibility issues. - KoboldCpp can also be run using the command line, for example
koboldcpp.exe --model "C:\Models\model.gguf"For more info, please checkkoboldcpp.exe --help - One liner setup for windows:
cmd /c "curl -fLo koboldcpp.exe https://github.com/LostRuins/koboldcpp/releases/latest/download/koboldcpp.exe && koboldcpp.exe"
- Download
koboldcpp-linux-x64for an x86-64 Linux system. Make it executable withchmod +x koboldcpp-linux-x64, then launch it from a terminal in the download folder with./koboldcpp-linux-x64. - Use
koboldcpp-linux-x64-nocudaif you do not need CUDA, or trykoboldcpp-linux-x64-oldpcfor older hardware. Substitute that filename in the commands above. - For command-line usage see
./koboldcpp-linux-x64 --help, models can be loaded directly with./koboldcpp-linux-x64 --model /path/to/model.gguf - For other hardware or distributions that cannot run the binaries, see building from source.
- One liner setup for linux:
curl -fLo koboldcpp-linux-x64 https://github.com/LostRuins/koboldcpp/releases/latest/download/koboldcpp-linux-x64 && chmod +x koboldcpp-linux-x64 && ./koboldcpp-linux-x64
- Download
koboldcpp-mac-arm64for an Apple Silicon (M-series) ARM64 Mac. Make it executable withchmod +x koboldcpp-mac-arm64, then launch it from a terminal in the download folder with./koboldcpp-mac-arm64. - If macOS blocks the app, follow Apple's instructions to whitelist it in security settings under System Settings and Privacy & Security. A macOS launch walkthrough is also available.
- For command-line usage see
./koboldcpp-mac-arm64 --help, models can be loaded directly with./koboldcpp-mac-arm64 --model /path/to/model.gguf - Intel Macs require a source build.
Android users can install through Termux. Source builds also support platforms such as OpenBSD and Raspberry Pi, see the build instructions and KoboldCpp wiki.
- KoboldCpp allows connecting the web UI directly with a supported external AI provider instead of using a local model. Currently, AI Horde, OpenAI Compatible, Anthropic, OpenRouter, Gemini, Grok, Mistral are among the supported services.
- To connect, start with
--nomodelor select Allow Launch Without Models in the launcher's Loaded Files tab. Alternatively, you can also use the online KoboldAI Lite WebUI directly.
- Run on Google Colab: Use the official KoboldCpp Colab GPU Notebook. This is an easy way to get started without installing anything in a minute or two. Your usage must comply with Colab's terms.
- Run on RunPod: Cloud rental GPUs with variable prices. Launch the KoboldCpp RunPod image, or try SimplePod for smaller models.
- Public demo: Try the KoboldCpp HuggingFace Space. Please be considerate of this free shared service.
- Caution: The official KoboldCpp Docker image is intended for experts only, primarily for cloud GPU rentals. It uses an x86-64 Ubuntu environment internally and expects an NVIDIA or AMD GPU.
- Docker may perform poorly on some Windows or macOS setups, and ARM systems may fail to run it. CPU feature detection can also incorrectly select slower fallback binaries on some systems.
- For local use, you're recommended to start with the prebuilt binaries.
KoboldCpp does not include model files. For local text generation, huggingface.co hosts many GGUF models, including Bartowski's model collection. Image generation, music and audio features use their own model files and settings, and CivitAI has a good source of image models. Start with a smaller model if you are unsure what your computer can run. Alternatively, click 'Get Help' in the GUI launcher and browse the 'Newbie Templates' (recommended)
KoboldCpp also retains backward compatibility with legacy GGML .bin models, though some newer features may be unavailable.
- General Text Generation: Qwen3-VL-8B (Most Recommended, best all rounder model)
- Add optional Qwen3-VL-8B MMproj file for vision recognition capabilities.
- Creative Writing and Roleplay: L3-8B-Stheno-v3.2
- Lightweight and Fast: Gemma3-4B
- Add optional Gemma3-4B MMproj file for vision recognition capabilities.
- Image Generation: PicX Real
- Speech Recognition: Whisper models for Speech-To-Text
- Text-To-Speech: TTS models for Narration
- This is just a list for noobs to get started! There are hundreds more GGUFs out there!
- More newbie templates - Contains premade KoboldCpp quick launch templates curated for newbies.
- More popular templates - Contains premade KoboldCpp quick launch templates curated for popularity.
- More extra templates - Other premade KoboldCpp templates
To enable vision, load the matching MMProj file in Mmproj File under the launcher's Loaded Files tab, or add --mmproj /path/to/mmproj.gguf to your launch command alongside the text model.
To convert your own models, use the GGUF conversion and quantization tools: run convert_hf_to_gguf.py, then quantize_gguf.exe to quantize the result.
KoboldCpp provides many hardware configurations that can affect performance. Generally the default configuration should work decently, however you can make some adjustments to optimize your experience.
- System runs out of RAM: Try a smaller model or a shorter context. Reducing GPU layers can increase system RAM usage by moving more model weights off the GPU.
- Adjusting Context Size: Use
--contextsize Nto set the maximum context length: how much text the model can work with at once, measured in tokens. Larger contexts need more memory. - Adjusting Batch Size: Use
--batchsize Nto adjust prompt-processing batch size. A smaller batch can use less memory but might be slower. Set-1to disable batching.
- Adjusting Context Size: Use
- GPU runs out of VRAM: Try fewer GPU layers, a smaller model, or a shorter context. Autofit is an estimate and may need manual adjustment.
- Generation is slow: Check that the intended GPU backend is selected and that layers are offloaded. CPU-only generation works, but speed depends on your hardware and model.
- GPU Acceleration: Windows and Linux users with GPUs can use
--usecudaflag (Nvidia Only), or--usevulkan(AMD, Nvidia, Intel GPUs) for GPU acceleration, make sure you select the correct .exe with CUDA support. This is also selectable in the hardware preset in the GUI launcher. - GPU Layer Offloading: Add
--gpulayers Nto offload model layers to the GPU. The default,-1, enables autofit;0disables GPU offloading. Lower the layer count if you run out of GPU memory.
- GPU Acceleration: Windows and Linux users with GPUs can use
- An older computer crashes at startup: Some devices lack newer CPU instruction support. Try an
oldpcrelease or use--noavx2. See the release notes for hardware compatibility. - The browser cannot connect: Wait for model loading to finish and check the address printed in the terminal. The default is localhost:5001; a custom
--portchanges it. - A model will not load: Check the terminal error, your available memory, and whether your KoboldCpp version supports the model. You can trigger debug mode with the
--debugmodeflag or launch toggle. Try the latest release and check the wiki or create a Github issue to report a bug. - Model is incoherent: You might be using an incorrect chat template. Try relaunch with
--jinjatoolsto use the included Jinja template, or enable the Jinja toggle in the GUI. - For more information, be sure to run the program with the
--helpflag.
KoboldCpp serves many APIs alongside multiple bundled web UIs. Simply connect your software With the default port:
| Interface | Base URL |
|---|---|
| KoboldCpp Default API base | http://localhost:5001 |
| OpenAI-compatible API | http://localhost:5001/v1 |
| Interactive API documentation | http://localhost:5001/api |
| KoboldAI Lite web UI | http://localhost:5001 |
| llama.cpp web UI | http://localhost:5001/lcpp |
| StableUI Image Gen UI | http://localhost:5001/sdui |
| MusicUI Music Gen UI | http://localhost:5001/musicui |
Additional APIs supported: KoboldAI, OpenAI, Anthropic, Ollama, AUTOMATIC1111, ComfyUI, XTTS
Image, speech, embedding and music generation require the corresponding models to be loaded. Replace the host and port when connecting to a remote server or using a custom --port. For other apps, select a compatible API type and point the app at your running KoboldCpp server.
- Supported endpoints and request formats are described in the KoboldCpp API reference
Use a source build if a prebuilt binary does not suit your platform or you want to develop KoboldCpp. Manual builds require Git, Python 3, a C/C++ toolchain, and the development libraries for your chosen GPU backend.
Optional Python runtime dependencies include customtkinter and Tk support for the GUI launcher, jinja2 for chat templates, and psutil for system information. Install the Python packages you need in your Python environment with python -m pip install -r requirements.txt; Tk may require a separate package from your operating system. The automated Linux build script manages its own dependencies.
Start by cloning the repository and entering its directory:
git clone https://github.com/LostRuins/koboldcpp.git
cd koboldcppRun the following build commands from that directory. Use make -jN to compile with N parallel jobs. For manual builds intended for other machines, add LLAMA_PORTABLE=1; this avoids optimizing only for the build machine, but platform and runtime requirements still apply.
Automated build: koboldcpp.sh uses a local micromamba/conda environment to obtain dependencies and build the libraries. Install curl and bzip2 first.
./koboldcpp.sh # Build as needed and open the launcher (requires X11)
./koboldcpp.sh --help # Show command-line options
./koboldcpp.sh rebuild # Refresh the environment and rebuild after updates
./koboldcpp.sh dist # Package a standalone binary in dist/To build for other machines, use KCPP_PORTABLE=1 ./koboldcpp.sh dist. The packaged binary still depends on the target system's compatibility with the Linux environment used to build it.
Manual build: Run make for a CPU build, or choose a backend below and install its prerequisites.
| Backend | Build command | Prerequisite |
|---|---|---|
| CPU | make |
C/C++ toolchain |
| Vulkan | make LLAMA_VULKAN=1 |
Vulkan SDK |
| NVIDIA CUDA | make LLAMA_CUBLAS=1 |
CUDA Toolkit |
| AMD ROCm | make LLAMA_HIPBLAS=1 |
ROCm development libraries |
| CUDA and Vulkan | make LLAMA_CUBLAS=1 LLAMA_VULKAN=1 |
Both toolkits |
After building, launch with python3 koboldcpp.py --model /path/to/model.gguf.
- Install the standard x64 version of w64devkit, not the i686 variant.
- Open its integrated terminal in the repository directory.
- Run
makefor a CPU build, ormake LLAMA_VULKAN=1for Vulkan. This produces the DLLs used bykoboldcpp.py. - Launch with
python koboldcpp.py --model "C:\Models\model.gguf".
CUDA builds require Visual Studio, CMake, and the CUDA Toolkit. Open the project's CMake configuration in Visual Studio, build it, and copy koboldcpp_cublas.dll beside koboldcpp.py. The Makefile's LLAMA_CUBLAS=1 option is for Linux. Portable CUDA executables must include matching cublas, cublasLt, and cudart libraries from the same CUDA Toolkit family used for the build.
Packaging an executable: Install PyInstaller and the Python modules collected by make_pyinstaller.bat. Build the CPU and Vulkan libraries with make LLAMA_VULKAN=1 LLAMA_PORTABLE=1 to provide the script's required DLLs, then run the batch file from a Windows command prompt. It produces dist/koboldcpp-nocuda.exe; see the Windows release workflow for CUDA packaging.
If replacing bundled Vulkan libraries, put the matching .lib files in kcpp_src/lib and their .dll files in the repository root, then rebuild. This is an advanced configuration.
Run make for a CPU build. For Metal GPU support, install the Apple command-line developer tools and build with:
make LLAMA_METAL=1
python3 koboldcpp.py --model /path/to/model.gguf --gpulayers -1Install GNU Make with pkg_add gmake, then use gmake for a CPU build. For Vulkan, install vulkan-loader (also included as a dependency of vulkan-tools) and shaderc:
pkg_add gmake vulkan-loader shaderc
ulimit -d 8388608
gmake LLAMA_VULKAN=1
python3 koboldcpp.py --model /path/to/model.ggufRun package installation with the required system privileges. The ulimit setting raises the data-size limit for compilation. If the build reports that ggml-vulkan-shaders.hpp is missing, check that glslc from shaderc is installed.
Install Termux from F-Droid, then choose an automated or manual setup.
Automated setup: Download and run the Android installer. Its interactive menu offers installation with a starter model or without a model.
curl -sSL https://raw.githubusercontent.com/LostRuins/koboldcpp/concedo/android_install.sh | shManual setup: In Termux, install the dependencies, clone the repository if you have not already done so, and build:
pkg update
pkg upgrade
pkg install openssl wget git python clang make
git clone https://github.com/LostRuins/koboldcpp.git
cd koboldcpp
make
python koboldcpp.py --model /path/to/model.ggufDownload a small GGUF model before the final command, and replace the model path with its location. Open localhost:5001 in your mobile browser once loading completes. For portable ARM builds, LLAMA_PORTABLE=1 disables native ARM instruction optimizations.
If package installation fails, update your packages or use termux-change-repo to choose another mirror. These instructions cover CPU builds; GPU acceleration depends on the device and its drivers.
Start with the KoboldCpp FAQ and knowledge base, then search existing issues and discussions. If you still need help, open an issue or join the KoboldAI Discord.
For troubleshooting, include your operating system, hardware, KoboldCpp version, model filename, launch settings, and relevant error output.
These community projects may be outdated or unmaintained. Contact their maintainers for support.
- Arch Linux: AUR packages for CUDA and HIPBLAS.
- Community Docker images: korewaChino and noneabove1182.
- Nix and NixOS: Add
koboldcpptoenvironment.systemPackagesorhome.packages. See the Nix setup example and report packaging problems to Nixpkgs. - AMD ROCm fork: YellowRoseCx/koboldcpp-rocm. Check its maintenance status; Vulkan in the main release is a starting point for AMD users.
- Microsoft Word integration: GPTLocalhost connects Word to a local KoboldCpp server.
KoboldCpp and KoboldAI Lite are licensed under the GNU AGPL v3.0, unless a file states otherwise. Bundled components retain their respective licenses, including the MIT license for GGML, llama.cpp, and stable-diffusion.cpp.
KoboldCpp builds on the work of these projects:
- GGML and llama.cpp (MIT)
- stable-diffusion.cpp (MIT)
- TTS.cpp (MIT)
- Qwen3-TTS.cpp (MIT)
- acestep.cpp (MIT)
- KoboldAI Lite (AGPL)
For enquiries, contact @concedo on discord, message u/HadesThrowaway on reddit, or find LostRuins on github.





