llama cpp router for scheduling and multi-GPU system
  • Python 67.7%
  • HTML 12.4%
  • JavaScript 11%
  • Nix 4.3%
  • CSS 4.2%
  • Other 0.4%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-09-27 13:43:41 -05:00
docker multi-GPU scheduling, GPU pinning, docker support 2026-07-02 11:01:08 +09:00
examples (refactor) split api extension from router code 2026-09-27 13:43:41 -05:00
frontend-demo [Chore] comment cleaning 2026-07-30 17:16:22 +09:00
scripts [Feature] Prompt Caching Support 2026-09-09 09:17:04 +09:00
src (refactor) split api extension from router code 2026-09-27 13:43:41 -05:00
tests [Fix] fixed model serving canceled requests and [Routine] added complementary testers for concurrent and canceled request 2026-09-22 15:46:12 +09:00
.envrc [Chore]: add direnv dev shell, changed ./src to match my style guidelines 2026-07-29 09:35:06 +09:00
.gitignore [Fix] fixed model serving canceled requests and [Routine] added complementary testers for concurrent and canceled request 2026-09-22 15:46:12 +09:00
CLAUDE.md [Fix] fixed model serving canceled requests and [Routine] added complementary testers for concurrent and canceled request 2026-09-22 15:46:12 +09:00
docker-compose.override.yml.example docker: make compose host-agnostic for multi-service orchestration 2026-07-17 04:15:17 +09:00
docker-compose.yml docker: make compose host-agnostic for multi-service orchestration 2026-07-17 04:15:17 +09:00
flake.lock moved code from kkuroma/nixos-configs 2026-07-02 09:48:05 +09:00
flake.nix [Fix] fixed model serving canceled requests and [Routine] added complementary testers for concurrent and canceled request 2026-09-22 15:46:12 +09:00
LICENSE Initial commit 2026-07-02 09:08:39 +09:00
module.nix (refactor) split api extension from router code 2026-09-27 13:43:41 -05:00
package.nix moved code from kkuroma/nixos-configs 2026-07-02 09:48:05 +09:00
pyrightconfig.json [Chore]: add direnv dev shell, changed ./src to match my style guidelines 2026-07-29 09:35:06 +09:00
README.md (refactor) split api extension from router code 2026-09-27 13:43:41 -05:00
run_tests.sh [Fix] fixed model serving canceled requests and [Routine] added complementary testers for concurrent and canceled request 2026-09-22 15:46:12 +09:00

LLaMa router

A router that sits in front of your llama.cpp's server that handles seamless model switching of many small GPUs for homelabs of a few users. It spawns and supervises llama-server processes on demand, routes an OpenAI-compaible traffic to the right one, and hot swaps model per GPU with VRAM constraint. Pure python, no build or compilation. Ships as a Nix flake, a NixOS module, and a Docker image.

What it does

  • It schedules your requests so nothing starves: Overload it and requests queue and drain, never a 429 error. The scheduler batches queued requests by resident model to swap as little as possible, spreads a hot model across num_instance replicas by least-busy routing, and force-loads any request that has waited too long so nothing gets starved out.
  • It rarely, if ever, crashes silently: Dead replicas are reaped and reload on their next request, and a load that runs out of VRAM fails fast with a reason instead of wedging the GPU in a half-dead state.
  • Explicit, declarative GPU placement: give each model a list of GPU ids, then the router manages eviction and device masking. No CUDA_VISIBLE_DEVICES needed to be set by hand. Models on disjoint GPUs stay resident and serve at the same time.
  • Ships a beautiful dashboard: /dash (per-GPU util/VRAM timeline, request history), /chat (embedded llama.cpp UI per replica for tok/s testing), /translate.

What it does better than LLaMa swap

llama-swap is another tool is aimed at a similar space of locally hosted LLMs on consumer hardware with VRAM constraint. It's gained more traction, but here's a quick comparison between the both of them:

llama-swap llama-router
Overload behavior reject (429) queue and drain, no drops
Multi-model concurrency manual groups matrix automatic, per-GPU residency
GPU placement manual CUDA_VISIBLE_DEVICES automatic masking from gpus
VRAM / GPU accounting none (you avoid oversubscribing) tracked per GPU, targeted eviction
Scheduling granularity whole-server per GPU (a swap on GPU 0 does not block GPU 1)
Replicas per model 1 num_instance, least-busy routing

llama-router only speaks llama.cpp, which is intentional. vLLM does not play well with hot-swapping and uses a messier environment, and we've decided against using it in a few-user homelab situation. If you want vLLM, whisper, and a single Go binary, use llama-swap. If you own a pile of small GPUs and want one endpoint that figure out where requests go without needing to manually manage VRAM states and never drops a request, use llama-server. Proven by months of stable use as "just an OpenAI compatible provider" you don't need to worry again. Perfectly compatible with LiteLLM, Librechat, OpenCode, or any local LLM tool.

Benchmarks

scripts/ holds standalone benchmarks (stdlib only, no venv needed) that write to outputs/benchmark-<task>.json:

  • benchmark-concurrency.py hammers a router with N short requests through a bounded worker pool and reports the success rate and latency distribution. It only needs the host and N (--host localhost:11434 -n 100). On a 3-GPU box serving three swapping models (gemma-4-26b on [0,2], gemma-4-12b on [1], a qwen-27b on [0,1,2]), N=100 concurrent requests completed at a 100% success rate: the router queued and drained every request across the swaps without dropping one.
  • benchmark-throughput.py measures prompt-processing and token-generation tok/s per model at concurrency 1 and each model's parallel-slot count.
  • benchmark-prompt-cache.py measures how much of a prompt llama.cpp reuses instead of reprocessing (--instance-url + --model): a reuse pass covering repeat, next turn, eviction and cache_prompt=false, then a capacity pass that holds N long conversations and revisits each one.

Tests

tests/ runs the scheduler against real weights, since ordering and concurrency do not survive a mock. Each case spawns src/main.py on its own port, drives it over HTTP, and reads concurrency off the arrival time of every streamed token.

  • test_request_order.py sends a long request to model 0, a request to model 1 a tenth of a second later, then another to model 0, and requires the replies in that same order. A later request for the resident model may not overtake the queued one that needs a swap. The test config sets QUEUE_HEAD_GRACE to 0 and parks QUEUE_FORCE_LOAD_TIMEOUT out of reach, so a pass can only come from the ordering and never from a timer firing.
  • test_canceled_order.py fills both slots with long requests, drops one client mid-stream, then sends two trivial requests behind it and requires them to finish before the survivor. A disconnect ends the request, so the slot has to come back immediately rather than when the abandoned generation would have ended.
  • test_concurrent_request.py fires N users at one model with parallel = P slots and requires min(N, P) of them to generate at the same instant, at a combined token rate well above a single stream. It reads the replica's own /slots alongside, which tells a stalled router apart from a replica that only came up with one slot.

Every request sends the same ~14k token essay from tests/prompt.txt and asks for a 2000 word continuation, differing only in its closing line, so the shared prefix and the long generation both behave the way a fleet of agents does. Ports, the llama-server binary, both model directories and the size of every case live in tests/configs.toml. Run them with ./run_tests.sh, which enters the nix dev shell first. They are marked gpu, so a bare pytest collects them and filters them out.

Prompt caching

Caching belongs to llama.cpp, not the router: request bodies are forwarded verbatim, so a prompt whose prefix a replica already processed is reused whether or not the client sets cache_prompt. It is on by default. Two things end a cached conversation: the host cache filling up, and the router evicting the model, which kills the process and everything it held.

Measured on one RTX 3090 with ctk/ctv = q4_0 and parallel = 1, at 16k tokens and again at each model's full context window:

model prompt cold cached host memory per conversation
Gemma-4-26B-A4B (mxfp4, MTP draft) 16k 5.0 s 0.38 s 268 MiB
Gemma-4-26B-A4B 131k, the full window 65.0 s 0.16 s 909 MiB
Wordslop-Qwen3.6-27B (iq2_m) 16k 17.4 s 0.42 s 1250 MiB
Wordslop-Qwen3.6-27B 262k, the full window 574 s 1.67 s 8500 MiB

A cached turn costs about the same whatever the prompt length, so the saving grows with the window: 13x at 16k on Gemma, 344x at 262k on the Qwen hybrid.

Do not size cram by extrapolating a bytes-per-token figure from a short prompt. Each held conversation costs a fixed part plus a part that grows with the prompt, and at 16k the fixed part dominates. Across the two lengths above:

model growing fixed per conversation
Gemma-4-26B-A4B 5.7 KiB/token 176 MiB
Wordslop-Qwen3.6-27B 30.2 KiB/token 766 MiB

The spread comes from attention geometry, which is readable from the GGUF header. Gemma sets sliding_window = 1024 and marks 25 of its 30 layers sliding in sliding_window_pattern, so only 5 full-attention layers hold a cache that grows with the prompt, and 5.7 KiB/token is what those 5 layers cost at q4_0. The Qwen hybrid sets full_attention_interval = 4 alongside its SSM parameters, so a quarter of its 65 layers grow and the rest hold a fixed-size recurrent state.

Sizing follows from the table: cram (MiB) is the cap on conversations that are not in a slot right now, per llama-server process, so it holds roughly cram / (cost of one conversation at the length you actually run). It is a cap and not an allocation. Past it llama.cpp drops the least recently used conversation, and a round robin over more conversations than fit degrades to no reuse at all rather than to partial reuse. At 16 GiB that is 18 Gemma conversations at 131k, but only 1 Qwen conversation at 262k.

One global cram therefore either starves the heavy model or overprovisions the light one. Give the heavy model its own cap with promptCache.ramMiBPerModel, which writes cram into that model's own preset section. Because the cap is per llama-server process, a large value there costs nothing while another model is loaded: with maxModelsPerGpu = 1 only one cap is ever live.

There is no time-based expiry anywhere in llama.cpp: size and LRU are the only controls. If you want conversations to stop occupying memory after some idle period, the lever is unloading the model, not the cache.

cache-reuse (salvaging a prompt whose middle changed, by KV shifting) is unavailable on sliding-window and hybrid contexts, which is most recent models. They log cache_reuse is not supported by this context at load and reprocess from the first changed token, so a client that trims old turns off the front of a conversation pays full price every turn.

Per request, the reuse is visible as timings.cache_n in the reply; per process, as llamacpp:prompt_tokens_cached_total on the replica's /metrics; per model over time, as the Cached column and the cached share of the Prompt tokens tile in /dash.

Scheduling model

One llama-server process per loaded model replica. Loading a model spawns num_instance llama-server processes (each hosting exactly that model, pinned to its GPUs); evicting kills them, which frees VRAM unconditionally. The router owns all placement decisions — llama-server's own models-max is irrelevant in this design and can be omitted.

Each model is pinned to a set of GPU ids. The router keeps at most MAX_MODELS_PER_GPU models resident per GPU (not globally): models pinned to disjoint GPUs stay in memory together, and loading a model only evicts residents on the GPUs it actually needs.

Example with MAX_MODELS_PER_GPU = 1: models A and B pinned to GPUs [0, 1], C and D pinned to [2]. A and C can be resident simultaneously (two llama-server processes). Requesting B kills only A's process (GPUs 0/1); C's process is untouched. Requesting D evicts only C.

  • No gpus field → the model counts against GPU 0 only. On a single-GPU host this reproduces plain global behavior exactly.
  • gpus = "all" or -1 → the model counts against every GPU (it will evict on all of them as needed).
  • Eviction policy: lru (default, evicts the model whose last request is oldest) or fifo (evicts the earliest-loaded model). Either way, residents with requests still waiting in the queue are only evicted when there is no other candidate on that GPU.
  • Anti-starvation: a head-of-queue request whose model isn't loaded holds the queue once it has waited QUEUE_HEAD_GRACE (default 0 s), so later requests for the resident model stop being served ahead of it and the swap happens as soon as the in-flight work drains. At 0 arrival order is absolute and two models in alternation pay a swap each time; raising it buys cache hits back at the cost of that ordering. QUEUE_FORCE_LOAD_TIMEOUT (default 300 s) stays as the last resort behind it.
  • Crash recovery: dead replicas are reaped automatically; the model simply reloads on its next request.
  • num_instance > 1 spawns that many replica processes of the model (requests balance across them by in-flight count). Each replica is a full copy of the weights — VRAM scales linearly.
  • GPU count is autodetected via NVML; override with ROUTER.NUM_GPUS (falls back to highest pinned id + 1 when NVML is unavailable).
  • Small overhead note: co-resident models are separate processes, so each pays its own CUDA context (~a few hundred MB per process per GPU it touches).

GPU pinning

The gpus field on the model's entry (config.json LLM section, or the gpus attribute in the NixOS module) is the single source of truth for both halves of pinning:

  1. Scheduler accounting — residency/eviction are tracked per GPU in gpus.
  2. Physical placement — the router spawns each llama-server with CUDA_VISIBLE_DEVICES set to that model's gpus, so the process only ever touches those devices.

You do not set a device key in presets.ini. Masking with CUDA_VISIBLE_DEVICES (rather than llama.cpp's --device) is deliberate: ggml initializes a CUDA context and reserves buffers on every visible device, so a model pinned via --device alone still holds hundreds of MB / a couple GB on the GPUs it isn't computing on — enough to OOM a co-resident model on a tight box. Masking keeps that overhead off other models' GPUs. Because the visible devices are renumbered from 0 inside each process, an absolute device = CUDA2 would not even resolve; omit it. Unpinned models default to GPU 0 (masked to GPU 0), so on multi-GPU hosts they no longer silently shard across all GPUs.

Usage as a NixOS module

Add the flake input (it follows your nixpkgs):

{
  inputs = {
    nixpkgs.url = "github:NixOS/nixpkgs/nixos-unstable";
    llama-router = {
      url = "git+https://git.kuroma.dev/kkuroma/llama-router";
      inputs.nixpkgs.follows = "nixpkgs";
    };
  };
}

Import llama-router.nixosModules.default into your host and configure:

{
  services.llama-router = {
    enable = true;
    port = 11434;

    # scheduler
    maxModelsPerGpu = 1;          # residency cap PER GPU
    evictionPolicy = "lru";       # or "fifo"
    queueForceLoadTimeout = 300;  # seconds before a starved request forces a load
    # gpuCount = 4;               # optional; autodetected via NVML otherwise

    # prompt cache, written into the "[*]" section (see Prompt caching above)
    promptCache = {
      ramMiB = 16384;    # host cap for idle conversations, per llama-server process
      ramMiBPerModel = { "Wordslop-Qwen3.6-27B" = 32768; };  # heavier model, own cap
      # reuseChunk = 0;  # KV shifting, ignored by sliding-window models
    };

    # llama.cpp settings applied to every preset (the "[*]" section)
    # these win over the promptCache keys, and a key on a model wins over everything
    presetGlobals = {
      jinja = true;
      fa = true;
      ngl = 99;
      ctk = "q4_0";
      ctv = "q4_0";
    };

    # each model becomes a presets.ini section; num_instance, gpus,
    # reasoning_effort, and cost are router-only (gpus masks the process via
    # CUDA_VISIBLE_DEVICES; no device key)
    models = {
      "Qwen3-4B" = {
        num_instance = 1;
        gpus = [ 0 1 ];  # omit = GPU 0; "all" or -1 = every GPU
        reasoning_effort = { options = [ "low" "medium" "xhigh" ]; default = "xhigh" };  # advertised via /v1/models; disable = "none" allows thinking-off
        cost = { input = 0.4; cached_input = 0.15; output = 2.5; };  # $/1M tokens, advertised as pricing
        model = "/data/llm-models/Qwen3-4B-Q8_0.gguf";
        c = 65536;
        b = 4096;
        parallel = 4;
      };
    };
  };
}

The router binds 127.0.0.1 by default (services.llama-router.host) — put a reverse proxy in front of it or scope your firewall accordingly before exposing it wider; spawned llama-server instances have no auth.

Use services.llama-router.llamaCpp to supply a CUDA/ROCm build of llama.cpp, e.g. pkgs.llama-cpp.override { cudaSupport = true; }. GPU masking sets both CUDA_VISIBLE_DEVICES and HIP_VISIBLE_DEVICES, so CUDA and ROCm both work with no extra configuration.

Usage with Docker

The image layers the router on top of ghcr.io/ggml-org/llama.cpp:server-cuda (which provides llama-server at /app/llama-server).

docker-compose.yml is a host-agnostic base: it defines how to build/run the router (port 11434, env, healthcheck, reserves all GPUs) but declares no mounts — where the weights/configs live is the deployer's decision. The router reads three paths inside the container: /models (ggufs, ro), /configs (config.json + presets.ini, ro) and /webui (request-history SQLite, rw). Supply them one of two ways:

Standalone — one router on one host:

cp docker-compose.override.yml.example docker-compose.override.yml
mkdir -p models configs webui
cp examples/config.json examples/presets.ini configs/
# drop your .gguf files into models/, edit the configs to match
docker compose up -d --build

docker compose auto-merges the override, which adds ./models, ./configs, ./webui. The override is gitignored, so git pull stays clean.

Orchestrated — this router as one service in a larger stack. A parent compose include:s this file and patches in the host mounts (and can point them at shared directories outside this repo, e.g. a central weights dir):

# ../docker-compose.yml   (run `docker compose up` from the parent dir)
include:
  - llama-router/docker-compose.yml
services:
  llama-router:
    volumes:
      - ../clanker-weights:/models:ro
      - ./config/llama-router:/configs:ro
      - ./webui:/webui:rw

include resolves this file's build: context relative to llama-router/, while the parent's added volumes: resolve relative to the parent — so upstream git pulls here never touch the host wiring.

To expose only some GPUs, override the device reservation (count: 2 or device_ids: ["0","1"]) from the deployer file — GPU ids inside the container renumber from 0, and each model's gpus pin refers to those container-local ids.

Running directly

nix run git+https://git.kuroma.dev/kkuroma/llama-router

Configuration is passed via environment variables:

Variable Default Purpose
ROUTER_CONFIG_PATH /configs/config.json Router config (models, scheduler timings, ports)
LLAMA_PRESETS_PATH /configs/presets.ini llama.cpp presets INI
ROUTER_HOST 0.0.0.0 API bind address
HISTORY_DB_PATH /webui/monitor/history.db SQLite request history

Scheduler settings live in the ROUTER section of config.json: MAX_MODELS_PER_GPU (default 1), EVICTION_POLICY (lru/fifo), QUEUE_HEAD_GRACE (seconds, default 0), QUEUE_FORCE_LOAD_TIMEOUT (seconds, default 300), NUM_GPUS (optional override), plus the health-check/load/unload timings shown in examples/config.json.