A GPU Scheduler for Local LLMs: What Broke on Real Hardware
Several things want my GPU at the same time: the chat I'm typing into, a coding agent, an indexer working through incoming documents, and a model that reads my screenshots. There is one card — an RTX 4060 Ti with 16 GB. Give each app its own llama-server and they know nothing about each other: while memory lasts, they split the GPU evenly, and when it runs out, something crashes. Ollama can unload models, but on a timer, with no idea that a chat matters more than indexing.
So in June I started writing GridCore, a scheduler that sits between applications and the GPU. The video shows what it does. This note is about how it's built — and what only broke once it met real hardware.
A control plane, not another inference engine
GridCore doesn't run inference itself. llama.cpp does that, one llama-server process per model. GridCore starts and stops those processes, accepts requests over the OpenAI API, and decides who gets the GPU right now.
For applications, only the address changes. The endpoints are the same, and aliases in the config let an existing client keep sending "model": "gpt-4o" without knowing that a local Gemma is behind it. The one extra thing an app can tell GridCore is its class — interactive, background or batch — through a header, a field in the request body (any OpenAI SDK can send it via extra_body), or a suffix on the model name:
curl localhost:8080/v1/embeddings \
-H 'X-GridCore-Class: background' \
-d '{"model": "nomic-embed", "input": ["...", "..."]}'A request without a class is treated as interactive.
The first question I get: why bother, when llama-server has had router mode since December 2025? Router mode also runs one process per model and also loads and unloads them, but it evicts the least recently used model once their number exceeds --models-max. It doesn't account for VRAM, it has no priorities, and its child processes are hidden inside it — so there is no way to measure how much memory each model actually takes. And those numbers are what every scheduling decision is built on.
GridCore is written in Go: a single static binary with no cgo, two dependencies (yaml.v3 and the Prometheus client), running as a systemd user service. The daemon is about 7,500 lines of code, with 120 tests.
A few hard rules instead of one clever algorithm
I wanted the scheduler's behavior to be predictable from the config alone. So at its core are a handful of simple rules:
- Strict class order. Interactive always goes first; within a class it's FIFO. No aging, where a job that has waited long enough gradually climbs in priority: sooner or later that would let a batch job overtake a chat.
- Preemption only between steps. A chat response is one indivisible step: GridCore never interrupts a generation halfway through. Background work, by contrast, is cut into small steps: a large
/v1/embeddingsrequest is split into chunks of 32 inputs, each scheduled separately, and the client still gets a single assembled response. Between steps, the GPU can go to the chat. - Interactive mode. While an interactive request is running, and for two seconds after it, no new background steps start. The two seconds keep the card from going to background work in the pause between two messages of the same conversation.
- Residency tiers. A model used interactively in the last five minutes is hot: background work can't evict it. The rest are cold and get evicted least-recently-used first. Pinned models — for me, a 0.4 GB embedder — are never evicted.
- Invariant number one. A process with requests in flight is never stopped. The scheduler first stops giving it work, waits for the running requests to finish, and only then unloads it.
All state belongs to a single goroutine: HTTP handlers only enqueue a job and wait for permission to run. A model that starts loading reserves its memory immediately, so two concurrent decisions can't claim the same space twice.
Here is what that looks like in the event log when a user starts chatting 300 ms after background indexing began:
11:20:27.194 enqueue background embedding nomic-embed steps=32
11:20:27.194 dispatch background step 1/32 -> nomic-embed
11:20:27.500 enqueue interactive chat gemma4-12b steps=1
11:20:27.500 dispatch interactive step 1/1 -> gemma4-12b
11:20:27.500 mode gpu interactive
11:20:28.731 complete interactive gemma4-12b
11:20:30.833 mode gpu idle
11:20:31.032 dispatch background step 32/32 -> nomic-embedThe chat got the GPU in the same millisecond, with no time in the queue. The background job waited for the answer to finish and the two-second window to pass, then picked up where it had left off.
Count first, load second
To decide whether a model fits, you need to know its VRAM footprint before loading it. The first version estimated it from the file size and was off by anywhere from −8% to +94%. The worst case was Gemma 4 E2B: an estimate of 5678 MB against an actual 2924. That model has a huge per-layer embedding table, which llama.cpp keeps in system RAM rather than on the card, so the file size tells you almost nothing.
Now GridCore reads the GGUF header and counts tensor by tensor: the weights, minus the embedding tables that stay in RAM; the KV cache, layer by layer, accounting for sliding-window layers, Gemma's shared-KV layers and hybrid SSM architectures; plus the vision projector and runtime buffers. On the models I tested, the error is between −1% and +10%. After the first load, the estimate is replaced by the memory the process actually took on the card.
Two mistakes worth knowing about if you do this math yourself:
--ctx-sizeinllama-serveris the total context pool shared by all slots, not a per-slot size. As long as I multiplied it by the number of slots, the Qwen3-14B estimate came out at 12.7 GB against a real 9.7.- Placement flags change everything. Ornith-1.5-35B-A3B with
--cpu-moekeeps 18.6 GB of experts in system RAM and takes 2.2 GB on the card — and still generates 46 tokens per second, as fast as a 9B model sitting entirely on the GPU. Without accounting for the flag, the estimate was 21 GB and the scheduler refused to load the model at all. Now it understands-ngl,--cpu-moe,--n-cpu-moeand-ot: 2229 MB estimated against 2244 measured.
This also explains a decision that looks odd at first glance: GridCore doesn't use the automatic placement fitting that llama.cpp has added. Auto-fit adapts to whatever memory is free at the moment of loading, while the scheduler needs a predictable footprint before it. So placement is set explicitly in each model's config.
For the same reason, GridCore has no KV cache saturation to fight and never frees llama.cpp's slots by hand. llama-server allocates the entire KV cache once, when the model loads: --ctx-size is the total pool, --parallel splits it evenly into slots, and that amount doesn't grow at runtime. GridCore doesn't touch the slots either: it tracks how many are busy itself and never sends a model more requests than it has slots. Everything else waits in GridCore's priority queue rather than in llama-server's internal queue — which is exactly what lets a chat overtake background work. If a request is longer than one slot's context, llama-server returns exceed_context_size_error, and GridCore passes that error back to the client as is, without truncating anything. The price of this approach is that KV cache memory is reserved for the worst case, even while slots sit idle. In exchange, you know in advance exactly how much memory a model will take.
What only broke on real hardware
On the simulated test GPU and simulated runtime, every scenario test passed. Then I ran a five-minute load test on the real card — 6 chat clients, 8 background, 4 batch and 2 embedding clients — and it found what the tests couldn't see.
- Complete background starvation. Six chat clients never once let the GPU leave interactive mode, and in five minutes the background jobs completed zero requests. Under constant load, "background waits while chat is quiet" means "background never runs". Now, if the head of the background queue hasn't made progress for 3 seconds, one background step may run alongside the chat. The first version of that rule didn't work either: the oldest starved job was waiting for a model that couldn't fit next to two hot models, and it blocked embedding jobs whose model was already loaded. Now the scheduler walks the starved jobs from oldest to newest and runs the first one that actually can run.
- Excess evictions. To free 200 MB, the scheduler unloaded an 8.5 GB model, then another 6.8 GB one. The first victim was still finishing its requests and its memory still counted as used — so the scheduler picked a second victim. Now the memory of models already marked for eviction counts as being freed.
- Phantom memory.
nvidia-smireports total used memory and the per-process list in two separate calls. If allama-serverexits between them, for one poll its memory looks like ~11 GB taken by some outside process. The scheduler concluded there was no room and, in the same millisecond, unloaded a second, perfectly healthy model. Now the estimate of outside memory is frozen while a process is shutting down and for two polls after. - 502s out of nowhere. Three errors in five minutes, and nothing in the
llama-serverlogs. The HTTP server inside llama.cpp closes a connection after five requests, and Go kept picking already-closed connections from its pool. One connection per request over loopback turned out to be cheaper than retries.
Each of these bugs became a regression test, and the simulated GPU learned things it couldn't do before — such as holding on to the memory of a process that is already gone from the list.
The numbers
All on the same RTX 4060 Ti 16 GB. During the load test, four models lived on the card at once — Gemma 4 12B for chat, Gemma 4 E4B and E2B for background and batch work, and the embedder — using 15.5 of 16 GB, with not a single unnecessary load and not a single error in five minutes:
| Class | Requests | p50 | p95 |
|---|---|---|---|
| interactive | 366 | 3.4 s | 4.1 s |
| background | 194 | 11.8 s | 13.3 s |
| batch | 96 | 11.8 s | 13.4 s |
| embeddings | 21 | 25 s | 34 s |
The honest cost: under mixed load, chat p50 went from 1.9 to 3.4 seconds. Part of the difference is six clients sharing two model slots, part is that one background step running next to the chat. If chat needs full exclusivity, background_max_starvation: 0s restores it — but then, under constant chat, background work doesn't move at all.
What v0.1 doesn't do
Deliberately left out: interrupting a chat mid-generation, multiple GPUs, AMD cards, vLLM and MLX runtimes, cloud fallback and authentication. Next in line are Radeon support, cost-based eviction instead of LRU (reload time × probability of reuse × priority), and quality tiers: under memory pressure, use a smaller variant of the same model. The Gemma 4 family fits that almost perfectly — 12B, E4B and E2B share one architecture and one prompt format.
The bottom line
The scheduling algorithm itself turned out to be the easy part. The hard part was accounting. The scheduler makes its decisions based on memory numbers, and nearly every serious bug came down to those numbers lying: an estimate based on file size, two inconsistent nvidia-smi calls, the memory of a process that hadn't finished unloading. A scheduler is only as good as its bookkeeping.
GridCore is in beta and available to Zero to MVP supporters on Patreon. If you're working on a similar problem — on a single GPU or a whole rack — I'd like to compare approaches: get in touch.
One AI signal, one tool, one MVP idea — a practical 5-minute email.
No spam, no link dumps. Unsubscribe in one click.

