colibri
Quickstart GitHub

Run open models.
Keep your hardware.

colibri streams a model's experts off the disk instead of loading them, so a 2.8-trillion-parameter model answers on a machine that could never hold it.

A real session, nothing staged: coli chat attached to a running server, then the same model in System One mode. The numbers are the ones it produced.

Measured today on machines people actually have

RTX 5090Graviton4Xeon Gold DGX SparkM5 MaxRyzen 9Mac Mini

Get started in one step

One file looks at your machine, recommends a model that fits, gets the engine, downloads the model with resume and opens the dashboard. You need 8 GB of RAM at the very least (16 GB or more is better) and 22 GB free on the disk for the smallest model. A graphics card is optional.

Windows

On GitHub click Code, then Download ZIP, unzip it and double-click START-HERE.bat. If Python is missing, it offers to install it for you.

Linux

Ubuntu and Debian; other distributions have the same packages under their own names.

sudo apt install git python3 build-essential git clone https://github.com/JustVugg/colibri cd colibri ./start-here.sh

macOS

With Homebrew.

xcode-select --install brew install libomp git python git clone https://github.com/JustVugg/colibri cd colibri ./start-here.sh
  1. It looks at your machine: RAM, free disk, CPU and GPUs.
  2. It recommends a model that fits, and Enter takes it. --list shows every model against your machine, --model ID picks another.
  3. It gets the engine: built for your machine when a compiler is there, or the prebuilt one from the release, with Vulkan built in on Linux and Windows. It uses your GPU when that pays: CUDA for an NVIDIA card on Linux with the CUDA toolkit, otherwise Vulkan; on an integrated GPU only for the models measured faster there. A missing package is printed as the exact command to run.
  4. It downloads the model with resume: stop it any time, run it again and it continues.
  5. It starts colibri and opens the dashboard, and prints the OpenAI and Anthropic addresses for other apps. Next time it starts straight away.

Or let your AI assistant do it

Ask your coding assistant the line below; it follows every step as a command and asks you before downloading anything. Assistants that speak the Model Context Protocol can use coli mcp instead.

Set up colibri on this machine following docs/AI_SETUP.md from https://github.com/JustVugg/colibri
If something goes wrong →

Models

The ones tagged setup are offered by the one-step setup, which picks among them by your RAM and disk. For the others, pull a converted container where one is published, or point the engine at the original checkpoint.

glm-5.2 / glm-5.3

The model colibri was built around. 78 layers, 256 routed experts, streamed from disk on a machine that could never hold them. GLM-5.3 is the same base model on the same engine, from its own container, which has no MTP head.

setupMoEint4MLAMTP
hf download mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp --local-dir ~/Models/glm52_i4
hf download Justvugg/GLM-5.3-colibri-int4-g64 --local-dir ~/Models/glm53_i4

glm-5.3-flash

Natively multimodal. Routed experts convert to int4-gs64, the dense set stays BF16 and its precision is a load-time choice. One resumable command downloads the 328 GB release and converts it a shard at a time.

MoEint4-gs64vision
321B total18B active convert, 195 GB int4 glm53.c
python3 c/tools/convert_glm53.py --outdir ~/Models/glm53_flash_i4 --min-free-gb 30

deepseek-v4.1-flash

Runs off the official checkpoint untouched: experts are already fp4, the dense set fp8-e4m3. Routed experts cost 4.5 GB per token against GLM-5.2's 12.7.

MoEfp4visiontoolsDSpark
552B total16B active official, no conversion deepseek_v41.c
hf download deepseek-ai/DeepSeek-V4.1-Flash --local-dir ~/Models/DeepSeek-V4.1-Flash

deepseek-v4-flash

Native fp4 experts, fp8-e4m3 dense. The REAP-pruned 150B cut, 132 of 256 experts and 85 GB, loads with the same engine and no conversion.

setupMoEfp4MTPREAP
284B total13B active official, no conversion deepseek_v4.c
hf download deepseek-ai/DeepSeek-V4-Flash-0731 --local-dir ~/Models/DeepSeek-V4-Flash

mimo-v2.6-flash

Xiaomi's hybrid: 39 of its 48 layers attend a 128-token window, so a long context costs the KV of 9. Experts stream as released in MXFP4, with no conversion. 2.34 to 3.37 tok/s on an 8-core Ryzen with 64 GB.

setupMoEMXFP4visiontools
309B total15B active official, no conversion mimo.c
hf download XiaomiMiMo/MiMo-V2.6-Flash-MOPD --local-dir ~/Models/MiMo-V2.6-Flash

mimo-v2.6-pro

The same engine and architecture as Flash at 70 layers and 384 experts. Checked against Xiaomi's own modeling code on 32 of its 70 layers; 0.66 to 0.79 tok/s on the same 64 GB box.

setupMoEMXFP4visiontools
1.02T total42B active official, no conversion mimo.c
hf download XiaomiMiMo/MiMo-V2.6-Pro-MOPD --local-dir ~/Models/MiMo-V2.6-Pro --exclude "dflash/*" --exclude "audio_tokenizer/*" --exclude "model_mtp.safetensors"

kimi-k3

The largest model colibri runs. 93 layers, 69 Kimi Delta Attention plus 24 gated MLA, and 16 of 896 experts per token. QAT-trained MXFP4 streams straight from the original shards.

setupMoEMXFP4KDAlargest
2.8T total104B active official, no conversion kimi_k3.c
hf download moonshotai/Kimi-K3 --local-dir ~/Models/Kimi-K3

inkling

Sliding-window attention with short convolutions. Audio input is supported when the audio tensors are present; the dense set fits a 25 GB host at 15.3 GB.

setupMoEint4audioApache 2.0
975B total41B active container, 514 GBupstream inkling.c
hf download nbeerbower/Inkling-colibri-int4 --local-dir ~/Models/inkling_i4

qwen3.8-flash-next

48 layers, 512 experts, 10 routed per token. 125B of backbone plus a 51B n-gram embedding that stays pageable and a 4B MTP head. The experts stay native block-FP8, or run as int4-g64 from an optional sidecar: 1.4-1.5x the decode speed at the same cache. Opt-in MTP drafting gives the same output, 12-14% faster.

setupMoEblock-FP8int4-g64DeltaNetPLEMTP
180B total6B active official, no conversion qwen38.c
hf download Qwen/Qwen3.8-Flash-Next-FP8 --local-dir ~/Models/Qwen3.8-Flash-Next-FP8

qwen3.6-35b-a3b

The one to start with: 23 GB on disk, gated attention interleaved with gated DeltaNet. 6.0 tok/s on an 8-core desktop CPU, 9.9 with Vulkan on its integrated Radeon 780M, and 30.0 with CUDA on an RTX 3070 with the DeltaNet layers on the card.

setupMoEint4-gs64DeltaNetCUDA tierVulkan
35B total3B active container, 23 GBupstream qwen36.c
hf download Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64 --local-dir ~/Models/qwen36_i4_gs64

qwen3.8-27b

Dense: one MLP per layer and no router, on the Qwen3.6 engine, text and images. Every weight is read for every token, so memory bandwidth sets the speed, not the disk. With every weight in int4 it needs 18.8 GB and decoded 3.45 tok/s on a 16-thread CPU server.

denseint8visiontools
27B totaldense, all active convert, 51 GB f16 qwen36.c
python3 c/tools/convert_qwen36.py --repo Qwen/Qwen3.8-27B --out ~/Models/qwen38_27b

qwen3-coder-30b-a3b

Qwen3-Coder on the Qwen3.6 engine: 48 attention layers, 128 experts with 8 per token, tool calls in its own XML form and no thinking. With all 128 experts cached the int4 container decoded at 8.5 to 9.6 tok/s on an 8-core Ryzen, in 15.2 GB.

setupMoEint4-gs64tools
30B total3B active container, 19 GBupstream qwen36.c
hf download Justvugg/Qwen3-Coder-30B-A3B-colibri-int4 --local-dir ~/Models/qwen3-coder-i4

olmoe

Fully open weights and data, 16 layers and 64 experts. The one to learn the tooling on: the whole container is 7 GB.

MoEint8open datasmall
7B total1B active convert, 7 GB int8 olmoe.c
python3 c/tools/convert_olmoe_merged.py --repo allenai/OLMoE-1B-7B-0125-Instruct --out ~/Models/olmoe_i8

qwen-image-2.1

Text to picture, read straight from the official diffusers checkpoint and quantized to int8 while loading. coli chat draws the image inside the terminal, coli serve answers the OpenAI images API. One 768x512 image peaks at 9.0 GB of RAM. Qwen Research License: non-commercial use only.

setupimageint8DiTnon-commercial
text to image official, 33 GB, no conversion qwenimage.c
hf download Qwen/Qwen-Image-2.1 --local-dir ~/Models/Qwen-Image-2.1

laya

A decision model: a state and typed questions in, a calibrated probability for every option out, in one forward pass. Served on POST /v1/systemone; 219 ms for one question on a laptop CPU, in 1.7 GB of RAM. English.

decisionModernBERTApache 2.0
hf download convaiinnovations/laya model.safetensors rl_agent_config.json "encoder/*" "tokenizer/*" --local-dir ~/Models/laya

gliner2.5-decide

A classifier for operational decisions: intent, routing, urgency, yes or no over a passage. All the questions of a request are read in one pass; 294 ms for one question on a busy laptop CPU, in 1.9 GB of RAM. English.

decisionDeBERTa-v3Apache 2.0
encoder + classifier official, 1.9 GB, no conversion gliner_decide.c
hf download fastino/GLiNER2.5-Decide model.safetensors config.json "encoder_config/*" tokenizer.json tokenizer_config.json special_tokens_map.json --local-dir ~/Models/GLiNER2.5-Decide

clef

Qwen3.8-27B post-trained with a joint schema head that scores every option of every question in one pass. POST /v1/systemone answers natively and the same process still chats. 20.4 s per request in int8 on an 8-core desktop CPU.

decisiondenseApache 2.0
27B total convert, 55 GB qwen36.c
python3 c/tools/convert_qwen36.py --model ~/Models/clef --out ~/Models/clef_c

Bigger than your RAM

A mixture-of-experts token touches a small fraction of the weights. colibri keeps that fraction resident and reads the rest as the router asks for it, so the model does not have to fit.

6x RTX 5090full residency9.0-9.2 tok/s
AWS Graviton4aarch64, CPU only8.0 tok/s
2x Xeon Gold 64301 TB DDR55.42 tok/s
NVIDIA DGX Spark121 GB unified3.33 tok/s
MacBook Pro M5 MaxMetal2.0 tok/s
Ryzen 9 9950X3DRTX 5090 + NVMe1.23 tok/s
25 GB dev boxcold0.05-0.1 tok/s

GLM-5.2 int4, decode speed reported by the people who ran it. The hardware changes where the experts live, not what the model answers.

A GPU is optional on every engine. With Vulkan, any GPU keeps the hottest experts and runs whole layers: on the integrated Radeon 780M of an 8-core desktop, Qwen3.6 went from 6.0 to 9.9 tok/s. A small model can lose on an integrated GPU, so the setup turns it on only where it was measured faster (coli setup --backend vulkan or --backend cpu decides for you); no discrete GPU has run the new expert tier yet. With CUDA, NVIDIA cards keep the hottest experts in VRAM. The release archives carry these backends ready to run: Vulkan on Linux and Windows, Metal on macOS, and a CUDA package for NVIDIA cards on Windows. Vulkan notes

See the benchmarks →

Frontier open models

Thirteen engines run today: ten for language models, one for pictures and two for decisions, one hand-written C file each. Not a wrapper around someone else's runtime.

kimi-k3104B active2.8T
mimo-v2.6-pro42B active1.02T
inkling41B active975B
glm-5.2 / glm-5.340B active744B
deepseek-v4.1-flash16B active552B
glm-5.3-flash18B active320B
mimo-v2.6-flash15B active309B
deepseek-v4-flash13B active284B
qwen3.8-flash-next6B active180B
qwen3.6-35b-a3b3B active35B
qwen3-coder-30b-a3b3B active30B
qwen3.8-27bdense, all active27B
olmoe1B active7B

Total parameters. Only the active ones are computed per token, and only the routed experts move. The engines, by the names the docs use: GLM-5.3-Flash, GLM-5.2/5.3, Inkling, Kimi K3, OLMoE, Qwen3.6-35B-A3B (which also runs Qwen3-Coder-30B-A3B and the dense Qwen3.8-27B), Qwen3.8-Flash-Next, DeepSeek V4 Flash, DeepSeek V4.1 Flash, MiMo-V2.6 Flash (and Pro, on the same engine). And for images, Qwen-Image-2.1: text to picture, drawn inside the terminal. And for decisions, Laya and GLiNER2.5-Decide, with Clef on the Qwen3.6 engine: typed questions in, a probability per option out. Every language engine serves up to 16 conversations at once, their next tokens decoded together.

Browse all models →

System One mode

Most of what people ask a model for is a choice, not an essay. Hand it the options it may pick and it reports the probability of each one. Nothing is generated, so there is nothing to parse and no answer outside the list. Typed decisions with calibrated probabilities on POST /v1/systemone, with the request and the reply of TypeSafe's Jev API: a Jev client switches by changing its base URL.

one question, 10 optionsSystem One101.2 s
chat125.6 s 1.2x
4 items, shared instructionsSystem One45.7 s
chat261.1 s 5.7x

Measured on qwen36 (22 GB), same engine, separate processes. Time from the prompt to the answer.

Triage and routing

Anything that ends in a bucket: a ticket to a queue, a pull request to a verdict, a message to a policy.

question what should the reviewer do?
options merge · request changes · close

Three decision models

Laya, GLiNER2.5-Decide and Clef answer the same request natively: one forward pass, no prompt and nothing generated. Any chat model answers it by scoring.

same API POST /v1/systemone
Laya 219 ms for one question on a laptop CPU

Many questions, one document

Read the contract once, then ask it twenty things. The first question pays for the reading, the rest cost their own words.

one snapshot 496 tokens for 4 items
two levels 176 tokens
POST /v1/systemone
{
  "state": "340 lines, 8 files, no tests.",
  "questions": {"review": {
    "type": "choice",
    "instructions": "What should the reviewer do?",
    "criteria": {"merge": null, "request changes": null, "close": null}
  }}
}
the reply
{
  "answers": {"review": {
    "choice": "request changes",
    "probabilities": {
      "merge": 0.023, "request changes": 0.974, "close": 0.004
    },
    "confidence": 0.961
  }},
  "usage": {"output_tokens": 4, "cost": 0}
}

And the thing generation cannot give

Same document, same model, three closed questions about it (measured when they were the fields of one JSON object, filled one at a time). On area it is sure and it is right. On the other two the entropy is near the maximum: it does not know, and it says so. The chat arm wrote "risk": "high" with exactly the same face.

areathe model is sure94.4% entropy 0.17
needs_testsit is not52% entropy 0.95
riskit is not46% entropy 0.99

That number is what you put a threshold on to send a case to a human. System One mode does not write prose, so anything open-ended is still the chat's job, and the speed multipliers above were measured on an engine reading experts from disk.

Read the System One guide →

Your weights stay yours

The engine is one C file per family with no runtime and no dependencies. It runs on a laptop with the network off, and the model files never leave the machine. Apache 2.0, and every container keeps the licence its authors gave it.

what you keep
The weightson your disk, in a format you can read
The promptsnever sent anywhere
The engineone file you can read in an afternoon
The numberstoken-exact against the reference implementation
Read the engine →

Get up and running on hardware you already own

Get started