Windows
On GitHub click Code, then Download ZIP, unzip it and double-click
START-HERE.bat. If Python is missing, it offers to install it for you.
colibri streams a model's experts off the disk instead of loading them, so a 2.8-trillion-parameter model answers on a machine that could never hold it.
coli chat attached to a running
server, then the same model in System One mode. The numbers are the ones it produced.Measured today on machines people actually have
One file looks at your machine, recommends a model that fits, gets the engine, downloads the model with resume and opens the dashboard. You need 8 GB of RAM at the very least (16 GB or more is better) and 22 GB free on the disk for the smallest model. A graphics card is optional.
On GitHub click Code, then Download ZIP, unzip it and double-click
START-HERE.bat. If Python is missing, it offers to install it for you.
Ubuntu and Debian; other distributions have the same packages under their own names.
sudo apt install git python3 build-essential
git clone https://github.com/JustVugg/colibri
cd colibri
./start-here.shWith Homebrew.
xcode-select --install
brew install libomp git python
git clone https://github.com/JustVugg/colibri
cd colibri
./start-here.sh--list shows every model against your machine, --model ID picks another.Ask your coding assistant the line below; it follows every step as a command and
asks you before downloading anything. Assistants that speak the Model Context Protocol can
use coli mcp instead.
Set up colibri on this machine following docs/AI_SETUP.md from https://github.com/JustVugg/colibriThe ones tagged setup are offered by the one-step setup, which picks among them by your RAM and disk. For the others, pull a converted container where one is published, or point the engine at the original checkpoint.
The model colibri was built around. 78 layers, 256 routed experts, streamed from disk on a machine that could never hold them. GLM-5.3 is the same base model on the same engine, from its own container, which has no MTP head.
hf download mastouri/GLM-5.2-colibri-int4-g64-with-int8-mtp --local-dir ~/Models/glm52_i4hf download Justvugg/GLM-5.3-colibri-int4-g64 --local-dir ~/Models/glm53_i4Natively multimodal. Routed experts convert to int4-gs64, the dense set stays BF16 and its precision is a load-time choice. One resumable command downloads the 328 GB release and converts it a shard at a time.
python3 c/tools/convert_glm53.py --outdir ~/Models/glm53_flash_i4 --min-free-gb 30Runs off the official checkpoint untouched: experts are already fp4, the dense set fp8-e4m3. Routed experts cost 4.5 GB per token against GLM-5.2's 12.7.
hf download deepseek-ai/DeepSeek-V4.1-Flash --local-dir ~/Models/DeepSeek-V4.1-FlashNative fp4 experts, fp8-e4m3 dense. The REAP-pruned 150B cut, 132 of 256 experts and 85 GB, loads with the same engine and no conversion.
hf download deepseek-ai/DeepSeek-V4-Flash-0731 --local-dir ~/Models/DeepSeek-V4-FlashXiaomi's hybrid: 39 of its 48 layers attend a 128-token window, so a long context costs the KV of 9. Experts stream as released in MXFP4, with no conversion. 2.34 to 3.37 tok/s on an 8-core Ryzen with 64 GB.
hf download XiaomiMiMo/MiMo-V2.6-Flash-MOPD --local-dir ~/Models/MiMo-V2.6-FlashThe same engine and architecture as Flash at 70 layers and 384 experts. Checked against Xiaomi's own modeling code on 32 of its 70 layers; 0.66 to 0.79 tok/s on the same 64 GB box.
hf download XiaomiMiMo/MiMo-V2.6-Pro-MOPD --local-dir ~/Models/MiMo-V2.6-Pro --exclude "dflash/*" --exclude "audio_tokenizer/*" --exclude "model_mtp.safetensors"The largest model colibri runs. 93 layers, 69 Kimi Delta Attention plus 24 gated MLA, and 16 of 896 experts per token. QAT-trained MXFP4 streams straight from the original shards.
hf download moonshotai/Kimi-K3 --local-dir ~/Models/Kimi-K3Sliding-window attention with short convolutions. Audio input is supported when the audio tensors are present; the dense set fits a 25 GB host at 15.3 GB.
hf download nbeerbower/Inkling-colibri-int4 --local-dir ~/Models/inkling_i448 layers, 512 experts, 10 routed per token. 125B of backbone plus a 51B n-gram embedding that stays pageable and a 4B MTP head. The experts stay native block-FP8, or run as int4-g64 from an optional sidecar: 1.4-1.5x the decode speed at the same cache. Opt-in MTP drafting gives the same output, 12-14% faster.
hf download Qwen/Qwen3.8-Flash-Next-FP8 --local-dir ~/Models/Qwen3.8-Flash-Next-FP8The one to start with: 23 GB on disk, gated attention interleaved with gated DeltaNet. 6.0 tok/s on an 8-core desktop CPU, 9.9 with Vulkan on its integrated Radeon 780M, and 30.0 with CUDA on an RTX 3070 with the DeltaNet layers on the card.
hf download Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64 --local-dir ~/Models/qwen36_i4_gs64Dense: one MLP per layer and no router, on the Qwen3.6 engine, text and images. Every weight is read for every token, so memory bandwidth sets the speed, not the disk. With every weight in int4 it needs 18.8 GB and decoded 3.45 tok/s on a 16-thread CPU server.
python3 c/tools/convert_qwen36.py --repo Qwen/Qwen3.8-27B --out ~/Models/qwen38_27bQwen3-Coder on the Qwen3.6 engine: 48 attention layers, 128 experts with 8 per token, tool calls in its own XML form and no thinking. With all 128 experts cached the int4 container decoded at 8.5 to 9.6 tok/s on an 8-core Ryzen, in 15.2 GB.
hf download Justvugg/Qwen3-Coder-30B-A3B-colibri-int4 --local-dir ~/Models/qwen3-coder-i4Fully open weights and data, 16 layers and 64 experts. The one to learn the tooling on: the whole container is 7 GB.
python3 c/tools/convert_olmoe_merged.py --repo allenai/OLMoE-1B-7B-0125-Instruct --out ~/Models/olmoe_i8Text to picture, read straight from the official diffusers checkpoint and quantized to int8 while loading. coli chat draws the image inside the terminal, coli serve answers the OpenAI images API. One 768x512 image peaks at 9.0 GB of RAM. Qwen Research License: non-commercial use only.
hf download Qwen/Qwen-Image-2.1 --local-dir ~/Models/Qwen-Image-2.1A decision model: a state and typed questions in, a calibrated probability for every option out, in one forward pass. Served on POST /v1/systemone; 219 ms for one question on a laptop CPU, in 1.7 GB of RAM. English.
hf download convaiinnovations/laya model.safetensors rl_agent_config.json "encoder/*" "tokenizer/*" --local-dir ~/Models/layaA classifier for operational decisions: intent, routing, urgency, yes or no over a passage. All the questions of a request are read in one pass; 294 ms for one question on a busy laptop CPU, in 1.9 GB of RAM. English.
hf download fastino/GLiNER2.5-Decide model.safetensors config.json "encoder_config/*" tokenizer.json tokenizer_config.json special_tokens_map.json --local-dir ~/Models/GLiNER2.5-DecideQwen3.8-27B post-trained with a joint schema head that scores every option of every question in one pass. POST /v1/systemone answers natively and the same process still chats. 20.4 s per request in int8 on an 8-core desktop CPU.
python3 c/tools/convert_qwen36.py --model ~/Models/clef --out ~/Models/clef_cNothing matches that.
A mixture-of-experts token touches a small fraction of the weights. colibri keeps that fraction resident and reads the rest as the router asks for it, so the model does not have to fit.
GLM-5.2 int4, decode speed reported by the people who ran it. The hardware changes where the experts live, not what the model answers.
A GPU is optional on every engine. With Vulkan, any GPU keeps the hottest
experts and runs whole layers: on the integrated Radeon 780M of an 8-core desktop,
Qwen3.6 went from 6.0 to 9.9 tok/s. A small model can lose on an integrated GPU, so the
setup turns it on only where it was measured faster (coli setup --backend vulkan
or --backend cpu decides for you); no discrete GPU has run the new
expert tier yet. With CUDA, NVIDIA cards keep the hottest experts in VRAM. The release
archives carry these backends ready to run: Vulkan on Linux and Windows, Metal on macOS,
and a CUDA package for NVIDIA cards on Windows.
Vulkan notes
Thirteen engines run today: ten for language models, one for pictures and two for decisions, one hand-written C file each. Not a wrapper around someone else's runtime.
Total parameters. Only the active ones are computed per token, and only the routed experts move. The engines, by the names the docs use: GLM-5.3-Flash, GLM-5.2/5.3, Inkling, Kimi K3, OLMoE, Qwen3.6-35B-A3B (which also runs Qwen3-Coder-30B-A3B and the dense Qwen3.8-27B), Qwen3.8-Flash-Next, DeepSeek V4 Flash, DeepSeek V4.1 Flash, MiMo-V2.6 Flash (and Pro, on the same engine). And for images, Qwen-Image-2.1: text to picture, drawn inside the terminal. And for decisions, Laya and GLiNER2.5-Decide, with Clef on the Qwen3.6 engine: typed questions in, a probability per option out. Every language engine serves up to 16 conversations at once, their next tokens decoded together.
Browse all models →Most of what people ask a model for is a choice, not an essay. Hand it the
options it may pick and it reports the probability of each one. Nothing is generated,
so there is nothing to parse and no answer outside the list. Typed decisions with
calibrated probabilities on POST /v1/systemone, with the request and the reply
of TypeSafe's Jev API: a Jev client switches by changing its base URL.
Measured on qwen36 (22 GB), same engine, separate processes. Time from the prompt to the answer.
Anything that ends in a bucket: a ticket to a queue, a pull request to a verdict, a message to a policy.
Laya, GLiNER2.5-Decide and Clef answer the same request natively: one forward pass, no prompt and nothing generated. Any chat model answers it by scoring.
Read the contract once, then ask it twenty things. The first question pays for the reading, the rest cost their own words.
{
"state": "340 lines, 8 files, no tests.",
"questions": {"review": {
"type": "choice",
"instructions": "What should the reviewer do?",
"criteria": {"merge": null, "request changes": null, "close": null}
}}
}{
"answers": {"review": {
"choice": "request changes",
"probabilities": {
"merge": 0.023, "request changes": 0.974, "close": 0.004
},
"confidence": 0.961
}},
"usage": {"output_tokens": 4, "cost": 0}
}Same document, same model, three closed
questions about it (measured when they were the fields of one JSON object, filled one
at a time). On area it is sure and it is right. On the other two the entropy is
near the maximum: it does not know, and it says so. The chat arm wrote
"risk": "high"
with exactly the same face.
That number is what you put a threshold on to send a case to a human. System One mode does not write prose, so anything open-ended is still the chat's job, and the speed multipliers above were measured on an engine reading experts from disk.
Read the System One guide →The engine is one C file per family with no runtime and no dependencies. It runs on a laptop with the network off, and the model files never leave the machine. Apache 2.0, and every container keeps the licence its authors gave it.
| what you keep | |
|---|---|
| The weights | on your disk, in a format you can read |
| The prompts | never sent anywhere |
| The engine | one file you can read in an afternoon |
| The numbers | token-exact against the reference implementation |