Both edge0-8b and edge0-35b load fine — LoRA and prerouter install cleanly, generation runs to completion and reports a token count — but what comes out the other end is incoherent, mixed-language noise instead of a real answer. Nowhere near the benchmark numbers in the README (91.5 HumanEval / 70.1 MMLU-Pro for edge0-8b, for example).
Environment: MacBook Neo, Apple A18 Pro chip, macOS 27.0. I hit this on the beta build (26A5421a) and it's identical after updating to the public release (26A428), so it's not a beta quirk. Python 3.12.14 via Homebrew. Fresh clone, pip install -e '.[dev,fetch]' (pulls in mlx==0.30.4, mlx-metal==0.30.4, mlx-lm==0.31.0 as pinned). Both models pulled with scripts/fetch_models.py --tier all — no .incomplete files left over, shard sizes match what's on HF.
Repro:
export EDGE0_8B_MODEL=/path/to/models/edge0-8b
edge0 demo edge0-8b --max-new 30 --prompt "Hello! Write one short sentence about the seaside."
gives me:
[lora] applied=153 not_found=0 skipped=0 scale=2.0
[edge0-8b] prerouter installed: 16 heads, start=7, K=8
[edge0-8b] built in 0.8s
# 30 tokens in 22.278s
user : Hello! Write one short sentence about the seaside.
edge0: ! sharper impacting工作者在图AvecAw quen asian-img平米 axes성这首歌Whole。作为 triang telchina考试pons光纤ząd洞天_ENTER的新闻知识 endpointolated直气
Same with the built-in Chinese demo prompt. Same with --no-lora --no-prerouter (plain int4 base, no adapters), so it's not an adapter problem. And it happens on edge0-35b too, just a different flavor of garbage (repeated ```python fences mixed into the noise), so it's not tied to one model either.
A few things I checked along the way: the download isn't corrupt (verified shard sizes against HF after the fact). LoRA/prerouter both report clean installs. Basic MLX ops work fine — mx.default_device() gives Device(gpu, 0) and a plain matmul on GPU comes back correct, so Metal itself isn't broken.
I also tried bumping mlx/mlx-metal to 0.31.2 and then 0.32.2, keeping mlx-lm pinned at 0.31.0 like the README says to. Both make it worse — outright crash instead of garbage:
File ".../edge0/streaming/layer.py", line 1310, in __call__
unique = sorted(set(int(v) for v in flat.tolist()))
^^^^^^^^^^^^^
RuntimeError: There is no Stream(gpu, 4) in current thread.
Looks like the streaming/expert-offload code depends on some MLX Stream behavior that changed after 0.30.4 — maybe the same category of issue the README already calls out for mlx-lm.
Any idea if this is a known incompatibility with this chip, or is there a setting I'm missing to get real output out of either tier? Can dig up more logs if useful.
Both edge0-8b and edge0-35b load fine — LoRA and prerouter install cleanly, generation runs to completion and reports a token count — but what comes out the other end is incoherent, mixed-language noise instead of a real answer. Nowhere near the benchmark numbers in the README (91.5 HumanEval / 70.1 MMLU-Pro for edge0-8b, for example).
Environment: MacBook Neo, Apple A18 Pro chip, macOS 27.0. I hit this on the beta build (26A5421a) and it's identical after updating to the public release (26A428), so it's not a beta quirk. Python 3.12.14 via Homebrew. Fresh clone,
pip install -e '.[dev,fetch]'(pulls in mlx==0.30.4, mlx-metal==0.30.4, mlx-lm==0.31.0 as pinned). Both models pulled withscripts/fetch_models.py --tier all— no.incompletefiles left over, shard sizes match what's on HF.Repro:
gives me:
Same with the built-in Chinese demo prompt. Same with
--no-lora --no-prerouter(plain int4 base, no adapters), so it's not an adapter problem. And it happens on edge0-35b too, just a different flavor of garbage (repeated```pythonfences mixed into the noise), so it's not tied to one model either.A few things I checked along the way: the download isn't corrupt (verified shard sizes against HF after the fact). LoRA/prerouter both report clean installs. Basic MLX ops work fine —
mx.default_device()givesDevice(gpu, 0)and a plain matmul on GPU comes back correct, so Metal itself isn't broken.I also tried bumping mlx/mlx-metal to 0.31.2 and then 0.32.2, keeping mlx-lm pinned at 0.31.0 like the README says to. Both make it worse — outright crash instead of garbage:
Looks like the streaming/expert-offload code depends on some MLX Stream behavior that changed after 0.30.4 — maybe the same category of issue the README already calls out for mlx-lm.
Any idea if this is a known incompatibility with this chip, or is there a setting I'm missing to get real output out of either tier? Can dig up more logs if useful.