Skip to content

Garbled/incoherent output from edge0-8b and edge0-35b on Apple A18 Pro (MacBook Neo), macOS 27 #8

Description

@msdyatt

Both edge0-8b and edge0-35b load fine — LoRA and prerouter install cleanly, generation runs to completion and reports a token count — but what comes out the other end is incoherent, mixed-language noise instead of a real answer. Nowhere near the benchmark numbers in the README (91.5 HumanEval / 70.1 MMLU-Pro for edge0-8b, for example).

Environment: MacBook Neo, Apple A18 Pro chip, macOS 27.0. I hit this on the beta build (26A5421a) and it's identical after updating to the public release (26A428), so it's not a beta quirk. Python 3.12.14 via Homebrew. Fresh clone, pip install -e '.[dev,fetch]' (pulls in mlx==0.30.4, mlx-metal==0.30.4, mlx-lm==0.31.0 as pinned). Both models pulled with scripts/fetch_models.py --tier all — no .incomplete files left over, shard sizes match what's on HF.

Repro:

export EDGE0_8B_MODEL=/path/to/models/edge0-8b
edge0 demo edge0-8b --max-new 30 --prompt "Hello! Write one short sentence about the seaside."

gives me:

[lora] applied=153 not_found=0 skipped=0 scale=2.0
[edge0-8b] prerouter installed: 16 heads, start=7, K=8
[edge0-8b] built in 0.8s
# 30 tokens in 22.278s
user : Hello! Write one short sentence about the seaside.
edge0: ! sharper impacting工作者在图AvecAw quen asian-img平米 axes성这首歌Whole。作为 triang telchina考试pons光纤ząd洞天_ENTER的新闻知识 endpointolated直气

Same with the built-in Chinese demo prompt. Same with --no-lora --no-prerouter (plain int4 base, no adapters), so it's not an adapter problem. And it happens on edge0-35b too, just a different flavor of garbage (repeated ```python fences mixed into the noise), so it's not tied to one model either.

A few things I checked along the way: the download isn't corrupt (verified shard sizes against HF after the fact). LoRA/prerouter both report clean installs. Basic MLX ops work fine — mx.default_device() gives Device(gpu, 0) and a plain matmul on GPU comes back correct, so Metal itself isn't broken.

I also tried bumping mlx/mlx-metal to 0.31.2 and then 0.32.2, keeping mlx-lm pinned at 0.31.0 like the README says to. Both make it worse — outright crash instead of garbage:

File ".../edge0/streaming/layer.py", line 1310, in __call__
    unique = sorted(set(int(v) for v in flat.tolist()))
                                        ^^^^^^^^^^^^^
RuntimeError: There is no Stream(gpu, 4) in current thread.

Looks like the streaming/expert-offload code depends on some MLX Stream behavior that changed after 0.30.4 — maybe the same category of issue the README already calls out for mlx-lm.

Any idea if this is a known incompatibility with this chip, or is there a setting I'm missing to get real output out of either tier? Can dig up more logs if useful.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions