-
Notifications
You must be signed in to change notification settings - Fork 1.2k
Pull requests: FlashML-org/FreeToken
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
feat(qwen3_5_moe): keep a natively-fp8 lm_head native instead of failing to load
#416
opened Sep 8, 2026 by
salekseev
Loading…
fix(qwen3_5_moe): accept a per-channel fp8 weight_scale in the dense loader
#415
opened Sep 8, 2026 by
salekseev
Loading…
fix(kernel): order the batch-memcpy probe against the current stream
#414
opened Sep 8, 2026 by
tspeaks
Loading…
fix(models): support llm-compressor NVFP4 MoE export variants
#413
opened Sep 8, 2026 by
Romeriz
Loading…
fix(server): honor configured output default across APIs
#411
opened Sep 7, 2026 by
earlvanze
Loading…
fix(models): serve qwen4_exp FTW checkpoints with PLE streamed from source
#405
opened Sep 7, 2026 by
Eng-Ahmd
Loading…
fix(models): support nvidia/Qwen3.8-Flash-Next-NVFP4 in qwen4_exp
#404
opened Sep 7, 2026 by
Eng-Ahmd
Loading…
fix(cpu-moe): honor padded fp8 scale strides, add avx512f tier and float64 parity (builds on #36)
#399
opened Sep 6, 2026 by
ChenyuHeee
Loading…
feat(kvcache): add ISO3/ISO4 KV-cache quantization (--kv-cache-iso)
#398
opened Sep 5, 2026 by
AsmanovLev
Loading…
qwen4_exp: serve the block-FP8 dense projections natively (+25% decode)
#392
opened Sep 5, 2026 by
gberasmus87
Loading…
fix(models): better support for mixed-precision compressed-tensors NVFP4
#390
opened Sep 5, 2026 by
Sam-Izdat
Loading…
fix(scheduler): reserve paged KV at allocation granularity
#367
opened Sep 3, 2026 by
taking-lying-flat
Contributor
Loading…
Previous Next
ProTip!
Filter pull requests by the default branch with base:main.