-
-
Notifications
You must be signed in to change notification settings - Fork 21.9k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Bugfix][Quantization] Match compressed-tensors MoE target aliases
bug
Something isn't working
quantization
#55814
opened Sep 8, 2026 by
Kangwenqiao
•
Draft
[BugFix] GLM-5: load msmodelslim-native checkpoints and drop MTP draft weights
bug
Something isn't working
glm
#55813
opened Sep 8, 2026 by
Liears
Loading…
[Metrics] Add admission rejection metrics
documentation
Improvements or additions to documentation
#55812
opened Sep 8, 2026 by
thuongvu
Loading…
[ROCm][Kimi-K3] Add large-M merged MoE front
k3
kimi
rocm
Related to AMD ROCm
#55811
opened Sep 8, 2026 by
jiacao-amd
Contributor
•
Draft
[BugFix] Group GLM-5 hybrid KV cache specs (indexer/state roles) correctly
bug
Something isn't working
glm
kv-cache-manager
#55810
opened Sep 8, 2026 by
Liears
Loading…
[Bench] Save server queue time and TTFT in detailed results
documentation
Improvements or additions to documentation
performance
Performance-related issues
rust
#55809
opened Sep 8, 2026 by
lzzzzzc
Loading…
4 tasks done
[ROCm][Perf] Remove AITER paged-MQA outputs guard for DeepSeek-V4
deepseek
Related to DeepSeek models
DSv4
rocm
Related to AMD ROCm
#55808
opened Sep 8, 2026 by
shen-shanshan
Collaborator
•
Draft
4 tasks
[Docs] Add Qwen3.5 text CausalLM to supported models
documentation
Improvements or additions to documentation
qwen
Related to Qwen models
#55807
opened Sep 8, 2026 by
imitater-dou
•
Draft
4 tasks done
[BugFix] Flatten tuple-valued layer caches in copy_kv_cache_blocks_inplace
bug
Something isn't working
#55806
opened Sep 8, 2026 by
Liears
Loading…
[Docs] Add DiffusionGemma, LongcatFlashNgram, and Step3Text to supported models
documentation
Improvements or additions to documentation
#55805
opened Sep 8, 2026 by
imitater-dou
•
Draft
[Bugfix][EPLB] Use group-local ranks for torch P2P transfers
bug
Something isn't working
#55804
opened Sep 8, 2026 by
LostFox11
Loading…
[Docs] Add FunAudioChatForConditionalGeneration to supported models
documentation
Improvements or additions to documentation
#55803
opened Sep 8, 2026 by
imitater-dou
•
Draft
[Bugfix][KV Connector] Apply KV recompute threshold to NIXL remote-prefill pulls
bug
Something isn't working
documentation
Improvements or additions to documentation
kv-connector
#55801
opened Sep 8, 2026 by
LPMing
Loading…
[Bugfix][CI] skip conftest for NPU compatibility test
bug
Something isn't working
ci/build
#55799
opened Sep 8, 2026 by
yzeyu71
Contributor
Loading…
1 of 6 tasks
[LoRA] Support PEFT trainable token rows alongside LoRA adapters
documentation
Improvements or additions to documentation
#55796
opened Sep 8, 2026 by
zupengwang
Contributor
Loading…
[Docs] Add LagunaForCausalLM to supported models
documentation
Improvements or additions to documentation
#55795
opened Sep 8, 2026 by
imitater-dou
•
Draft
4 tasks done
[Docs] Add NemotronH_Nano_VL_V2 to supported models
documentation
Improvements or additions to documentation
#55794
opened Sep 8, 2026 by
imitater-dou
•
Draft
3 tasks
[Docs] Add Inkling to supported models
documentation
Improvements or additions to documentation
inkling
#55793
opened Sep 8, 2026 by
imitater-dou
•
Draft
4 tasks
[Docs] Add NemotronParse to supported models
documentation
Improvements or additions to documentation
#55791
opened Sep 8, 2026 by
imitater-dou
•
Draft
3 tasks
[Bugfix][Frontend] Clean up child tasks when request wrappers are cancelled
bug
Something isn't working
frontend
#55790
opened Sep 8, 2026 by
git-jxj
Loading…
Revert "[Bugfix][Gemma] Conditionally create KV projections/norms on KV-shared layers" (#54917)
bug
Something isn't working
#55789
opened Sep 8, 2026 by
vllm-agent
Contributor
•
Draft
[Bugfix][Tool Parser] Stream all parallel hy_v4 tool calls arriving in one delta
bug
Something isn't working
tool-calling
#55788
opened Sep 8, 2026 by
nikhilkulkarni1755
Contributor
Loading…
[Bugfix] FlashInfer derive CUDA-graph head count from KV group layers
bug
Something isn't working
ci/build
documentation
Improvements or additions to documentation
frontend
kv-cache-manager
kv-connector
needs-rebase
new-model
Requests to new models
nvidia
performance
Performance-related issues
qwen
Related to Qwen models
scheduler
speculative-decoding
tool-calling
#55787
opened Sep 8, 2026 by
XFDG
Loading…
[Bugfix][Frontend] Apply the model's transcription post-processing on the realtime path
bug
Something isn't working
frontend
#55786
opened Sep 8, 2026 by
twu3202
Loading…
3 of 4 tasks
Previous Next
ProTip!
Type g p on any issue or pull request to go back to the pull request listing page.