-
-
Notifications
You must be signed in to change notification settings - Fork 21.9k
Pull requests: vllm-project/vllm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[New Model][Nvidia] Add SM12x support for DeepSeek V4 Flash with essential fixes
ci/build
deepseek
Related to DeepSeek models
frontend
kv-connector
mrv2
Model Runner V2 specific
needs-rebase
new-model
Requests to new models
nvidia
quantization
speculative-decoding
structured-output
tool-calling
v1
#41834
opened May 6, 2026 by
jasl
Contributor
Loading…
[Core] Extensible (growable) KV cache
ci/build
cohere
Related to Cohere models
cpu
Related to CPU backends
deepseek
Related to DeepSeek models
documentation
Improvements or additions to documentation
DSv4
frontend
glm
gpt-oss
Related to GPT-OSS models
inkling
intel-gpu
Related to Intel GPU
k3
kimi
kv-cache-manager
kv-connector
llama
Related to Llama models
minimax
mistral
Related to Mistral models
mrv2
Model Runner V2 specific
multi-modality
Related to multi-modality (#4194)
needs-rebase
nvidia
performance
Performance-related issues
quantization
qwen
Related to Qwen models
ray
anything related with ray
ready
ONLY add when PR is ready to merge/full CI is needed
rocm
Related to AMD ROCm
speculative-decoding
structured-output
tool-calling
torch.compile
Fix Mamba copy metadata high-bit pointers
needs-rebase
v1
#41995
opened May 7, 2026 by
dipenbhuva
Loading…
[Spec Decode][Hybrid] Add ngram-eagle SD method
documentation
Improvements or additions to documentation
needs-rebase
performance
Performance-related issues
speculative-decoding
v1
#24344
opened Sep 5, 2025 by
ekagra-ranjan
Contributor
Loading…
3 tasks done
Add DeepSeek-V4 DCP decode support
deepseek
Related to DeepSeek models
DSv4
kv-cache-manager
needs-rebase
nvidia
v1
#44573
opened Jun 4, 2026 by
baonudesifeizhai
Contributor
Loading…
4 tasks
[MoE][Offload] Run MoE models exceeding VRAM via expert CPU offloading with GPU cache (--moe-expert-cache-size)
ci/build
documentation
Improvements or additions to documentation
frontend
performance
Performance-related issues
quantization
verified
Run pre-commit for new contributors without triggering other tests
#37190
opened Mar 16, 2026 by
e1n00r
Loading…
[ROCm] Fix two RDNA4 correctness bugs in the LLMM1 skinny-GEMM kernel
performance
Performance-related issues
ready
ONLY add when PR is ready to merge/full CI is needed
rocm
Related to AMD ROCm
#40827
opened Apr 24, 2026 by
wjabbour
Contributor
Loading…
fix(kv-cache): allow TurboQuant on hybrid models
quantization
v1
#41123
opened Apr 28, 2026 by
cderinbogaz
Loading…
8 of 9 tasks
[EC Connector] Mooncake EC Connector for Distributed Encoder-Cache Transfer
documentation
Improvements or additions to documentation
frontend
kv-connector
needs-rebase
v1
#40695
opened Apr 23, 2026 by
fake0fan
Contributor
Loading…
[ROCm][Bugfix] Densify strided activations for wvSplitKQ
bug
Something isn't working
rocm
Related to AMD ROCm
#50618
opened Jul 31, 2026 by
JohnQinAMD
Contributor
Loading…
feat(openai): add per-request timing metrics and completion_tokens_de…
documentation
Improvements or additions to documentation
frontend
needs-rebase
[Bugfix] Handle None compilation times in v1 Executor initialize_from_config
bug
Something isn't working
needs-rebase
stale
Over 90 days of inactivity
v1
#37628
opened Mar 20, 2026 by
ec-jt
Loading…
[Feature] Support batch invariance on ROCm
ci/build
needs-rebase
nvidia
quantization
rocm
Related to AMD ROCm
#52231
opened Aug 14, 2026 by
mawong-amd
Contributor
Loading…
4 tasks
feat: FlashInfer CuteDSL MegaMoE integration
deepseek
Related to DeepSeek models
DSv4
mrv2
Model Runner V2 specific
nvidia
quantization
#54049
opened Aug 27, 2026 by
jdebache
Contributor
Loading…
[Bugfix][AsyncScheduler] Reserve placeholders for the final prefill chunk of a streaming update
bug
Something isn't working
needs-rebase
scheduler
streaming-input
#50692
opened Aug 1, 2026 by
Aayush7352
Loading…
[ROCm][AITER] Add support for new per-stream workspace requirement by AITER v0.1.20
mrv2
Model Runner V2 specific
needs-rebase
nvidia
rocm
Related to AMD ROCm
[ROCm] optimize memory for GLM 5.2
glm
needs-rebase
rocm
Related to AMD ROCm
v1
#50459
opened Jul 30, 2026 by
Concurrensee
Contributor
•
Draft
[Rust Frontend] Harden resumable session lifecycle in engine-core-client
rust
streaming-input
#53454
opened Aug 23, 2026 by
TanNgocDo
Contributor
Loading…
2 tasks done
[Perf][PCP] Shard decode requests across PCP ranks
deepseek
Related to DeepSeek models
mrv2
Model Runner V2 specific
needs-rebase
ready
ONLY add when PR is ready to merge/full CI is needed
#52162
opened Aug 13, 2026 by
pisceskkk
Contributor
Loading…
4 tasks done
[Bugfix][Spec Decode] DFlash2: accept unquantized linear LM heads in the candidate selector
bug
Something isn't working
dflash
mrv2
Model Runner V2 specific
needs-rebase
new-model
Requests to new models
quantization
qwen
Related to Qwen models
speculative-decoding
#52883
opened Aug 19, 2026 by
oceanplexian
Loading…
[ROCm][aiter] officially supporting aiter Triton kernels+dsv4 on RDNA3
DSv4
quantization
rocm
Related to AMD ROCm
torch.compile
#52970
opened Aug 19, 2026 by
amd-xavierwang
Loading…
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.