Skip to content

Pull requests: vllm-project/vllm

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

[Core] Extensible (growable) KV cache ci/build cohere Related to Cohere models cpu Related to CPU backends deepseek Related to DeepSeek models documentation Improvements or additions to documentation DSv4 frontend glm gpt-oss Related to GPT-OSS models inkling intel-gpu Related to Intel GPU k3 kimi kv-cache-manager kv-connector llama Related to Llama models minimax mistral Related to Mistral models mrv2 Model Runner V2 specific multi-modality Related to multi-modality (#4194) needs-rebase nvidia performance Performance-related issues quantization qwen Related to Qwen models ray anything related with ray ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm speculative-decoding structured-output tool-calling torch.compile
#50779 opened Aug 2, 2026 by njhill Member Draft
[Spec Decode][Hybrid] Add ngram-eagle SD method documentation Improvements or additions to documentation needs-rebase performance Performance-related issues speculative-decoding v1
#24344 opened Sep 5, 2025 by ekagra-ranjan Contributor Loading…
3 tasks done
Add DeepSeek-V4 DCP decode support deepseek Related to DeepSeek models DSv4 kv-cache-manager needs-rebase nvidia v1
#44573 opened Jun 4, 2026 by baonudesifeizhai Contributor Loading…
4 tasks
[MoE][Offload] Run MoE models exceeding VRAM via expert CPU offloading with GPU cache (--moe-expert-cache-size) ci/build documentation Improvements or additions to documentation frontend performance Performance-related issues quantization verified Run pre-commit for new contributors without triggering other tests
#37190 opened Mar 16, 2026 by e1n00r Loading…
[Perf][Kernel][ROCm] Add Triton split-KV paged decode fallback for gfx12 mrv1-only Issues/PRs which apply only to Model Runner V1 (not applicable to Model Runner V2) rocm Related to AMD ROCm v1
#45916 opened Jun 17, 2026 by feiyehua Loading…
4 tasks done
[ROCm] Fix two RDNA4 correctness bugs in the LLMM1 skinny-GEMM kernel performance Performance-related issues ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm
#40827 opened Apr 24, 2026 by wjabbour Contributor Loading…
[ROCm][Bugfix] Densify strided activations for wvSplitKQ bug Something isn't working rocm Related to AMD ROCm
#50618 opened Jul 31, 2026 by JohnQinAMD Contributor Loading…
[Bugfix] Handle None compilation times in v1 Executor initialize_from_config bug Something isn't working needs-rebase stale Over 90 days of inactivity v1
#37628 opened Mar 20, 2026 by ec-jt Loading…
feat: FlashInfer CuteDSL MegaMoE integration deepseek Related to DeepSeek models DSv4 mrv2 Model Runner V2 specific nvidia quantization
#54049 opened Aug 27, 2026 by jdebache Contributor Loading…
[ROCm][AITER] Add support for new per-stream workspace requirement by AITER v0.1.20 mrv2 Model Runner V2 specific needs-rebase nvidia rocm Related to AMD ROCm
#54095 opened Aug 27, 2026 by qli88 Contributor Draft
[ROCm] optimize memory for GLM 5.2 glm needs-rebase rocm Related to AMD ROCm v1
#50459 opened Jul 30, 2026 by Concurrensee Contributor Draft
[Perf][PCP] Shard decode requests across PCP ranks deepseek Related to DeepSeek models mrv2 Model Runner V2 specific needs-rebase ready ONLY add when PR is ready to merge/full CI is needed
#52162 opened Aug 13, 2026 by pisceskkk Contributor Loading…
4 tasks done
[Bugfix][Spec Decode] DFlash2: accept unquantized linear LM heads in the candidate selector bug Something isn't working dflash mrv2 Model Runner V2 specific needs-rebase new-model Requests to new models quantization qwen Related to Qwen models speculative-decoding
#52883 opened Aug 19, 2026 by oceanplexian Loading…
[Bugfix] Format kernel-import errors eagerly so warning_once does not retain them bug Something isn't working cpu Related to CPU backends nvidia
#54098 opened Aug 27, 2026 by zwang86 Loading…
ProTip! Add no:assignee to see everything that’s not assigned.