Skip to content
View gavinkvx's full-sized avatar

Block or report gavinkvx

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Showing results

A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM

Python 823 217 Updated Sep 9, 2026

A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations

Python 19,486 1,571 Updated Sep 9, 2026

Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling

Go 263 83 Updated Sep 9, 2026

CUDA Templates and Python DSLs for High-Performance Linear Algebra

C++ 10,406 2,076 Updated Sep 8, 2026

cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

Python 933 278 Updated Sep 10, 2026

NVIDIA Inference Xfer Library (NIXL)

C++ 1,242 436 Updated Sep 9, 2026

A programmable Mixture-of-Models router for heterogeneous LLM inference

Go 5,701 928 Updated Sep 9, 2026

A Datacenter Scale Distributed Inference Serving Framework

Rust 1 Updated Sep 1, 2026

Building the Virtuous Cycle for AI-driven LLM Systems

Python 281 48 Updated May 1, 2026

A Datacenter Scale Distributed Inference Serving Framework

Rust 8,008 1,562 Updated Sep 10, 2026

FlashInfer: Kernel Library for LLM Serving

Python 1 Updated Sep 3, 2026

slime is an LLM post-training framework for RL Scaling.

Python 8,426 1,235 Updated Sep 3, 2026

FlashInfer: Kernel Library for LLM Serving

Cuda 6,362 1,414 Updated Sep 10, 2026

An LLM post-training framework with vLLM for RL Scaling

Python 456 90 Updated Sep 7, 2026
1 Updated Aug 19, 2026

SGLang is a high-performance serving framework for large language models and multimodal models.

Python 35,717 8,715 Updated Sep 10, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 1 Updated Aug 19, 2026

A high-throughput and memory-efficient inference and serving engine for LLMs

Python 91,368 21,988 Updated Sep 10, 2026

Fork of the classic

Rust 1 Updated Jul 30, 2026

Development repository for the Triton language and compiler

MLIR 20,117 3,177 Updated Sep 10, 2026

Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.

Rust 15,890 1,041 Updated Sep 9, 2026

LLM inference in C/C++

C++ 127,657 22,956 Updated Sep 9, 2026