-
Meta
- Stanford
-
18:18
(UTC -07:00) - in/gavinwxx
- @GavinXWang4899
Stars
A unified library for building, evaluating, and storing speculative decoding algorithms for LLM inference in vLLM
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
Kubernetes enhancements for Network Topology Aware Gang Scheduling & Autoscaling
CUDA Templates and Python DSLs for High-Performance Linear Algebra
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
A programmable Mixture-of-Models router for heterogeneous LLM inference
gavinkvx / dynamo
Forked from ai-dynamo/dynamoA Datacenter Scale Distributed Inference Serving Framework
Building the Virtuous Cycle for AI-driven LLM Systems
A Datacenter Scale Distributed Inference Serving Framework
gavinkvx / flashinfer
Forked from flashinfer-ai/flashinferFlashInfer: Kernel Library for LLM Serving
slime is an LLM post-training framework for RL Scaling.
FlashInfer: Kernel Library for LLM Serving
An LLM post-training framework with vLLM for RL Scaling
SGLang is a high-performance serving framework for large language models and multimodal models.
gavinkvx / vllm
Forked from vllm-project/vllmA high-throughput and memory-efficient inference and serving engine for LLMs
A high-throughput and memory-efficient inference and serving engine for LLMs
gavinkvx / llama2-fork
Forked from karpathy/llama2.cFork of the classic
Development repository for the Triton language and compiler
Burn is a next generation tensor library and Deep Learning Framework that doesn't compromise on flexibility, efficiency and portability.