Highlights
- Pro
Pinned Loading
-
MoE-Kernel-Agent
MoE-Kernel-Agent Public🥈 2nd Place (Full-Agent) · MLSys 2026 FlashInfer AI Kernel Generation Contest — MoE Track (NVIDIA) · 25.7× median speedup over PyTorch reference
-
Triton-Attention-Kernels
Triton-Attention-Kernels PublicGPU attention kernels for LLM inference, implemented from scratch in Triton.
Jupyter Notebook
-
DeepSeek-V3-FP8-MoE-Kernel
DeepSeek-V3-FP8-MoE-Kernel PublicHigh-performance fused MoE kernel for DeepSeek-V3 FP8 inference on Blackwell (B200), written in Triton. 1.08× faster than FlashInfer on total workload latency.
Python
-
SVD-Flash
SVD-Flash PublicSVD-Flash: Efficient LLM inference via SVD Compression and Tiling on AWS Trainium
Python
-
ThunderAgent
ThunderAgent PublicForked from ThunderAgent-org/ThunderAgent
A simple, fast and robust program-aware agentic inference system.
HTML
-
If the problem persists, check the GitHub status page or contact support.