Skip to content
View Jerry2423's full-sized avatar

Highlights

  • Pro

Block or report Jerry2423

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse

Pinned Loading

  1. MoE-Kernel-Agent MoE-Kernel-Agent Public

    🥈 2nd Place (Full-Agent) · MLSys 2026 FlashInfer AI Kernel Generation Contest — MoE Track (NVIDIA) · 25.7× median speedup over PyTorch reference

    Python 8 1

  2. Triton-Attention-Kernels Triton-Attention-Kernels Public

    GPU attention kernels for LLM inference, implemented from scratch in Triton.

    Jupyter Notebook

  3. DeepSeek-V3-FP8-MoE-Kernel DeepSeek-V3-FP8-MoE-Kernel Public

    High-performance fused MoE kernel for DeepSeek-V3 FP8 inference on Blackwell (B200), written in Triton. 1.08× faster than FlashInfer on total workload latency.

    Python

  4. SVD-Flash SVD-Flash Public

    SVD-Flash: Efficient LLM inference via SVD Compression and Tiling on AWS Trainium

    Python

  5. ThunderAgent ThunderAgent Public

    Forked from ThunderAgent-org/ThunderAgent

    A simple, fast and robust program-aware agentic inference system.

    HTML

  6. nki-llama nki-llama Public

    Python