Skip to content

Popular repositories Loading

  1. openfang openfang Public

    Open-source Agent Operating System

    Rust 18.2k 2.3k

  2. picolm picolm Public

    Run a 1-billion parameter LLM on a $10 board with 256MB RAM

    C 1.9k 241

  3. autokernel autokernel Public

    Autoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.

    Python 1.6k 163

  4. AutoMegaKernel AutoMegaKernel Public

    An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682

    Python 137 13

  5. qwen3.5-triton qwen3.5-triton Public

    Pure Triton kernels for Qwen3.5-27B inference on NVIDIA B200

    Python 123 10

  6. auto auto Public

    the agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for microdollars. nothing figured out twice, paper: https://arxiv.…

    Rust 123 11

Repositories

Showing 10 of 27 repositories
  • runinfra-sdk Public

    Official RunInfra SDK (TypeScript + Python) | optimized inference deployments

    RightNow-AI/runinfra-sdk's past year of commit activity
    TypeScript 4 5 3 17 Updated Aug 14, 2026
  • autoevolve Public

    Agent-native evolutionary optimization. Say the goal in english, evolve code toward a measured target with a population of coding agents.

    RightNow-AI/autoevolve's past year of commit activity
    Python 4 Apache-2.0 0 0 0 Updated Aug 6, 2026
  • inference-cost-truth Public

    What LLM inference actually costs. 874 verified price rows across 24 providers, 378 GPU rental rates, 395 cited throughput datapoints, and a self-hosting break-even model. Every number carries a source URL, a retrieval date, and passed a mechanical evidence gate. Dated immutable snapshots. CC BY 4.0.

    RightNow-AI/inference-cost-truth's past year of commit activity
    Python 6 1 0 0 Updated Jul 31, 2026
  • local-kimi Public

    Optimized local serving engine for Kimi-Linear-48B: INT4 quantizer, fused decode kernels for a measured 3.18x, and an OpenAI-compatible server. Ships with k3, a bridge that detects the client per request so Claude Code, Codex, Cline, Aider and opencode all work unchanged.

    RightNow-AI/local-kimi's past year of commit activity
    Python 39 Apache-2.0 6 0 0 Updated Jul 30, 2026
  • runinfra-cli Public
    RightNow-AI/runinfra-cli's past year of commit activity
    Shell 1 0 0 0 Updated Jul 29, 2026
  • inkling-turbo Public

    Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.

    RightNow-AI/inkling-turbo's past year of commit activity
    Python 5 Apache-2.0 1 0 0 Updated Jul 26, 2026
  • sglang Public Forked from sgl-project/sglang

    SGLang is a high-performance serving framework for large language models and multimodal models.

    RightNow-AI/sglang's past year of commit activity
    Python 1 Apache-2.0 8,741 0 0 Updated Jul 23, 2026
  • Memoir Public

    Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792

    RightNow-AI/Memoir's past year of commit activity
    Python 15 Apache-2.0 2 0 0 Updated Jul 23, 2026
  • autotree Public

    Tree execution engine for LLM inference: fork, merge, prune KV cache at token granularity

    RightNow-AI/autotree's past year of commit activity
    Python 5 Apache-2.0 0 0 0 Updated Jul 20, 2026
  • bonsai-turbo Public

    Single-launch batch-1 decode engine for PrismML Bonsai 27B (ternary and 1-bit) on NVIDIA GPUs. 1.76x the vendor llama.cpp fork on H100, same outputs.

    RightNow-AI/bonsai-turbo's past year of commit activity
    Cuda 16 Apache-2.0 6 2 0 Updated Jul 16, 2026