Skip to content
#

inference

Here are 322 public repositories matching this topic...

Production-grade C++20/CUDA distributed LLM inference system with TCP networking, MPI scheduling, and content-addressed storage. Features comprehensive benchmarking (p50/p95/p99 latencies), epoll async I/O, and optimizations for low-VRAM GPUs (1GB+). 2500+ LOC demonstrating systems engineering expertise.

  • Updated Dec 27, 2025
  • C++
highperf-ai-ml-inference

High-perf C++ AI/ML inference engine with ONNX Runtime & LibTorch. CPU default, GPU opt-in. CLI + REST API, cross-platform via CMake/Docker. Auto-fetch models/assets, CUDA runtime check, CI merge gating, future multi-model & GPU training support.

  • Updated Aug 16, 2025
  • C++

Add this topic to your repo

To associate your repository with the inference topic, visit your repo's landing page and select "manage topics."

Learn more