Skip to content
#

model-serving

Here are 15 public repositories matching this topic...

highperf-ai-ml-inference

High-perf C++ AI/ML inference engine with ONNX Runtime & LibTorch. CPU default, GPU opt-in. CLI + REST API, cross-platform via CMake/Docker. Auto-fetch models/assets, CUDA runtime check, CI merge gating, future multi-model & GPU training support.

  • Updated Aug 16, 2025
  • C++

Add this topic to your repo

To associate your repository with the model-serving topic, visit your repo's landing page and select "manage topics."

Learn more