Community maintained hardware plugin for vLLM on Ascend
-
Updated
Sep 8, 2026 - C++
Community maintained hardware plugin for vLLM on Ascend
A scalable inference server for models optimized with OpenVINO™
A highly optimized LLM inference acceleration engine for Llama and its variants.
Large-scale Auto-Distributed Training/Inference Unified Framework | Memory-Compute-Control Decoupled Architecture | Multi-language SDK & Heterogeneous Hardware Support
🔥 Alternative to Ollama — multi-model serving with sub-ms model switching · CPU-only 20B inference for Edge AI · llama.cpp + stablediffusion.cpp
Serve pytorch / torch models using Drogon
High-Throughput Batch Inference
PipelineScheduler optimizes workload distribution between servers and edge devices, setting optimal batch sizes to maximize throughput and minimize latency amid content dynamics and network instability. It also addresses resource contention with spatiotemporal inference scheduling to reduce co-location interference.
A streaming action runtime for AI agents, model serving, and multimodal APIs
FastText inference optimization with trie-backed n-gram ids, mark-compact vector storage, and mmap-friendly retrieval.
Distributed ML training and serving platform
Serving object detection models on different hardware.
High-perf C++ AI/ML inference engine with ONNX Runtime & LibTorch. CPU default, GPU opt-in. CLI + REST API, cross-platform via CMake/Docker. Auto-fetch models/assets, CUDA runtime check, CI merge gating, future multi-model & GPU training support.
To associate your repository with the model-serving topic, visit your repo's landing page and select "manage topics."