A high-throughput and memory-efficient inference and serving engine for LLMs
-
Updated
Sep 8, 2026 - Python
A high-throughput and memory-efficient inference and serving engine for LLMs
Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
Port of OpenAI's Whisper model in C/C++
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
Making large AI models cheaper, faster and more accessible
Cross-platform, customizable ML solutions for live and streaming media.
Faster Whisper transcription with CTranslate2
ncnn is a high-performance neural network inference framework optimized for the mobile platform
Machine Learning Engineering Open Book
🎨 The exhaustive Pattern Matching library for TypeScript, with smart type inference.
NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker.
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Large Language Model Text Generation Inference
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
PyTorch implementation of YOLOv3, YOLOv3-SPP, and YOLOv3-tiny for real-time object detection with training, validation, inference, and multi-format export.
To associate your repository with the inference topic, visit your repo's landing page and select "manage topics."