An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
-
Updated
Sep 8, 2026 - Python
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
This is suite of the hands-on training materials that shows how to scale CV, NLP, time-series forecasting workloads with Ray.
Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GPUs, one gateway. Powered by Ray Serve.
Batch LLM Inference with Ray Data LLM: From Simple to Advanced
Building Real-Time Inference Pipelines with Ray Serve
BioEngine is a distributed AI platform that brings the power of cloud computing to bioimage analysis.
A Production-Ready, Scalable RAG-powered LLM-based Context-Aware QA App
Multi-backend LLM serving and training platform — vLLM/Triton/Ray Serve/KServe/BentoML behind one contract, Kueue/Karpenter GPU orchestration, Ray Train/FSDP/DeepSpeed with LoRA/PEFT and DVC, MLflow/W&B tracking, and a tool-grounded LangGraph advisor — CI-validated without real GPU cost.
Create Context-Aware Q&A Interfaces from Your Own Data with LLMs and Vector Embeddings - Includes an automated embedding pipeline and a model-powered Q&A interface
Plugin-first framework for modular Python services with FastAPI ingress and optional Ray execution.
Distributed LLM inference with Ray Serve — Ray Summit 2026 workshop notebooks and deployable examples.
Distributed RAG platform on Kubernetes using Ray Serve, FastAPI, vector databases, and LLM orchestration.
Multi-tenant context caching and serving for tabular in-context-learning foundation models (TabICL, TabPFN, and similar)
A distributed ML recommendation system — real-time streaming, multi-node distributed training, and fault-tolerant, scalable serving.
A comprehensive guide to setting up and managing Raspberry Pi, Ray Clusters, and distributed AI workloads. Includes network troubleshooting, IP configuration, Ray Dashboard, and Python script execution for scalable AI applications.
Production-grade scalable embedding API server using SentenceTransformers "intfloat/multilingual-e5-base" model, powered by Ray Serve for multi-GPU orchestration, with Prometheus & Grafana monitoring.
A drop-in replacement of fastapi to enable scalable and fault tolerant deployments with ray serve
Production-grade Grafana dashboards, Prometheus recording & alerting rules for vLLM, Ray Serve, Kubernetes and NVIDIA GPU LLM inference observability — TTFT, TPOT, KV cache, scheduler and GPU analytics.
Production LLM serving infrastructure using Triton Inference Server, vLLM, and Ray Serve with OpenAI-compatible endpoints. Includes Kubernetes autoscaling configs driven by DCGM GPU metrics and a BentoML packaging path for portable model deployment.
To associate your repository with the ray-serve topic, visit your repo's landing page and select "manage topics."