Skip to content
#

ray-serve

Here are 34 public repositories matching this topic...

Self-hosted, OpenAI-compatible inference for the agentic era: reasoning LLMs, universal tool calling, and the Responses API alongside embeddings, speech, and image models — many models sharing your GPUs, one gateway. Powered by Ray Serve.

  • Updated Sep 7, 2026
  • Python

Multi-backend LLM serving and training platform — vLLM/Triton/Ray Serve/KServe/BentoML behind one contract, Kueue/Karpenter GPU orchestration, Ray Train/FSDP/DeepSpeed with LoRA/PEFT and DVC, MLflow/W&B tracking, and a tool-grounded LangGraph advisor — CI-validated without real GPU cost.

  • Updated Jul 27, 2026
  • Python

A comprehensive guide to setting up and managing Raspberry Pi, Ray Clusters, and distributed AI workloads. Includes network troubleshooting, IP configuration, Ray Dashboard, and Python script execution for scalable AI applications.

  • Updated Feb 23, 2025
  • Shell

Add this topic to your repo

To associate your repository with the ray-serve topic, visit your repo's landing page and select "manage topics."

Learn more