Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
-
Updated
Sep 8, 2026 - Python
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
Operational doctrine for practical AI systems design.
Benchmarking Open-Ended Inference Optimization by AI Agents
pytest for AI Apps — Fast, free, offline testing and self-improving prompts for Python developers.
Evaluation Infrastructure for AI Agents
Core engine behind Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Learn to evaluate AI products for production — 21 hands-on lessons on evals, metrics, fairness, agents, red teaming, and release decisions for working PMs.
Collection of frameworks and tools for AI evalations, including tool-use, agentic AI, MCP, and multimodal
Agent-ready knowledge architecture, run daily in Claude Code: turn a coding agent into a system you can hand work to and trust while you're away. 18 load-bearing patterns, one-page workspace map, interactive tour, guided learn track, teardowns of real systems, roles, typed memory, hooks, delegation queue, self-audits. Fork-ready samples.
An sdk and framework for evaluating and comparing multiple model outputs using configurable LLM-based jurors
Eval-first AI agent that triages property maintenance emails. The real work is the eval system around it: trace-driven error analysis, code graders and validated LLM-as-judge (TPR/TNR), component and end-to-end evals, a failure taxonomy, and a CI regression gate. LangGraph, FastAPI, Langfuse.
Governed multi-LLM Responsible AI control plane for prior-authorization decision support, with PHI-safe audit trails, HITL review, evidence packets, and deterministic governance evals.
FDE Consultants Protocoles — open-source Forward Deployed Engineer + DeepSCR protocol: turn any coding agent (Claude Code, Codex, Cursor) into a certifying engineer with verifiable AI Assurance Scores, MCP tools, and a public trust registry. Apache-2.0.
Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Benchmark LLM jailbreak resilience across providers with standardized tests, adversarial mode, rich analytics, and a clean Web UI.
Backend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations
Public research on LLM internals, Jacobian lenses, SAE steering, nonlinear dynamics, evaluation, and inspectable AI systems.
Open-source toolkit for assessing whether an AI workflow is ready for production: governance, RAG quality, evals, observability, human review, cost, risk, and business value.
To associate your repository with the ai-evals topic, visit your repo's landing page and select "manage topics."