Skip to content
#

ai-evals

Here are 90 public repositories matching this topic...

Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.

  • Updated Sep 8, 2026
  • Python
agent-workspace-architecture

Agent-ready knowledge architecture, run daily in Claude Code: turn a coding agent into a system you can hand work to and trust while you're away. 18 load-bearing patterns, one-page workspace map, interactive tour, guided learn track, teardowns of real systems, roles, typed memory, hooks, delegation queue, self-audits. Fork-ready samples.

  • Updated Sep 8, 2026
  • Python

Eval-first AI agent that triages property maintenance emails. The real work is the eval system around it: trace-driven error analysis, code graders and validated LLM-as-judge (TPR/TNR), component and end-to-end evals, a failure taxonomy, and a CI regression gate. LangGraph, FastAPI, Langfuse.

  • Updated Jun 7, 2026
  • Python

FDE Consultants Protocoles — open-source Forward Deployed Engineer + DeepSCR protocol: turn any coding agent (Claude Code, Codex, Cursor) into a certifying engineer with verifiable AI Assurance Scores, MCP tools, and a public trust registry. Apache-2.0.

  • Updated Jul 3, 2026
  • HTML

Frontend for Calibrate, a framework for evaluating AI agents: speech-to-text, text-to-speech, LLM evaluation, end-to-end simulations

  • Updated Sep 3, 2026
  • TypeScript

Add this topic to your repo

To associate your repository with the ai-evals topic, visit your repo's landing page and select "manage topics."

Learn more