Skip to content

Pull requests: danielrosehill/Awesome-AI-Evaluations-Tools

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Add PaperEdits agentic video benchmark
#28 opened Sep 2, 2026 by sarvob Loading…
Add digline
#27 opened Sep 2, 2026 by alexpran Loading…
Add RepoTrials to Coding Evals
#24 opened Aug 15, 2026 by PozziTiv4ik Loading…
Add RewardHarness to multimodal evals
#23 opened Aug 15, 2026 by reacher-z Loading…
Add Coder Eval to Coding Evals
#20 opened Jul 28, 2026 by uipreliga Loading…
Add StructEval benchmark
#19 opened Jul 27, 2026 by reacher-z Loading…
Add ClawBench to Browser & Web Agent Evals
#17 opened Jul 26, 2026 by reacher-z Loading…
Add Awesome AI Testing to Awesome Lists / Resources
#15 opened Jul 13, 2026 by tugkanboz Loading…
Add AI Governance Benchmarks to README
#14 opened Jul 8, 2026 by tombudd Loading…
Add agenttrace
#3 opened May 10, 2026 by luoyuctl Loading…
ProTip! Updated in the last three days: updated:>2026-09-05.