LLM-EVAL-LAB
Popular repositories Loading
-
-
inspect_ai
inspect_ai PublicForked from UKGovernmentBEIS/inspect_ai
Inspect: A framework for large language model evaluations
Python
-
evals
evals PublicForked from openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
Python
-
-
opik
opik PublicForked from comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
Python
-
promptfoo
promptfoo PublicForked from promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command li…
TypeScript
Repositories
- ShapingBench Public Forked from deep-diver/ShapingBench
Executable behavioral contracts and hidden-evaluation harness for coding agents across nine software domains.
- deepteam Public Forked from confident-ai/deepteam
DeepTeam is a framework to red team LLMs and AI agents.
- opik Public Forked from comet-ml/opik
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
- iFixAi Public Forked from ifixai-ai/iFixAi
Independent Auditing of AI Agents. Run by human or the agent itself, to answer the most crucial question in the AI Agent Economy. Is the agent doing what is supposed to do? With iFixAi you can have this answer in less than 120 seconds.
- promptfoo Public Forked from promptfoo/promptfoo
Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command line and CI/CD integration. Used by OpenAI and Anthropic.
- inspect_ai Public Forked from UKGovernmentBEIS/inspect_ai
Inspect: A framework for large language model evaluations
- mslearn-ai-language Public Forked from MicrosoftLearning/mslearn-ai-language
Lab files for Azure AI Language modules
- Infosys-Responsible-AI-Toolkit Public Forked from Infosys/Infosys-Responsible-AI-Toolkit
The Infosys Responsible AI toolkit incorporates various features including safety, security, explainability, fairness, bias and hallucination detection to ensure AI solutions are trustworthy and transparent.
- evals Public Forked from openai/evals
Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
People
This organization has no public members. You must be a member to see who’s a part of this organization.
Top languages
Loading…
Most used topics
Loading…