Skip to content
@LLM-EVAL-LAB

LLM-EVAL-LAB

Popular repositories Loading

  1. eval_awareness_1 eval_awareness_1 Public

    Forked from honeypots-mats7/eval_awareness_1

    Python

  2. inspect_ai inspect_ai Public

    Forked from UKGovernmentBEIS/inspect_ai

    Inspect: A framework for large language model evaluations

    Python

  3. evals evals Public

    Forked from openai/evals

    Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

    Python

  4. deepeval deepeval Public

    Forked from confident-ai/deepeval

    The LLM Evaluation Framework

    Python

  5. opik opik Public

    Forked from comet-ml/opik

    Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.

    Python

  6. promptfoo promptfoo Public

    Forked from promptfoo/promptfoo

    Test your prompts, agents, and RAGs. Red teaming/pentesting/vulnerability scanning for AI. Compare performance of GPT, Claude, Gemini, DeepSeek, and more. Simple declarative configs with command li…

    TypeScript

Repositories

Showing 10 of 13 repositories

People

This organization has no public members. You must be a member to see who’s a part of this organization.

Top languages

Loading…

Most used topics

Loading…