Legal Agent Benchmark (LAB): An open-source benchmark for evaluating agents on real legal work.
Harvey LAB is an open-source project aimed at benchmarking LLM agents' abilities to perform legal work in realistic environments.
LAB consists of two parts: a dataset of tasks containing agent instructions, documents, and rubrics as well as an execution harness for running and evaluating agents against those tasks.
LAB is an ongoing project and we expect to consistently add to and refine the task set and execution harness.
Read the announcement post: Introducing Harvey's Legal Agent Benchmark
Start with the full walkthrough in docs/tutorial.md — it takes one realistic M&A data-room assignment end to end: setup, task inspection, agent run, scoring, report review, and comparison dashboards.
The harness and grader ship as the lab-core Python package (import name
lab_core), built as a wheel on every release:
uv add "lab-core @ https://github.com/harveyai/harvey-labs/releases/download/v1.2.0/lab_core-1.2.0-py3-none-any.whl"Point it at a checkout of the tasks with LAB_ROOT=/path/to/harvey-labs; the
python -m lab_core.… commands then work outside the repository. See CONTRIBUTING.md.
Grading .docx output also needs the pandoc CLI, 3.5 or later, on PATH. On Linux, this
runs the installer bundled in the wheel, the same one the sandbox image and CI use:
sudo sh "$(uv run python -c 'import lab_core, pathlib; print(pathlib.Path(lab_core.__file__).parent / "sandbox" / "install_pandoc.sh")')"| Guide | Description |
|---|---|
| Architecture | Task model, harness, tools, adapters, reports, and sweeps |
| Evaluation Methodology | All-pass rubric scoring and LLM judge behavior |
| Contributing | Add tasks, model adapters, evaluation improvements, and docs |
If you use Harvey LAB in your research, please cite it as:
@misc{harveylab2026,
title = {Harvey LAB: The Legal Agent Benchmark},
author = {{Harvey AI}},
year = {2026},
version = {v1.0},
url = {https://github.com/harveyai/harvey-labs/tree/v1.0},
note = {Announcement: \url{https://www.harvey.ai/blog/introducing-harveys-legal-agent-benchmark}}
}