MBPP is a dataset for evaluating the ability of models to synthesize short Python programs from natural language descriptions. It contains programming tasks that are designed to be solvable by entry-level programmers. We evaluate on the sanitized test split of the dataset, which was hand-verified to remove samples that lacked detail or were ambiguous.
Contributed by @jddantes
Install with pip install inspect-evals, or uv sync from a checkout of this repository.
uv run inspect eval inspect_evals/mbpp --model openai/gpt-5-nanoYou can also import tasks as normal Python objects and run them from python:
from inspect_ai import eval
from inspect_evals.mbpp import mbpp
eval(mbpp)Drop uv run if you manage dependencies yourself. Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.
You can control a variety of options from the command line. For example:
uv run inspect eval inspect_evals/mbpp --limit 10 --sample-shuffle
uv run inspect eval inspect_evals/mbpp --max-connections 10
uv run inspect eval inspect_evals/mbpp --temperature 0.5See uv run inspect eval --help for all available options.
temperature(float): The temperature when generating completions (default:0.5)
Here is a sample from the dataset:
Write a python function to find the volume of a triangular prism.
test_list:
```
[ "assert find_Volume(10,8,6) == 240",
"assert find_Volume(3,2,2) == 6",
"assert find_Volume(1,2,1) == 1" ]
```
The model is tasked to write Python code that will pass the assert statements in the test_list.
The model is prompted to solve the problem, given its description and test cases to pass. Additionally, few shot examples can be included in the prompt, with prompts patterned after an AgentCoder implementation which topped the leaderboard.
Evaluation metrics follow similar benchmarks such as HumanEval. The benchmark uses the
where we sample
- Completion extraction now also handles a
python3label, a fence that is never closed (a truncated completion), a fence after prose on the same line, a closing fence glued to the last line of code, an empty block before the real one, and a stray fence before the labelled block. Each of these used to fall back to the raw completion, Markdown included, or to an empty block. Blocks labelledpyandpythonare now taken in order of position, where apythonblock used to win wherever it was. A replay of 5,456 unique MBPP completions from 12 models extracted the same code as 3-A for all of them.
- Fix completion extraction so Python-labelled and unlabelled Markdown code fences are handled consistently, including case-insensitive labels and CRLF line endings. (@sylvesterkaczmarek)
- Migrate version to new scheme. See #907.
- Adds backoff policy for functions that connect to huggingface servers.