Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 

README.md

MBPP: Mostly Basic Python Problems

MBPP is a dataset for evaluating the ability of models to synthesize short Python programs from natural language descriptions. It contains programming tasks that are designed to be solvable by entry-level programmers. We evaluate on the sanitized test split of the dataset, which was hand-verified to remove samples that lacked detail or were ambiguous.

Contributed by @jddantes

Usage

Installation

Install with pip install inspect-evals, or uv sync from a checkout of this repository.

Running evaluations

uv run inspect eval inspect_evals/mbpp --model openai/gpt-5-nano

You can also import tasks as normal Python objects and run them from python:

from inspect_ai import eval
from inspect_evals.mbpp import mbpp
eval(mbpp)

Drop uv run if you manage dependencies yourself. Log viewing (inspect view) and default-model setup are documented in the Inspect Evals README.

Options

You can control a variety of options from the command line. For example:

uv run inspect eval inspect_evals/mbpp --limit 10 --sample-shuffle
uv run inspect eval inspect_evals/mbpp --max-connections 10
uv run inspect eval inspect_evals/mbpp --temperature 0.5

See uv run inspect eval --help for all available options.

Parameters

mbpp

  • temperature (float): The temperature when generating completions (default: 0.5)

Dataset

Here is a sample from the dataset:

Write a python function to find the volume of a triangular prism.

test_list:

```
[ "assert find_Volume(10,8,6) == 240",
   "assert find_Volume(3,2,2) == 6",
   "assert find_Volume(1,2,1) == 1" ]
```

The model is tasked to write Python code that will pass the assert statements in the test_list.

Scoring

The model is prompted to solve the problem, given its description and test cases to pass. Additionally, few shot examples can be included in the prompt, with prompts patterned after an AgentCoder implementation which topped the leaderboard.

Evaluation metrics follow similar benchmarks such as HumanEval. The benchmark uses the $\text{pass}@k$ metric to measure functional correctness. In brief terms, this is the per-problem probability of at least 1 correct sample generation given $k$ generations. It is defined using the following expectation:

$$ \text{pass@}k := \underset{\text{Problems}}{\mathbb{E}}\left[1-\frac{{n-c}\choose{k}}{{n}\choose{k}}\right] $$

where we sample $n \geq k$ generations to reduce variance. Note that the default in this benchmark implementation is $n = 5$, and we evaluate $\text{pass}@k$ for $k \in \{1, 2, 5\}$.

Changelog

[4-A] - 2026-10-01

  • Completion extraction now also handles a python3 label, a fence that is never closed (a truncated completion), a fence after prose on the same line, a closing fence glued to the last line of code, an empty block before the real one, and a stray fence before the labelled block. Each of these used to fall back to the raw completion, Markdown included, or to an empty block. Blocks labelled py and python are now taken in order of position, where a python block used to win wherever it was. A replay of 5,456 unique MBPP completions from 12 models extracted the same code as 3-A for all of them.

[3-A] - 2026-08-19

  • Fix completion extraction so Python-labelled and unlabelled Markdown code fences are handled consistently, including case-insensitive labels and CRLF line endings. (@sylvesterkaczmarek)

[2-A] - 2026-02-16

  • Migrate version to new scheme. See #907.

[1.0.1] - 2025-12-18

  • Adds backoff policy for functions that connect to huggingface servers.