Skip to content

Commit 2f23497

Browse files
correction: the grep method measures naming, not coverage
Coverage instrumentation on five evals I did not contribute to (agentharm, xstest, ifeval, worldsense, mask; 18 metric definitions) shows every one executes under test. They are reached through end-to-end task tests without being named in a test file, which is the failure mode the grep method cannot see and which dominates the others. Flags the headline as under revision rather than editing it quietly. The taxonomy and the stereoset finding are unaffected: both are reproduced against installed packages.
1 parent cb2aafc commit 2f23497

2 files changed

Lines changed: 64 additions & 0 deletions

File tree

‎README.md‎

Lines changed: 32 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,37 @@
11
# The metrics are the least-tested code in the eval stack
22

3+
> ## Correction in progress, 2026-08-12
4+
>
5+
> **The headline claim on this page is not supported by the method used to produce it, and I am
6+
> in the middle of replacing it.** Stated plainly rather than quietly edited, because the
7+
> argument of this repository is that wrong numbers survive when nobody says anything.
8+
>
9+
> The figures below measure **"no test anywhere refers to this metric by either of its
10+
> names."** I presented that as a proxy for "untested". It is not a good one. Re-running
11+
> `inspect_evals` under **coverage instrumentation**, which records what actually executes,
12+
> every metric in a sample of five evals I had not contributed to (`agentharm`, `xstest`,
13+
> `ifeval`, `worldsense`, `mask`, 18 definitions) **did execute under test**. They are reached
14+
> through end-to-end task tests without ever being named in a test file.
15+
>
16+
> So the grep-based figure is measuring naming convention, not test coverage, and it
17+
> substantially overstates the problem. A full coverage-instrumented measurement is running and
18+
> this page will be rewritten around it.
19+
>
20+
> Two things survive the correction, and they are the parts that mattered:
21+
>
22+
> - **The taxonomy below**, which is derived from defects reproduced against installed
23+
> packages, not from this metric.
24+
> - **The finding that started it**: `stereotype_score` in `inspect_evals` had no test that
25+
> exercised it, a run where every answer failed to parse reported StereoSet's *ideal* score,
26+
> and [the fix](https://github.com/UKGovernmentBEIS/inspect_evals/pull/2123) ships a test
27+
> that fails on `main`. That is reproducible regardless of what the coverage rate turns out
28+
> to be.
29+
>
30+
> Executed under test still is not the same as asserted on. The replacement measurement will
31+
> use mutation testing, which is the only method here that checks assertions rather than
32+
> execution.
33+
34+
335
Measured with each framework's own registration marker, counting a metric as touched if either
436
its function name or its registered name appears anywhere in any test file:
537

‎research/metric-test-coverage.md‎

Lines changed: 32 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,37 @@
11
# The metrics are the least-tested code in the eval stack
22

3+
> ## Correction in progress, 2026-08-12
4+
>
5+
> **The headline claim on this page is not supported by the method used to produce it, and I am
6+
> in the middle of replacing it.** Stated plainly rather than quietly edited, because the
7+
> argument of this repository is that wrong numbers survive when nobody says anything.
8+
>
9+
> The figures below measure **"no test anywhere refers to this metric by either of its
10+
> names."** I presented that as a proxy for "untested". It is not a good one. Re-running
11+
> `inspect_evals` under **coverage instrumentation**, which records what actually executes,
12+
> every metric in a sample of five evals I had not contributed to (`agentharm`, `xstest`,
13+
> `ifeval`, `worldsense`, `mask`, 18 definitions) **did execute under test**. They are reached
14+
> through end-to-end task tests without ever being named in a test file.
15+
>
16+
> So the grep-based figure is measuring naming convention, not test coverage, and it
17+
> substantially overstates the problem. A full coverage-instrumented measurement is running and
18+
> this page will be rewritten around it.
19+
>
20+
> Two things survive the correction, and they are the parts that mattered:
21+
>
22+
> - **The taxonomy below**, which is derived from defects reproduced against installed
23+
> packages, not from this metric.
24+
> - **The finding that started it**: `stereotype_score` in `inspect_evals` had no test that
25+
> exercised it, a run where every answer failed to parse reported StereoSet's *ideal* score,
26+
> and [the fix](https://github.com/UKGovernmentBEIS/inspect_evals/pull/2123) ships a test
27+
> that fails on `main`. That is reproducible regardless of what the coverage rate turns out
28+
> to be.
29+
>
30+
> Executed under test still is not the same as asserted on. The replacement measurement will
31+
> use mutation testing, which is the only method here that checks assertions rather than
32+
> execution.
33+
34+
335
Measured with each framework's **own** registration marker, so "this is a metric" is the
436
framework's judgement and not mine. A metric counts as touched if **either** its function name
537
**or** the name it is registered under appears anywhere in any test file.

0 commit comments

Comments
 (0)