Skip to content

First-class Score.reason and normalized scorer failure policy - #4629

Merged
dragonstyle merged 31 commits into
UKGovernmentBEIS:mainfrom
MattFisher:feature/score-reason
Aug 26, 2026
Merged

dragonstyle merged 31 commits into
UKGovernmentBEIS:mainfrom
MattFisher:feature/score-reason

Conversation

@MattFisher

@MattFisher MattFisher commented Jul 26, 2026 •

Copy link
Copy Markdown
Contributor

This PR contains:

  • New features
  • Changes to dev-tools e.g. CI config / github tooling
  • Docs
  • Bug fixes
  • Code refactor

What is the current behavior? (You can also link to an open issue here)

Closes #4567.

Built-in scorers record "couldn't get a usable reading" three different ways: pattern()/answer() return NOANSWER, math() returns INCORRECT with an answer="None" string artefact, and model_graded_* returns Score.unscored() with an ad hoc metadata["unscored_reason"] key. There's no first-class field for why a score is abnormal, separate from the value (which decides what the metric sees) and the human-readable explanation. The interim metadata["reason"] convention is fragile: apply_score_edit-style edits replace metadata wholesale and silently drop it, and it squats on the author's own metadata namespace.

What is the new behavior?

  • Score.reason: ScoreReason | str | None — a new field recording a machine-readable failure mode, separate from value/explanation/metadata. ScoreReason is an open-union Literal with a 5-value attribution taxonomy — model under test: invalid_response_format, refusal, no_response (kept fine-grained; refusal rate is a first-class quantity); measurement instrument: grader_failed (grader model failed), scoring_failed (scorer could not run) — coarse deliberately, with the detail in explanation. The standard vocabulary is discoverable via IDE completion, but any custom string is legal, and adding values later is trivially non-breaking while removing them never is.
  • Score.unscored(reason=...) — the primary construction site for instrument (grader/harness) failures.
  • ScoreEdit.reason — follows the same UNCHANGED-sentinel pattern as answer/explanation/metadata, so reason survives score edits by construction (edits no longer silently drop it the way wholesale-metadata-replacement does).
  • A normalized attribution policy across built-in scorers — whose output failed?
    • Model-under-test format violations stay INCORRECT (in the metric denominator) with reason="invalid_response_format": pattern()/answer() revert from NOANSWER (Pattern scorer noanswer #3628) back to INCORRECT+reason. (math()'s equivalent fix is deliberately not in this PR — see coordination note below.)
    • Grader/instrument failures return Score.unscored(reason=...), excluded from the denominator: model_graded_qa/model_graded_fact move their grade-parse-failure reason from metadata["unscored_reason"] (fix(scorer): mark model_graded grade-parse failure as unscored (not INCORRECT) #4048) into the reason field as grader_failed; perplexity()/target_perplexity() return Score.unscored(reason="scoring_failed") instead of a raw NaN Score (no_response when the completion is empty); multi_scorer() sets reason="scoring_failed" when every sub-scorer declines.
  • Backward compatibility: a Pydantic before-validator lifts legacy metadata["unscored_reason"] (written by 0.3.245+) into reason at read time, without touching the metadata key — old logs read cleanly with no migration needed. A follow-up fix ensures an edit that explicitly clears reason also drops the stale metadata["unscored_reason"] key, so a log round-trip (write → read) can't resurrect a reason the user just cleared.
  • Epoch reducers (mean_score, median_score, mode_score, at_least, pass_at, pass_k) and multi_scorer's reducer path now retain reason across epochs (same "equal across all scores, else None" rule already used for answer/explanation).
  • score_<name>_reason columns now appear in samples_df()/score_details() output.
  • inspect-openapi.json regenerated to include reason on Score, ScoreEdit, and EvalSampleScore.

Does this PR introduce a breaking change? (What changes might users need to make in their application due to this PR?)

Yes, one deliberate behavior change flagged in the design discussion on #4567: pattern() and answer() now score unmatched model output as INCORRECT instead of NOANSWER (reverting #3628), with the failure mode now recorded in reason="invalid_response_format". Note this changes no default metric numbers: the default value_to_float() already maps NOANSWER to 0, so format violations were already in the denominator at value 0. The change matters only for analyses that filter on value == "N" and for custom value_to_float() mappings that treat noanswer differently from incorrect — those should key on reason instead.

Everything else is additive: new optional fields default to None/UNCHANGED, and the read shim keeps old logs compatible with no migration step.

Other information:

Coordination note: math()'s answer-extraction-failure fix is deliberately not included here, per the agreed resolution on #4091 — @antnewman will re-target that PR's math() fix onto Score.reason once this PR merges, rather than this PR carrying a competing edit to _math.py/test_math.py.

Known follow-ups (out of scope for this PR):

  1. The ts-mono viewer submodule's generated TypeScript types need a separate regeneration PR in that repo — done: Regenerate types for Score.reason field meridianlabs-ai/ts-mono#513 (see the coordinated ts-mono change section below).
  2. docs: add Scoring Policy page (score vs raise vs unscored) #4518 (dedicated scoring-policy docs page) is tracked separately and should land with or after this field per the issue's sequencing — see dragonstyle's review endorsing this PR's approach. Its docs table needs a pass for the narrowed 5-value taxonomy.
  3. Distinguish math() answer-extraction failures from wrong answers via Score.reason #4091 (math() reason, held for this PR per the coordination note above).
  4. The log viewer UI doesn't yet surface reason in the score detail view — separate follow-up.
  5. choice(): unparseable multiple-choice answers score INCORRECT with no reason today — parsing happens in the multiple_choice() solver, so attribution needs plumbing from solver to scorer.

Verification: tests/scorer + tests/analysis (targeted) and the full test suite (8176 passed, 0 failed, environment-only skips for --runtrio/vllm/slow markers) both pass; ruff format/ruff check/mypy --exclude tests/test_package are clean. The branch went through 11 plan tasks with a task-scoped implementer+reviewer pass each, plus a whole-branch review that found and fixed two Important issues before this PR was opened (a read-shim round-trip edge case that could resurrect a user-cleared reason, and epoch reducers dropping reason) — see commits 073b4f0bf and 9f7cbfe7d.

Implemented by Claude Code on Matt's behalf, per his design discussion on #4567.

Coordinated ts-mono change

Paired with meridianlabs-ai/ts-mono#513 (regenerated generated.ts for the new reason field). Two-phase landing: the gitlink currently pins the ts-mono branch head, so submodule-on-main stays red by design until #513 merges; then the gitlink is bumped to the merged main SHA and the gate re-run.

MattFisher and others added 13 commits July 24, 2026 15:41
Implements the first task of the score-reason feature:
- Adds ScoreReason type with 7 standard vocabulary values
- Adds reason field to Score model (optional, accepts ScoreReason | str)
- Updates Score.unscored() to accept reason parameter
- Exports ScoreReason from inspect_ai.scorer

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ponse_format reason (UKGovernmentBEIS#4567)

Reverts the NOANSWER behavior from UKGovernmentBEIS#3628 per the attribution policy in UKGovernmentBEIS#4567.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…f metadata (UKGovernmentBEIS#4567)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…NaN (UKGovernmentBEIS#4567)

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
)

- perplexity/target_perplexity unscorable-state tests: add `assert
  result is not None` before accessing `.reason` (Score | None union).
- test_df score_details test: annotate the literal dict as JsonValue
  so it matches score_details()'s declared parameter type instead of
  the narrower dict[str, dict[str, str]] mypy infers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two Important findings from the whole-branch review of Score.reason (UKGovernmentBEIS#4567):

- edit_score() now strips the legacy metadata["unscored_reason"] key
  whenever an edit explicitly touches `reason` (including clearing it to
  None). Without this, the exclude_none=True JSON round-trip made an
  explicit reason=None indistinguishable from "never set", so the next
  Score.model_validate() would resurrect the old reason via the
  _lift_unscored_reason read shim - silently reverting the user's edit.
  The metadata is rebuilt as a fresh dict rather than mutated in place,
  since it can alias edit.metadata / the emitted ScoreEditEvent / the
  synthesized pre-edit history snapshot.

- _reduced_score() and _nan_score() in the epoch reducers now carry
  `reason` through using the same "retain only if equal across all
  Scores" pattern already used for answer/explanation/metadata, so
  mean/median/mode/at_least/pass_at/pass_k no longer silently drop
  reason when every epoch is unscored for the same cause.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- test_multi_match_failure_with_target was missing the `result.reason is
  None` assertion its single-group sibling already has.
- _builtin-scorers.md still described pattern()'s unmatched-output case
  as returning NOANSWER; it now reads INCORRECT + reason, matching the
  behavior reverted in this branch (UKGovernmentBEIS#4567) and already updated in
  custom-scorers.qmd but missed here.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
MattFisher and others added 2 commits August 7, 2026 13:32
# Conflicts:
#	CHANGELOG.md
#	src/inspect_ai/scorer/_metric.py
#	tests/scorer/test_metric.py
…reed conflict resolution

Per the resolution on UKGovernmentBEIS#4091 (PR review thread), antnewman will re-target
math()'s reason handling onto Score.reason directly in UKGovernmentBEIS#4091 once this
branch's field lands, rather than this branch carrying the math()/test_math.py
slice. Reverts the math() portion of 7f0cb9f; pattern()/answer()/model_graded/
perplexity/multi_scorer changes are unaffected.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@MattFisher
MattFisher marked this pull request as ready for review August 7, 2026 03:57
MattFisher added a commit to MattFisher/inspect_ai that referenced this pull request Aug 7, 2026
Per dragonstyle's review (UKGovernmentBEIS#4518#pullrequestreview-4830631635): ScoreReason
was superseded upward into a typed Score field by UKGovernmentBEIS#4567/UKGovernmentBEIS#4629, so this page
(and the custom-scorers/builtin-scorers examples it cross-references) should
describe the final state rather than the interim metadata["reason"]
convention. Also updates the pattern()/answer() bullets in
_builtin-scorers.md to the post-revert INCORRECT+reason behavior (UKGovernmentBEIS#4629).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@dragonstyle dragonstyle self-assigned this Aug 9, 2026
# Conflicts:
#	CHANGELOG.md
#	tests/analysis/test_df.py
MattFisher and others added 2 commits August 12, 2026 10:33
Points the submodule at the head of meridianlabs-ai/ts-mono#513 and
rebuilds the embedded viewer bundle at that commit.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# Conflicts:
#	CHANGELOG.md
#	src/inspect_ai/_view/ts-mono
MattFisher and others added 3 commits August 24, 2026 19:02
…esponse

Model-attributed reasons (no_response, refusal) belong on scores that
stay in the denominator, where reason-consuming metrics can read them;
instrument reasons belong on unscored() exclusions. The empty-completion
detail stays in explanation.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

@dragonstyle dragonstyle left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving. The rework landed both asks from my last review — the 5-value taxonomy is right, every emission site lines up with it, and the corrected CHANGELOG/denominator claim is accurate. The edit/round-trip tests (especially the exclude_none resurrection guard) are excellent.

One heads-up: I pushed a small commit to the branch (ecbcfde) rather than another review round-trip — it settles the one question I'd flagged, and since the call was mine to make it seemed unfair to make you do the typing. It flips the empty-choices emission in perplexity()/target_perplexity() from no_response to scoring_failed (both scorers, both test assertions, one CHANGELOG clause). Feel free to push back if you disagree with either the change or the liberty.

The reasoning, since it generalizes beyond this PR: the reason groups carry denominator semantics. Model-attributed reasons (invalid_response_format, refusal, no_response) belong on scores that stay in the denominator — that's where reason-consuming metrics (e.g. a SimpleQA-style attempted/not-attempted trichotomy) can actually read them; a reason attached to Score.unscored() is invisible to every metric, since unscored samples never reach them. Instrument reasons (grader_failed, scoring_failed) belong on exclusions. Your no_response label was factually true — the completion is empty — but the treatment is an instrument-style exclusion, so the label and the treatment told different stories, and the factual detail lives happily in explanation. This coherence rule (reason group matches denominator treatment) is going into the #4518 policy page as the organizing principle, along with the note that reason=None is the success case by contract — we considered and rejected a success value, since totality is unenforceable for directly-constructed Scores. refusal staying at zero emitters is deliberate reserved vocabulary under the same frame.

Merge choreography as you described: ts-mono#513 first, then bump the gitlink to the merged SHA and re-run the gate — submodule-on-main is the expected red until then. I'll follow up on #4518 with the policy items and file two issues (scoring conformance corpus, mixed-version analysis warnings) that came out of this review. #4091 retargets after this merges.

dragonstyle pushed a commit to meridianlabs-ai/ts-mono that referenced this pull request Aug 26, 2026
* Regenerate types for Score.reason field

Adds the optional machine-readable reason field to Score, ScoreEvent,
and ScoreEdit, generated from inspect_ai's updated Pydantic models
(UKGovernmentBEIS/inspect_ai#4629).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Narrow ScoreReason enum to 5-value taxonomy

Regenerated after inspect_ai review feedback narrowed the vocabulary:
the grader_* values collapse to grader_failed, and scorer-cannot-run
cases become scoring_failed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
@dragonstyle
dragonstyle enabled auto-merge (squash) August 26, 2026 20:39
@dragonstyle
dragonstyle merged commit 4022900 into UKGovernmentBEIS:main Aug 26, 2026
18 of 20 checks passed
epatey added a commit to meridianlabs-ai/inspect_scout that referenced this pull request Aug 27, 2026
…es (#581)

Fixes the failing Nightly CI run
https://github.com/meridianlabs-ai/inspect_scout/actions/runs/33067365391
(`check-schema-and-types` and `py-mypy` red against inspect_ai git
main).

## Root cause: upstream drift

Three inspect_ai PRs landed on its main Aug 26–27:

- UKGovernmentBEIS/inspect_ai#4991 — a model role may now bind a
**list** of models. This widened `resolve_model_roles`,
`model_roles_to_model_roles_config`,
`model_roles_config_to_model_roles`, and `init_model_context` to
`dict[str, Model | list[Model]]` / `dict[str, ModelConfig |
list[ModelConfig]]` → the 11 mypy errors.
- UKGovernmentBEIS/inspect_ai#4624 — sandbox snapshot strategies → new
`SandboxSnapshotConfig`/`ArchiveSnapshots`/`ResticSnapshots` schemas.
- UKGovernmentBEIS/inspect_ai#4629 — first-class `Score.reason` → new
`reason` field in score schemas.

The latter two show up in the regenerated `openapi.json` via embedded
inspect_ai types (structural drift, not tolerated as docstring-only).

## Fix

- Widen scout's `model_roles` plumbing to match upstream: public params
take `ModelRoles` (inspect_ai's alias, a supertype of the old `dict[str,
str | Model]`), `ScanJob.model_roles` / `init_scan_model_context` return
`dict[str, Model | list[Model]]`, and `ScanSpec.model_roles` /
`IPCContext.model_roles` hold `ModelConfig | list[ModelConfig]`. Scout
only plumbs roles through to inspect_ai (resolution and
`get_model(role=...)` live upstream), so list roles work through scout
with no behavior change for existing callers.
- Regenerate `openapi.json` and `generated.ts` against inspect_ai main.
- Bump the ts-mono gitlink and rebuild the viewer `dist` (verified the
old pinned commit rebuilds byte-identically locally before trusting the
new build).

Paired ts-mono PR: meridianlabs-ai/ts-mono#587 (two-phase landing:
`submodule-on-main` stays red here until that merges).

## Verification

Local venv with inspect_ai @ git main:
- `mypy src examples tests` — clean (was 11 errors, exact repro of CI)
- `.github/scripts/check_openapi_drift.py` — passes
- `pytest` — 3041 passed, 107 skipped
- ts-mono `pnpm typecheck` — clean; `pnpm test` has a pre-existing
failure in `@tsmono/inspect-components` timeline fixtures that
reproduces on clean ts-mono main locally (ts-mono CI on the same SHA is
green — environmental, unrelated to this change)
epatey added a commit that referenced this pull request Aug 27, 2026
…t leg printing output for passed tests (#5075)

* ci-perf: snapshot for 2026-08-27

200 completed PR workflow runs, 2026-08-26 14:44 .. 2026-08-27 10:51 UTC,
with per-job and per-step timings, pytest --durations from 20 test-job
logs and pytest summary lines for suite-size trends.

Collected on the third attempt: the first two returned an identical bad
window (a stale page 1 serving a 2026-08-17/18 clump ahead of a correct
page 2, leaving a 204h hole in a 220h span) while a direct call to the
same endpoint in between returned a contiguous page 1.

* ci-perf: 2026-08-27 report

Headline: -rA costs ~46s of every test leg (14% of the pytest step)
printing captured stdout for tests that passed. Measured by timestamp
from the raw job logs of 16 test legs; filed as meridianlabs-ai#342.

Also resolves two open questions from the last report: the test-leg
creep is suite growth plus noise (+674 collected items in two days,
against a newly measured +/-60s-of-worker-time noise floor), and the
viewer component-test regression is gone as of the ts-mono bump that
rode in with #4629.

* ci-perf: ledger entries for this run and the 2026-08-25 PR

Records the upstream promotion of meridianlabs-ai#312 as #5057 (merged),
and this run's entry.

* ci-perf: record this run's PR and the upstream-write probe result

Upstream PR creation returned 403 again, so the run's output is fork PR
meridianlabs-ai#343 awaiting maintainer promotion.

* ci-perf: record this PR's own #299 observation (Build wall 102s, 7.2 runner-min)

* pytest -rA -> -ra: drop passed-test captured-output dump (~46s/test leg)

Fixes #342. Both the build.yml command line and pyproject addopts (command
line wins over addopts in CI; addopts covers local runs). Also the
sandbox-tools job's -rA for consistency. -ra still reports failures,
errors, skips and xfails.

Pushed from a human-driven session on Eric's behalf — the scheduled run's
token cannot push workflow files.

---------

Co-authored-by: i-am-marvin <294938113+i-am-marvin@users.noreply.github.com>
antnewman added a commit to antnewman/inspect_ai that referenced this pull request Aug 28, 2026
Re-targets this PR's math() fix from the interim metadata["reason"] convention onto the first-class Score.reason field landed by UKGovernmentBEIS#4629, per the agreed split that keeps math() as this PR's slice of UKGovernmentBEIS#4567. The fine-grained math_scorer_status metadata from UKGovernmentBEIS#4361 is left untouched alongside it.
dragonstyle pushed a commit to tonydzi/inspect_ai that referenced this pull request Aug 31, 2026
The panel audit trail read only metadata["unscored_reason"] (UKGovernmentBEIS#4048). UKGovernmentBEIS#4629
promotes that to a first-class Score.reason and stops writing the legacy key,
which left panel.failures[].reason as None for every failed grader.

Read the first-class field where present, falling back to the legacy key, so
this lands correctly whichever of the two PRs merges first. Behavior on
today's main is unchanged: Score has no reason attribute, so the fallback is
the only live path.

Checked with `is None` rather than truthiness, so an explicitly-empty reason
is not overridden by a stale legacy value.

Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
dragonstyle added a commit that referenced this pull request Sep 1, 2026
* fix(scorer): decide a grader panel by strict majority, not mode

A grader whose completion has no parseable grade returns Score.unscored(),
which the mode reducer filters out of the vote. An odd panel therefore
becomes even at runtime, and Counter.most_common breaks the resulting tie
by insertion order, so the grade follows the order of model=[...].

Add a `majority` reducer: a value wins only with more than half of the
scores reduced, and unscored members withhold a vote without lowering that
threshold, so a shrunken panel is unscored rather than decided by position.
The reduced score records the votes and the failures under a `panel`
metadata key. model_graded_qa()/model_graded_fact() default to it and take
`reducer="mode"` to reproduce existing scores.

Fixes #4721

* review: copy container votes, pin dict/list denominator, strengthen control test

* review pass 2: keep the failed grader's explanation, document the panel key, cover both aliasing directions

* Read the panel failure reason from Score.reason where it exists

The panel audit trail read only metadata["unscored_reason"] (#4048). #4629
promotes that to a first-class Score.reason and stops writing the legacy key,
which left panel.failures[].reason as None for every failed grader.

Read the first-class field where present, falling back to the legacy key, so
this lands correctly whichever of the two PRs merges first. Behavior on
today's main is unchanged: Score has no reason attribute, so the fallback is
the only live path.

Checked with `is None` rather than truthiness, so an explicitly-empty reason
is not overridden by a stale legacy value.

Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Post-rebase reconciliation with main's model-role and reason taxonomy

Thread reducer through both panel paths main gained since this branch
(the model-list branch and the role-bound-list fan-out, which hardcoded
mode); expect the grader_failed taxonomy value for unparseable verdicts
(main's convention post-#4908); list majority_score in the hand-
maintained reference docs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Tighten panel/reducer docs

Keep majority-specific detail out of the general reducers callout,
align the model_role precedence item with the reducer phrasing, and
compress the Multiple Models and multi_scorer passages.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* tests: narrow types instead of type: ignore suppressions

---------

Co-authored-by: Palo-Alto-AI-Research-Lab <194927794+Palo-Alto-AI-Research-Lab@users.noreply.github.com>
Co-authored-by: Anton Dziatkovskii <194927794+tonydzi@users.noreply.github.com>
Co-authored-by: dragonstyle <cteague@gmail.com>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Charles Teague <charles@meridianlabs.ai>
dragonstyle added a commit that referenced this pull request Sep 4, 2026
…Score.reason (#4091)

* fix(scorer): math() returns NOANSWER (not INCORRECT) when extract_answer fails

Sister fix to #4048 (model_graded_qa / model_graded_fact), addressing
the same silent-failure family flagged in #4026.

When extract_answer() returns None — the model output contains no
parseable mathematical expression — the scorer's else-branch returned
Score(value=INCORRECT, answer="None"). This conflated two semantically
different cases:

- "the model answered, and was wrong"
- "the model didn't produce a parseable answer at all"

In aggregate metrics the two became indistinguishable: a sample where
the model failed to format a math answer was silently counted as a
wrong answer, understating both true accuracy and parse-failure rate.

The fix returns NOANSWER (the same sentinel pattern() uses for
unparseable output, and the same convention #4048 adopts for
model_graded scorers) and preserves the original completion as the
`answer` field for downstream observability.

Two existing tests asserted the pre-fix behaviour and have been
updated to expect NOANSWER instead:
- test_no_boxed "not a number" case: directly exercises the parse-fail
  path, now correctly asserts NOANSWER.
- test_infinite_float_does_not_crash: the intent was "no crash on inf"
  — extract_answer returns None for infinity, so the parse-fail branch
  applies; NOANSWER is the right semantic.

Added test_parse_failure_yields_noanswer (parametrised over 4
representative parse-fail completions) and
test_parse_success_with_wrong_value_yields_incorrect (sanity check
that genuine wrong answers still get INCORRECT).

* fix(scorer): tag math() answer-extraction failures with reason metadata

When extract_answer() cannot parse a mathematical expression from the
completion, keep value=INCORRECT (a format violation is the model
failing the task and must not inflate accuracy by leaving the
denominator) and record metadata={"reason": "invalid_response_format"}
so analysis can separate "wrong answer" from "couldn't parse an
answer". Follows the convention discussed in #4091 in place of the
earlier NOANSWER approach.

Also stop recording the literal string "None" as the extracted answer
on parse failure (answer is now None with an explanatory message).

* fix(scorer): surface the full completion as answer on math() parse failure

Per review: the scorer docs ask extraction scorers to always return an answer, so the parse-failure branch now sets answer=state.output.completion and drops the completion from the (now static) explanation.

* fix(scorer): reapply parse-failure reason tagging onto rewritten math scorer

The #4361 rewrite already keeps INCORRECT on extraction failure with an internal math_scorer_status key; this adds the #4091/#4567 policy tag (reason=invalid_response_format) alongside it on the answer_parse_error/answer_limit branch, and surfaces the completion as the answer when no candidate was extracted. CHANGELOG entry restored to Unreleased after the merge relocated it.

* fix(scorer): record math() extraction failures in Score.reason

Re-targets this PR's math() fix from the interim metadata["reason"] convention onto the first-class Score.reason field landed by #4629, per the agreed split that keeps math() as this PR's slice of #4567. The fine-grained math_scorer_status metadata from #4361 is left untouched alongside it.

* docs(changelog): put the math() entry under a new Unreleased section

Merging upstream/main applied the entry cleanly but in the wrong place: the
neighbouring #4567 entries it was anchored against moved into the 0.3.261
section when that version was cut, and this line travelled with them.

The changelog-lint workflow requires new entries to sit under a '## Unreleased'
heading that is first in the file. The 0.3.262 release cut renamed the previous
Unreleased heading, so this restores one.

---------

Co-authored-by: Charles Teague <cteague@gmail.com>
dragonstyle added a commit that referenced this pull request Sep 9, 2026
* Recommend model-roles in custom-scorers docs

* Scoring policy docs v1

* Document pattern vs match outcomes for non-matching output

match/includes/exact score a non-match INCORRECT; pattern/answer score
NOANSWER when the pattern does not match at all. Note both map to 0.0 and
cross-reference Scoring Policy for the wrong-answer vs no-answer distinction.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Score grader parse failures as unscored in custom-scorers example

The model_graded_qa example returned INCORRECT when the grader's verdict
could not be parsed, charging an instrument failure to the model. Return
Score.unscored() instead, and link Scoring Policy for when to score,
raise, or unscore.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Link scoring policy for unparseable grader verdicts

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Link scoring policy from handling-errors

Pairs the run-level error options with the scorer authoring decision.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Restructure scoring policy to lead with concept, and cross-reference siblings

Lead with what scorers do and why before the prescriptive "if you're
writing a scorer" guidance, and cross-reference custom-scorers,
model-graded, and metrics rather than restating them.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Include answer in examples

* Fix typos

* Describe Score.reason as a first-class field, not a metadata convention

Per dragonstyle's review (#4518#pullrequestreview-4830631635): ScoreReason
was superseded upward into a typed Score field by #4567/#4629, so this page
(and the custom-scorers/builtin-scorers examples it cross-references) should
describe the final state rather than the interim metadata["reason"]
convention. Also updates the pattern()/answer() bullets in
_builtin-scorers.md to the post-revert INCORRECT+reason behavior (#4629).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Align scoring policy with final ScoreReason taxonomy

* Clarify scorer outcomes and metric coverage

Preserve the author's scoring policy structure while correcting scorer behavior, unscored values, epoch coverage, and error handling. Keep custom-scorer advice in custom-scorers.qmd and remove recommendations about evaluation design and reporting.

Includes the intended choice and math behavior from #5310 and #5311, pending when verified. Validation: lint, local documentation links and anchors, Markdown fences, and git diff --check passed. Targeted behavior tests against local main: 197 passed, 17 skipped. Quarto unavailable.

* Fix unparsable spelling in scoring documentation

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: dragonstyle <cteague@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

First-class Score.reason and normalized scorer failure policy

3 participants