Repository navigation
First-class Score.reason and normalized scorer failure policy - #4629
Conversation
Implements the first task of the score-reason feature: - Adds ScoreReason type with 7 standard vocabulary values - Adds reason field to Score model (optional, accepts ScoreReason | str) - Updates Score.unscored() to accept reason parameter - Exports ScoreReason from inspect_ai.scorer Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…KGovernmentBEIS#4567) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rnmentBEIS#4567) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ponse_format reason (UKGovernmentBEIS#4567) Reverts the NOANSWER behavior from UKGovernmentBEIS#3628 per the attribution policy in UKGovernmentBEIS#4567. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…overnmentBEIS#4567, UKGovernmentBEIS#4091) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…f metadata (UKGovernmentBEIS#4567) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…NaN (UKGovernmentBEIS#4567) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ernmentBEIS#4567) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rnmentBEIS#4567) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…vernmentBEIS#4567) Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
) - perplexity/target_perplexity unscorable-state tests: add `assert result is not None` before accessing `.reason` (Score | None union). - test_df score_details test: annotate the literal dict as JsonValue so it matches score_details()'s declared parameter type instead of the narrower dict[str, dict[str, str]] mypy infers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two Important findings from the whole-branch review of Score.reason (UKGovernmentBEIS#4567): - edit_score() now strips the legacy metadata["unscored_reason"] key whenever an edit explicitly touches `reason` (including clearing it to None). Without this, the exclude_none=True JSON round-trip made an explicit reason=None indistinguishable from "never set", so the next Score.model_validate() would resurrect the old reason via the _lift_unscored_reason read shim - silently reverting the user's edit. The metadata is rebuilt as a fresh dict rather than mutated in place, since it can alias edit.metadata / the emitted ScoreEditEvent / the synthesized pre-edit history snapshot. - _reduced_score() and _nan_score() in the epoch reducers now carry `reason` through using the same "retain only if equal across all Scores" pattern already used for answer/explanation/metadata, so mean/median/mode/at_least/pass_at/pass_k no longer silently drop reason when every epoch is unscored for the same cause. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- test_multi_match_failure_with_target was missing the `result.reason is None` assertion its single-group sibling already has. - _builtin-scorers.md still described pattern()'s unmatched-output case as returning NOANSWER; it now reads INCORRECT + reason, matching the behavior reverted in this branch (UKGovernmentBEIS#4567) and already updated in custom-scorers.qmd but missed here. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md # src/inspect_ai/scorer/_metric.py # tests/scorer/test_metric.py
…reed conflict resolution Per the resolution on UKGovernmentBEIS#4091 (PR review thread), antnewman will re-target math()'s reason handling onto Score.reason directly in UKGovernmentBEIS#4091 once this branch's field lands, rather than this branch carrying the math()/test_math.py slice. Reverts the math() portion of 7f0cb9f; pattern()/answer()/model_graded/ perplexity/multi_scorer changes are unaffected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Per dragonstyle's review (UKGovernmentBEIS#4518#pullrequestreview-4830631635): ScoreReason was superseded upward into a typed Score field by UKGovernmentBEIS#4567/UKGovernmentBEIS#4629, so this page (and the custom-scorers/builtin-scorers examples it cross-references) should describe the final state rather than the interim metadata["reason"] convention. Also updates the pattern()/answer() bullets in _builtin-scorers.md to the post-revert INCORRECT+reason behavior (UKGovernmentBEIS#4629). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md # tests/analysis/test_df.py
Points the submodule at the head of meridianlabs-ai/ts-mono#513 and rebuilds the embedded viewer bundle at that commit. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md # src/inspect_ai/_view/ts-mono
…esponse Model-attributed reasons (no_response, refusal) belong on scores that stay in the denominator, where reason-consuming metrics can read them; instrument reasons belong on unscored() exclusions. The empty-completion detail stays in explanation. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
dragonstyle
left a comment
There was a problem hiding this comment.
Approving. The rework landed both asks from my last review — the 5-value taxonomy is right, every emission site lines up with it, and the corrected CHANGELOG/denominator claim is accurate. The edit/round-trip tests (especially the exclude_none resurrection guard) are excellent.
One heads-up: I pushed a small commit to the branch (ecbcfde) rather than another review round-trip — it settles the one question I'd flagged, and since the call was mine to make it seemed unfair to make you do the typing. It flips the empty-choices emission in perplexity()/target_perplexity() from no_response to scoring_failed (both scorers, both test assertions, one CHANGELOG clause). Feel free to push back if you disagree with either the change or the liberty.
The reasoning, since it generalizes beyond this PR: the reason groups carry denominator semantics. Model-attributed reasons (invalid_response_format, refusal, no_response) belong on scores that stay in the denominator — that's where reason-consuming metrics (e.g. a SimpleQA-style attempted/not-attempted trichotomy) can actually read them; a reason attached to Score.unscored() is invisible to every metric, since unscored samples never reach them. Instrument reasons (grader_failed, scoring_failed) belong on exclusions. Your no_response label was factually true — the completion is empty — but the treatment is an instrument-style exclusion, so the label and the treatment told different stories, and the factual detail lives happily in explanation. This coherence rule (reason group matches denominator treatment) is going into the #4518 policy page as the organizing principle, along with the note that reason=None is the success case by contract — we considered and rejected a success value, since totality is unenforceable for directly-constructed Scores. refusal staying at zero emitters is deliberate reserved vocabulary under the same frame.
Merge choreography as you described: ts-mono#513 first, then bump the gitlink to the merged SHA and re-run the gate — submodule-on-main is the expected red until then. I'll follow up on #4518 with the policy items and file two issues (scoring conformance corpus, mixed-version analysis warnings) that came out of this review. #4091 retargets after this merges.
* Regenerate types for Score.reason field Adds the optional machine-readable reason field to Score, ScoreEvent, and ScoreEdit, generated from inspect_ai's updated Pydantic models (UKGovernmentBEIS/inspect_ai#4629). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Narrow ScoreReason enum to 5-value taxonomy Regenerated after inspect_ai review feedback narrowed the vocabulary: the grader_* values collapse to grader_failed, and scorer-cannot-run cases become scoring_failed. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md
…es (#581) Fixes the failing Nightly CI run https://github.com/meridianlabs-ai/inspect_scout/actions/runs/33067365391 (`check-schema-and-types` and `py-mypy` red against inspect_ai git main). ## Root cause: upstream drift Three inspect_ai PRs landed on its main Aug 26–27: - UKGovernmentBEIS/inspect_ai#4991 — a model role may now bind a **list** of models. This widened `resolve_model_roles`, `model_roles_to_model_roles_config`, `model_roles_config_to_model_roles`, and `init_model_context` to `dict[str, Model | list[Model]]` / `dict[str, ModelConfig | list[ModelConfig]]` → the 11 mypy errors. - UKGovernmentBEIS/inspect_ai#4624 — sandbox snapshot strategies → new `SandboxSnapshotConfig`/`ArchiveSnapshots`/`ResticSnapshots` schemas. - UKGovernmentBEIS/inspect_ai#4629 — first-class `Score.reason` → new `reason` field in score schemas. The latter two show up in the regenerated `openapi.json` via embedded inspect_ai types (structural drift, not tolerated as docstring-only). ## Fix - Widen scout's `model_roles` plumbing to match upstream: public params take `ModelRoles` (inspect_ai's alias, a supertype of the old `dict[str, str | Model]`), `ScanJob.model_roles` / `init_scan_model_context` return `dict[str, Model | list[Model]]`, and `ScanSpec.model_roles` / `IPCContext.model_roles` hold `ModelConfig | list[ModelConfig]`. Scout only plumbs roles through to inspect_ai (resolution and `get_model(role=...)` live upstream), so list roles work through scout with no behavior change for existing callers. - Regenerate `openapi.json` and `generated.ts` against inspect_ai main. - Bump the ts-mono gitlink and rebuild the viewer `dist` (verified the old pinned commit rebuilds byte-identically locally before trusting the new build). Paired ts-mono PR: meridianlabs-ai/ts-mono#587 (two-phase landing: `submodule-on-main` stays red here until that merges). ## Verification Local venv with inspect_ai @ git main: - `mypy src examples tests` — clean (was 11 errors, exact repro of CI) - `.github/scripts/check_openapi_drift.py` — passes - `pytest` — 3041 passed, 107 skipped - ts-mono `pnpm typecheck` — clean; `pnpm test` has a pre-existing failure in `@tsmono/inspect-components` timeline fixtures that reproduces on clean ts-mono main locally (ts-mono CI on the same SHA is green — environmental, unrelated to this change)
…t leg printing output for passed tests (#5075) * ci-perf: snapshot for 2026-08-27 200 completed PR workflow runs, 2026-08-26 14:44 .. 2026-08-27 10:51 UTC, with per-job and per-step timings, pytest --durations from 20 test-job logs and pytest summary lines for suite-size trends. Collected on the third attempt: the first two returned an identical bad window (a stale page 1 serving a 2026-08-17/18 clump ahead of a correct page 2, leaving a 204h hole in a 220h span) while a direct call to the same endpoint in between returned a contiguous page 1. * ci-perf: 2026-08-27 report Headline: -rA costs ~46s of every test leg (14% of the pytest step) printing captured stdout for tests that passed. Measured by timestamp from the raw job logs of 16 test legs; filed as meridianlabs-ai#342. Also resolves two open questions from the last report: the test-leg creep is suite growth plus noise (+674 collected items in two days, against a newly measured +/-60s-of-worker-time noise floor), and the viewer component-test regression is gone as of the ts-mono bump that rode in with #4629. * ci-perf: ledger entries for this run and the 2026-08-25 PR Records the upstream promotion of meridianlabs-ai#312 as #5057 (merged), and this run's entry. * ci-perf: record this run's PR and the upstream-write probe result Upstream PR creation returned 403 again, so the run's output is fork PR meridianlabs-ai#343 awaiting maintainer promotion. * ci-perf: record this PR's own #299 observation (Build wall 102s, 7.2 runner-min) * pytest -rA -> -ra: drop passed-test captured-output dump (~46s/test leg) Fixes #342. Both the build.yml command line and pyproject addopts (command line wins over addopts in CI; addopts covers local runs). Also the sandbox-tools job's -rA for consistency. -ra still reports failures, errors, skips and xfails. Pushed from a human-driven session on Eric's behalf — the scheduled run's token cannot push workflow files. --------- Co-authored-by: i-am-marvin <294938113+i-am-marvin@users.noreply.github.com>
Re-targets this PR's math() fix from the interim metadata["reason"] convention onto the first-class Score.reason field landed by UKGovernmentBEIS#4629, per the agreed split that keeps math() as this PR's slice of UKGovernmentBEIS#4567. The fine-grained math_scorer_status metadata from UKGovernmentBEIS#4361 is left untouched alongside it.
The panel audit trail read only metadata["unscored_reason"] (UKGovernmentBEIS#4048). UKGovernmentBEIS#4629 promotes that to a first-class Score.reason and stops writing the legacy key, which left panel.failures[].reason as None for every failed grader. Read the first-class field where present, falling back to the legacy key, so this lands correctly whichever of the two PRs merges first. Behavior on today's main is unchanged: Score has no reason attribute, so the fallback is the only live path. Checked with `is None` rather than truthiness, so an explicitly-empty reason is not overridden by a stale legacy value. Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(scorer): decide a grader panel by strict majority, not mode A grader whose completion has no parseable grade returns Score.unscored(), which the mode reducer filters out of the vote. An odd panel therefore becomes even at runtime, and Counter.most_common breaks the resulting tie by insertion order, so the grade follows the order of model=[...]. Add a `majority` reducer: a value wins only with more than half of the scores reduced, and unscored members withhold a vote without lowering that threshold, so a shrunken panel is unscored rather than decided by position. The reduced score records the votes and the failures under a `panel` metadata key. model_graded_qa()/model_graded_fact() default to it and take `reducer="mode"` to reproduce existing scores. Fixes #4721 * review: copy container votes, pin dict/list denominator, strengthen control test * review pass 2: keep the failed grader's explanation, document the panel key, cover both aliasing directions * Read the panel failure reason from Score.reason where it exists The panel audit trail read only metadata["unscored_reason"] (#4048). #4629 promotes that to a first-class Score.reason and stops writing the legacy key, which left panel.failures[].reason as None for every failed grader. Read the first-class field where present, falling back to the legacy key, so this lands correctly whichever of the two PRs merges first. Behavior on today's main is unchanged: Score has no reason attribute, so the fallback is the only live path. Checked with `is None` rather than truthiness, so an explicitly-empty reason is not overridden by a stale legacy value. Assisted-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * Post-rebase reconciliation with main's model-role and reason taxonomy Thread reducer through both panel paths main gained since this branch (the model-list branch and the role-bound-list fan-out, which hardcoded mode); expect the grader_failed taxonomy value for unparseable verdicts (main's convention post-#4908); list majority_score in the hand- maintained reference docs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Tighten panel/reducer docs Keep majority-specific detail out of the general reducers callout, align the model_role precedence item with the reducer phrasing, and compress the Multiple Models and multi_scorer passages. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * tests: narrow types instead of type: ignore suppressions --------- Co-authored-by: Palo-Alto-AI-Research-Lab <194927794+Palo-Alto-AI-Research-Lab@users.noreply.github.com> Co-authored-by: Anton Dziatkovskii <194927794+tonydzi@users.noreply.github.com> Co-authored-by: dragonstyle <cteague@gmail.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Co-authored-by: Charles Teague <charles@meridianlabs.ai>
…Score.reason (#4091) * fix(scorer): math() returns NOANSWER (not INCORRECT) when extract_answer fails Sister fix to #4048 (model_graded_qa / model_graded_fact), addressing the same silent-failure family flagged in #4026. When extract_answer() returns None — the model output contains no parseable mathematical expression — the scorer's else-branch returned Score(value=INCORRECT, answer="None"). This conflated two semantically different cases: - "the model answered, and was wrong" - "the model didn't produce a parseable answer at all" In aggregate metrics the two became indistinguishable: a sample where the model failed to format a math answer was silently counted as a wrong answer, understating both true accuracy and parse-failure rate. The fix returns NOANSWER (the same sentinel pattern() uses for unparseable output, and the same convention #4048 adopts for model_graded scorers) and preserves the original completion as the `answer` field for downstream observability. Two existing tests asserted the pre-fix behaviour and have been updated to expect NOANSWER instead: - test_no_boxed "not a number" case: directly exercises the parse-fail path, now correctly asserts NOANSWER. - test_infinite_float_does_not_crash: the intent was "no crash on inf" — extract_answer returns None for infinity, so the parse-fail branch applies; NOANSWER is the right semantic. Added test_parse_failure_yields_noanswer (parametrised over 4 representative parse-fail completions) and test_parse_success_with_wrong_value_yields_incorrect (sanity check that genuine wrong answers still get INCORRECT). * fix(scorer): tag math() answer-extraction failures with reason metadata When extract_answer() cannot parse a mathematical expression from the completion, keep value=INCORRECT (a format violation is the model failing the task and must not inflate accuracy by leaving the denominator) and record metadata={"reason": "invalid_response_format"} so analysis can separate "wrong answer" from "couldn't parse an answer". Follows the convention discussed in #4091 in place of the earlier NOANSWER approach. Also stop recording the literal string "None" as the extracted answer on parse failure (answer is now None with an explanatory message). * fix(scorer): surface the full completion as answer on math() parse failure Per review: the scorer docs ask extraction scorers to always return an answer, so the parse-failure branch now sets answer=state.output.completion and drops the completion from the (now static) explanation. * fix(scorer): reapply parse-failure reason tagging onto rewritten math scorer The #4361 rewrite already keeps INCORRECT on extraction failure with an internal math_scorer_status key; this adds the #4091/#4567 policy tag (reason=invalid_response_format) alongside it on the answer_parse_error/answer_limit branch, and surfaces the completion as the answer when no candidate was extracted. CHANGELOG entry restored to Unreleased after the merge relocated it. * fix(scorer): record math() extraction failures in Score.reason Re-targets this PR's math() fix from the interim metadata["reason"] convention onto the first-class Score.reason field landed by #4629, per the agreed split that keeps math() as this PR's slice of #4567. The fine-grained math_scorer_status metadata from #4361 is left untouched alongside it. * docs(changelog): put the math() entry under a new Unreleased section Merging upstream/main applied the entry cleanly but in the wrong place: the neighbouring #4567 entries it was anchored against moved into the 0.3.261 section when that version was cut, and this line travelled with them. The changelog-lint workflow requires new entries to sit under a '## Unreleased' heading that is first in the file. The 0.3.262 release cut renamed the previous Unreleased heading, so this restores one. --------- Co-authored-by: Charles Teague <cteague@gmail.com>
* Recommend model-roles in custom-scorers docs * Scoring policy docs v1 * Document pattern vs match outcomes for non-matching output match/includes/exact score a non-match INCORRECT; pattern/answer score NOANSWER when the pattern does not match at all. Note both map to 0.0 and cross-reference Scoring Policy for the wrong-answer vs no-answer distinction. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Score grader parse failures as unscored in custom-scorers example The model_graded_qa example returned INCORRECT when the grader's verdict could not be parsed, charging an instrument failure to the model. Return Score.unscored() instead, and link Scoring Policy for when to score, raise, or unscore. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Link scoring policy for unparseable grader verdicts Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Link scoring policy from handling-errors Pairs the run-level error options with the scorer authoring decision. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Restructure scoring policy to lead with concept, and cross-reference siblings Lead with what scorers do and why before the prescriptive "if you're writing a scorer" guidance, and cross-reference custom-scorers, model-graded, and metrics rather than restating them. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Include answer in examples * Fix typos * Describe Score.reason as a first-class field, not a metadata convention Per dragonstyle's review (#4518#pullrequestreview-4830631635): ScoreReason was superseded upward into a typed Score field by #4567/#4629, so this page (and the custom-scorers/builtin-scorers examples it cross-references) should describe the final state rather than the interim metadata["reason"] convention. Also updates the pattern()/answer() bullets in _builtin-scorers.md to the post-revert INCORRECT+reason behavior (#4629). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Align scoring policy with final ScoreReason taxonomy * Clarify scorer outcomes and metric coverage Preserve the author's scoring policy structure while correcting scorer behavior, unscored values, epoch coverage, and error handling. Keep custom-scorer advice in custom-scorers.qmd and remove recommendations about evaluation design and reporting. Includes the intended choice and math behavior from #5310 and #5311, pending when verified. Validation: lint, local documentation links and anchors, Markdown fences, and git diff --check passed. Targeted behavior tests against local main: 197 passed, 17 skipped. Quarto unavailable. * Fix unparsable spelling in scoring documentation --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: dragonstyle <cteague@gmail.com>
This PR contains:
What is the current behavior? (You can also link to an open issue here)
Closes #4567.
Built-in scorers record "couldn't get a usable reading" three different ways:
pattern()/answer()returnNOANSWER,math()returnsINCORRECTwith ananswer="None"string artefact, andmodel_graded_*returnsScore.unscored()with an ad hocmetadata["unscored_reason"]key. There's no first-class field for why a score is abnormal, separate from the value (which decides what the metric sees) and the human-readableexplanation. The interimmetadata["reason"]convention is fragile:apply_score_edit-style edits replacemetadatawholesale and silently drop it, and it squats on the author's own metadata namespace.What is the new behavior?
Score.reason: ScoreReason | str | None— a new field recording a machine-readable failure mode, separate fromvalue/explanation/metadata.ScoreReasonis an open-union Literal with a 5-value attribution taxonomy — model under test:invalid_response_format,refusal,no_response(kept fine-grained; refusal rate is a first-class quantity); measurement instrument:grader_failed(grader model failed),scoring_failed(scorer could not run) — coarse deliberately, with the detail inexplanation. The standard vocabulary is discoverable via IDE completion, but any custom string is legal, and adding values later is trivially non-breaking while removing them never is.Score.unscored(reason=...)— the primary construction site for instrument (grader/harness) failures.ScoreEdit.reason— follows the sameUNCHANGED-sentinel pattern asanswer/explanation/metadata, soreasonsurvives score edits by construction (edits no longer silently drop it the way wholesale-metadata-replacement does).INCORRECT(in the metric denominator) withreason="invalid_response_format":pattern()/answer()revert fromNOANSWER(Pattern scorer noanswer #3628) back toINCORRECT+reason. (math()'s equivalent fix is deliberately not in this PR — see coordination note below.)Score.unscored(reason=...), excluded from the denominator:model_graded_qa/model_graded_factmove their grade-parse-failure reason frommetadata["unscored_reason"](fix(scorer): mark model_graded grade-parse failure as unscored (not INCORRECT) #4048) into thereasonfield asgrader_failed;perplexity()/target_perplexity()returnScore.unscored(reason="scoring_failed")instead of a raw NaNScore(no_responsewhen the completion is empty);multi_scorer()setsreason="scoring_failed"when every sub-scorer declines.metadata["unscored_reason"](written by 0.3.245+) intoreasonat read time, without touching the metadata key — old logs read cleanly with no migration needed. A follow-up fix ensures an edit that explicitly clearsreasonalso drops the stalemetadata["unscored_reason"]key, so a log round-trip (write → read) can't resurrect a reason the user just cleared.mean_score,median_score,mode_score,at_least,pass_at,pass_k) andmulti_scorer's reducer path now retainreasonacross epochs (same "equal across all scores, elseNone" rule already used foranswer/explanation).score_<name>_reasoncolumns now appear insamples_df()/score_details()output.inspect-openapi.jsonregenerated to includereasononScore,ScoreEdit, andEvalSampleScore.Does this PR introduce a breaking change? (What changes might users need to make in their application due to this PR?)
Yes, one deliberate behavior change flagged in the design discussion on #4567:
pattern()andanswer()now score unmatched model output asINCORRECTinstead ofNOANSWER(reverting #3628), with the failure mode now recorded inreason="invalid_response_format". Note this changes no default metric numbers: the defaultvalue_to_float()already mapsNOANSWERto 0, so format violations were already in the denominator at value 0. The change matters only for analyses that filter onvalue == "N"and for customvalue_to_float()mappings that treat noanswer differently from incorrect — those should key onreasoninstead.Everything else is additive: new optional fields default to
None/UNCHANGED, and the read shim keeps old logs compatible with no migration step.Other information:
Coordination note:
math()'s answer-extraction-failure fix is deliberately not included here, per the agreed resolution on #4091 — @antnewman will re-target that PR'smath()fix ontoScore.reasononce this PR merges, rather than this PR carrying a competing edit to_math.py/test_math.py.Known follow-ups (out of scope for this PR):
The ts-mono viewer submodule's generated TypeScript types need a separate regeneration PR in that repo— done: Regenerate types for Score.reason field meridianlabs-ai/ts-mono#513 (see the coordinated ts-mono change section below).math()reason, held for this PR per the coordination note above).reasonin the score detail view — separate follow-up.choice(): unparseable multiple-choice answers scoreINCORRECTwith noreasontoday — parsing happens in themultiple_choice()solver, so attribution needs plumbing from solver to scorer.Verification:
tests/scorer+tests/analysis(targeted) and the full test suite (8176 passed,0 failed, environment-only skips for--runtrio/vllm/slow markers) both pass;ruff format/ruff check/mypy --exclude tests/test_packageare clean. The branch went through 11 plan tasks with a task-scoped implementer+reviewer pass each, plus a whole-branch review that found and fixed two Important issues before this PR was opened (a read-shim round-trip edge case that could resurrect a user-clearedreason, and epoch reducers droppingreason) — see commits073b4f0bfand9f7cbfe7d.Implemented by Claude Code on Matt's behalf, per his design discussion on #4567.
Coordinated ts-mono change
Paired with meridianlabs-ai/ts-mono#513 (regenerated
generated.tsfor the newreasonfield). Two-phase landing: the gitlink currently pins the ts-mono branch head, sosubmodule-on-mainstays red by design until #513 merges; then the gitlink is bumped to the merged main SHA and the gate re-run.