Skip to content

Commit 83fe609

Browse files
authored
[grading] Retry API errors, record grading errors, and read Word comments (#173)
* [judge] retry transient API errors and stop sending temperature to gpt-5.5 Claude judge requests now stream with a 64,000-token output cap and retry 429, 500, 503, and 529 responses, overload errors, and structured-output grammar compilation timeouts with linear backoff, up to six attempts. After repeated 500s the existing fallback still retries without output_config. OpenAI judges use the same output cap. gpt-5.5, gpt-5.5-pro, gpt-5.4-pro, and their dated snapshots reject the temperature parameter on the Responses API, so the judge omits it for those models; gpt-5.5 is the second judge in the default pair. * [grading] record judge failures per criterion and read Word comments - A criterion whose judge call fails gets an `error` verdict instead of aborting the task's grading; `RubricResult.n_grading_errors` and the scores file count them, and an error never counts as a pass. The dual evaluation keeps each judge's scores but writes no aggregate while any criterion is ungraded. - .docx text now ends with the document's Word margin comments, which pandoc drops in every track-changes mode. - A .pptx that is not a zip package is reported unreadable instead of being converted as plain text. - `read_file_as_text` is public. * [grading] add the score-impact changelog entry for #173 * [grading] list only the Word comments that pandoc's output leaves out pandoc prints comments in `--track-changes=all` mode but skips a comment whose range starts inside a tracked insertion, and prints none in `accept` mode. Appending every comment duplicated the printed ones for `include_docx_redlines` criteria. The margin-comments block now lists the referenced comments whose ids pandoc did not print, each with the text of the passage it is attached to. Comments the document body never references are left out, since the file format lets readers ignore them.
1 parent e08b58d commit 83fe609

10 files changed

Lines changed: 608 additions & 44 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 18 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -54,6 +54,24 @@ maps any entry to the merge commit that introduced it.
5454

5555
# Changes
5656

57+
## 2026-09-29 · PR #173 · [grading]
58+
A judge call that fails now gives that criterion an `error` verdict instead of
59+
aborting the task's grading; Claude judges retry transient API errors; judges
60+
using `gpt-5.5` no longer send the `temperature` parameter that model rejects;
61+
`.docx` text ends with the Word margin comments that pandoc's output leaves out,
62+
each with the passage it is attached to; and a `.pptx` that is not a zip package
63+
is reported unreadable instead of graded as raw bytes.
64+
Impact: grading, all providers, only in those cases. pandoc prints no comments
65+
for default criteria and skips comments anchored inside tracked insertions for
66+
`include_docx_redlines` criteria, so criteria on `.docx` deliverables with
67+
margin comments can now pass on the comment text. Comments the document body
68+
never references are not listed, since the file format lets Word ignore them.
69+
Runs that hit transient judge errors complete, with any ungraded criterion
70+
shown as `error` and never counted as a pass; `scores_dual.json` is still
71+
written only when every criterion is graded by both judges. Documents without
72+
margin comments or corruption extract byte-identically, so other results across
73+
this line are comparable.
74+
5775
## 2026-09-10 · PR #163 · [harness]
5876
Moved `lab_core/harness/`, `lab_core/evaluation/`, `lab_core/sandbox/`, `lab_core/utils/` under `lab_core/` and made the
5977
repo an installable package (`lab-core`). Commands are `python -m lab_core.<module>`;

‎lab_core/evaluation/judge.py‎

Lines changed: 53 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -8,6 +8,7 @@
88

99
import json
1010
import re
11+
import time
1112
from pathlib import Path
1213

1314
import anthropic
@@ -19,6 +20,22 @@
1920

2021
PROMPTS_DIR = Path(__file__).parent / "prompts"
2122

23+
# Output-token cap for Anthropic and OpenAI judges. The verdict itself is short, but with a
24+
# large deliverable the judge's reasoning can run past a lower cap and truncate.
25+
_JUDGE_MAX_OUTPUT_TOKENS = 64000
26+
27+
# Anthropic judge requests are retried with linear backoff on overload (529), rate limits
28+
# (429), brief unavailability (500, 503), and structured-output grammar compilation
29+
# timeouts, which the API reports as 400s.
30+
_JUDGE_API_MAX_ATTEMPTS = 6
31+
_JUDGE_API_BACKOFF_SECONDS = 20.0
32+
_JUDGE_RETRYABLE_STATUS = frozenset({429, 500, 503, 529})
33+
34+
# OpenAI judge models that reject the `temperature` parameter on the Responses API.
35+
_OPENAI_TEMPERATURE_UNSUPPORTED_MODELS = frozenset({"gpt-5.5", "gpt-5.5-pro", "gpt-5.4-pro"})
36+
# Dated snapshot suffix, e.g. "gpt-5.5-2026-07-01".
37+
_OPENAI_SNAPSHOT_SUFFIX_RE = re.compile(r"-\d{4}-\d{2}-\d{2}$")
38+
2239
_VERDICT_SCHEMA = {
2340
"type": "object",
2441
"properties": {
@@ -44,6 +61,12 @@ def _detect_provider(model: str) -> str:
4461
return "mistral"
4562
raise ValueError(f"Unknown judge provider for model: {model!r}")
4663

64+
65+
def _openai_temperature_unsupported(model: str) -> bool:
66+
"""Return whether an OpenAI model, or the base model of its dated snapshot, rejects `temperature`."""
67+
return _OPENAI_SNAPSHOT_SUFFIX_RE.sub("", model) in _OPENAI_TEMPERATURE_UNSUPPORTED_MODELS
68+
69+
4770
class Judge:
4871
"""LLM-as-judge that evaluates agent outputs against rubric criteria."""
4972

@@ -87,12 +110,35 @@ def evaluate(
87110
return self._evaluate_openai(prompt, temperature, _retries)
88111
return self._evaluate_mistral(prompt, temperature, _retries)
89112

113+
def _stream_with_transient_retry(self, kwargs: dict) -> anthropic.types.Message:
114+
"""Stream one Anthropic judge request and return its final message, retrying transient errors.
115+
116+
Statuses in `_JUDGE_RETRYABLE_STATUS`, overload errors, and grammar compilation
117+
timeouts are retried up to `_JUDGE_API_MAX_ATTEMPTS` times; any other error, or the
118+
error from the last attempt, is raised.
119+
"""
120+
for api_attempt in range(1, _JUDGE_API_MAX_ATTEMPTS + 1):
121+
try:
122+
# Streaming avoids the SDK's 10-minute limit on non-streaming requests at this output cap.
123+
with self.client.messages.stream(**kwargs) as stream:
124+
return stream.get_final_message()
125+
except anthropic.APIStatusError as e:
126+
retryable = (
127+
e.status_code in _JUDGE_RETRYABLE_STATUS
128+
or "overloaded" in str(e)
129+
or "Grammar compilation timed out" in str(e)
130+
)
131+
if not retryable or api_attempt == _JUDGE_API_MAX_ATTEMPTS:
132+
raise
133+
time.sleep(_JUDGE_API_BACKOFF_SECONDS * api_attempt)
134+
raise AssertionError("unreachable")
135+
90136
def _evaluate_anthropic(self, prompt: str, temperature: float, _retries: int) -> dict:
91137
last_err: Exception | None = None
92138
for attempt in range(_retries):
93139
kwargs = {
94140
"model": self.model,
95-
"max_tokens": 16384,
141+
"max_tokens": _JUDGE_MAX_OUTPUT_TOKENS,
96142
"temperature": temperature,
97143
"messages": [{"role": "user", "content": prompt}],
98144
}
@@ -105,7 +151,7 @@ def _evaluate_anthropic(self, prompt: str, temperature: float, _retries: int) ->
105151
}
106152
}
107153
try:
108-
response = self.client.messages.create(**kwargs)
154+
response = self._stream_with_transient_retry(kwargs)
109155
except anthropic.InternalServerError as e:
110156
# 500s on the structured-output path have been observed to
111157
# succeed when retried without output_config.
@@ -116,16 +162,12 @@ def _evaluate_anthropic(self, prompt: str, temperature: float, _retries: int) ->
116162
input_tokens = response.usage.input_tokens if response.usage else "unknown"
117163
raise ValueError(
118164
f"Judge response truncated (stop_reason=max_tokens, "
119-
f"input_tokens={input_tokens}, max_tokens={16384}). "
165+
f"input_tokens={input_tokens}, max_tokens={_JUDGE_MAX_OUTPUT_TOKENS}). "
120166
f"The agent output is likely too large for the judge context window. "
121167
f"Ensure criteria have deliverables lists to scope output."
122168
)
123169

124-
text = next(
125-
block.text
126-
for block in response.content
127-
if block.type == "text"
128-
)
170+
text = next((block.text for block in response.content if block.type == "text"), "")
129171
try:
130172
return self._parse_json(text)
131173
except (ValueError, json.JSONDecodeError) as e:
@@ -169,9 +211,10 @@ def _evaluate_openai(self, prompt: str, temperature: float, _retries: int) -> di
169211
kwargs = {
170212
"model": self.model,
171213
"input": prompt,
172-
"max_output_tokens": 16384,
173-
"temperature": temperature,
214+
"max_output_tokens": _JUDGE_MAX_OUTPUT_TOKENS,
174215
}
216+
if not _openai_temperature_unsupported(self.model):
217+
kwargs["temperature"] = temperature
175218
if attempt < _retries - 1:
176219
kwargs["text"] = {
177220
"format": {

‎lab_core/evaluation/report.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -92,7 +92,7 @@ def generate_report(run_id: str) -> Path:
9292
for c in criteria:
9393
verdict = c["verdict"]
9494
badge_cls = "badge-found" if verdict == "pass" else "badge-missed"
95-
badge_text = "PASS" if verdict == "pass" else "FAIL"
95+
badge_text = {"pass": "PASS", "error": "ERROR"}.get(verdict, "FAIL")
9696
reasoning = c.get("reasoning", "")
9797

9898
criteria_html.append(f"""

‎lab_core/evaluation/run_eval.py‎

Lines changed: 12 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -114,6 +114,7 @@ def evaluate_run(run_id: str, task: str, judge: Judge, parallel: int = 6) -> dic
114114
summary = (
115115
f"{n_passed}/{n_criteria} criteria passed."
116116
+ (" ALL-PASS." if all_pass else f" Missed {n_criteria - n_passed} — task FAIL.")
117+
+ (f" {result.n_grading_errors} could not be graded." if result.n_grading_errors else "")
117118
)
118119

119120
scores = {
@@ -123,6 +124,7 @@ def evaluate_run(run_id: str, task: str, judge: Judge, parallel: int = 6) -> dic
123124
"all_pass": all_pass,
124125
"n_criteria": n_criteria,
125126
"n_passed": n_passed,
127+
"n_grading_errors": result.n_grading_errors,
126128
"criteria_results": result.criteria_results,
127129
"run_id": run_id,
128130
"task": task,
@@ -164,7 +166,8 @@ def evaluate_run_dual(
164166
165167
Each judge grades every criterion independently. Per-judge results are
166168
preserved alongside the aggregate so single-judge artifacts are not
167-
overwritten.
169+
overwritten. If any criterion could not be graded by either judge, the
170+
per-judge results are kept, no aggregate is written, and RuntimeError is raised.
168171
"""
169172
if len(judge_models) != 2:
170173
raise ValueError("Dual evaluation requires exactly two judge models")
@@ -192,6 +195,14 @@ def evaluate_run_dual(
192195
if scores_path.exists():
193196
scores_path.rename(run_dir / f"scores_{judge_model}.json")
194197

198+
ungraded = {model: scores["n_grading_errors"] for model, scores in per_judge.items() if scores["n_grading_errors"]}
199+
if ungraded:
200+
detail = ", ".join(f"{model}: {count}" for model, count in ungraded.items())
201+
raise RuntimeError(
202+
f"Criteria could not be graded ({detail}); {out_path.name} was not written. "
203+
"Re-run the evaluation to grade them."
204+
)
205+
195206
def crit_frac(scores: dict) -> float:
196207
return (
197208
scores["n_passed"] / scores["n_criteria"]

0 commit comments

Comments
 (0)