forked from UKGovernmentBEIS/inspect_ai
-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathsetting-limits.qmd
More file actions
448 lines (327 loc) · 13.7 KB
/
Copy pathsetting-limits.qmd
File metadata and controls
448 lines (327 loc) · 13.7 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
---
title: "Setting Limits"
llms-description: Setting time, message, token, and cost limits on evaluation tasks, samples, and agent execution.
---
## Overview
In open-ended model conversations (for example, an agent evaluation with tool usage) it's possible that a model will get "stuck" attempting to perform a task with no realistic prospect of completing it. Further, sometimes models will call commands in a sandbox that take an extremely long time (or worst case, hang indefinitely).
For this type of evaluation it's normally a good idea to set limits on some combination of total time, total messages, turns, tokens used, and/or cost. This article covers:
1. [Sample Limits](#sample-limits) --- limits applied to individual samples within a task.
2. [Scoped Limits](#scoped-limits) --- limits applied to arbitrary blocks of code.
3. [Agent Limits](#agent-limits) --- limits applied to agent execution.
## Sample Limits {#sample-limits}
Sample limits don't result in errors, but rather an early exit from execution (samples that encounter limits are still scored, albeit nearly always as "incorrect").
### Time Limit
Here we set a `time_limit` of 15 minutes (15 x 60 seconds) for each sample within a task:
``` python
@task
def intercode_ctf():
return Task(
dataset=read_dataset(),
solver=[
system_message("system.txt"),
use_tools([bash(timeout=3 * 60)]),
generate(),
],
time_limit=15 * 60,
scorer=includes(),
sandbox="docker",
)
```
Note that we also set a timeout of 3 minutes for the `bash()` command. This isn't required but is often a good idea so that a single wayward bash command doesn't consume the entire `time_limit`.
We can also specify a time limit at the CLI or when calling `eval()`:
``` bash
inspect eval ctf.py --time-limit 900
```
Appropriate timeouts will vary depending on the nature of your task so please view the above as examples only rather than recommend values.
### Working Limit
{{< include _working_limits.md >}}
Here we set an `working_limit` of 10 minutes (10 x 60 seconds) for each sample within a task:
``` python
@task
def intercode_ctf():
return Task(
dataset=read_dataset(),
solver=[
system_message("system.txt"),
use_tools([bash(timeout=3 * 60)]),
generate(),
],
working_limit=10 * 60,
scorer=includes(),
sandbox="docker",
)
```
### Message Limit
{{< include _message_limits.md >}}
Here we set a `message_limit` of 30 for each sample within a task:
``` python
@task
def intercode_ctf():
return Task(
dataset=read_dataset(),
solver=[
system_message("system.txt"),
use_tools([bash(timeout=120)]),
generate(),
],
message_limit=30,
scorer=includes(),
sandbox="docker",
)
```
This sets a limit of 30 total messages in a conversation before the model is forced to give up. At that point, whatever `output` happens to be in the `TaskState` will be scored (presumably leading to a score of incorrect).
### Token Limit
{{< include _token_limits.md >}}
Here we set a `token_limit` of 500K for each sample within a task:
``` python
@task
def intercode_ctf():
return Task(
dataset=read_dataset(),
solver=[
system_message("system.txt"),
use_tools([bash(timeout=120)]),
generate(),
],
token_limit=(1024*500),
scorer=includes(),
sandbox="docker",
)
```
::: callout-important
It's important to note that the `token_limit` is for all tokens used within the execution of a sample. If you want to limit the number of tokens that can be yielded from a single call to the model you should use the `max_tokens` generation option.
:::
### Turn Limit
{{< include _turn_limits.md >}}
Here we set a `turn_limit` of 300 for each sample within a task:
``` python
@task
def intercode_ctf():
return Task(
dataset=read_dataset(),
solver=[
system_message("system.txt"),
use_tools([bash(timeout=120)]),
generate(),
],
turn_limit=300,
scorer=includes(),
sandbox="docker",
)
```
This limits the agent to 300 model generations before it is forced to give up. As with other sample limits, whatever `output` happens to be in the `TaskState` at that point will be scored.
### Cost Limit
{{< include _cost_limits.md >}}
Here we set a `cost_limit` of $2.00 for each sample within a task:
``` python
@task
def intercode_ctf():
return Task(
dataset=read_dataset(),
solver=[
system_message("system.txt"),
use_tools([bash(timeout=120)]),
generate(),
],
cost_limit=2.00,
scorer=includes(),
sandbox="docker",
)
```
::: callout-important
The `cost_limit` requires model cost data to be configured via `set_model_cost()` or `--model-cost-config`. An error will be raised if a cost limit is set without cost data for all models used in the evaluation.
:::
#### Model Cost {#model-cost}
Cost tracking requires cost data for each model present in the eval or eval set. There are two ways to set cost data:
**Python API:**
``` python
from inspect_ai.model import set_model_cost, ModelCost
set_model_cost("openai/gpt-4o", ModelCost(
input=2.50, output=10.00,
input_cache_write=0, input_cache_read=1.25,
))
```
**CLI (YAML or JSON file):**
Each model needs a price set for `input`, `output`, `input_cache_write`, and `input_cache_read`. Prices should be given in dollars per million tokens. Set unused fields to `0`.
Below is an example cost config file given in YAML:
``` yaml
openai/gpt-4o:
input: 2.50
output: 10.00
input_cache_write: 0
input_cache_read: 1.25
anthropic/claude-sonnet-4-5-20250514:
input: 3.00
output: 15.00
input_cache_write: 3.75
input_cache_read: 0.30
```
(As of Feb 9 2026, all major model providers count reasoning tokens as output tokens, so no separate price needs to be provided for reasoning tokens. If your use case requires separate calculation of reasoning token prices, contact us.)
When model cost data is configured, costs will be tracked for the sample as a whole, as well as any events within the sample that have a ModelUsage field.
Additionally, configuring model cost data allows setting sample cost limits:
``` bash
inspect eval ctf.py --model-cost-config pricing.yaml --cost-limit 2.00
```
### Custom Limit
When limits are exceeded, a `LimitExceededError` is raised and caught by the main Inspect sample execution logic. If you want to create custom limit types, you can enforce them by raising a `LimitExceededError` as follows:
``` python
from inspect_ai.util import LimitExceededError
raise LimitExceededError(
"custom",
value=value,
limit=limit,
message=f"A custom limit was exceeded: {value}"
)
```
### Query Usage
We can determine how much of a sample limit has been used, what the limit is, and how much of the resource is remaining:
``` python
sample_time_limit = sample_limits().time
print(f"{sample_time_limit.remaining:.0f} seconds remaining")
```
Note that `sample_limits()` only retrieves the sample-level limits, not [scoped limits](#scoped-limits) or [agent limits](#agent-limits).
## Scoped Limits {#scoped-limits}
You can also apply limits at arbitrary scopes, independent of the sample or agent-scoped limits. For instance, applied to a specific block of code. For example:
``` python
with token_limit(1024*500):
...
```
A `LimitExceededError` will be raised if the limit is exceeded. The `source` field on `LimitExceededError` will be set to the `Limit` instance that was exceeded.
When catching `LimitExceededError`, ensure that your `try` block encompasses the usage of the limit context manager as some `LimitExceededError` exceptions are raised at the scope of closing the context manager:
``` python
try:
with token_limit(1024*500):
...
except LimitExceededError:
...
```
The `apply_limits()` function accepts a list of `Limit` instances. If any of the limits passed in are exceeded, the `limit_error` property on the `LimitScope` yielded when opening the context manager will be set to the exception. By default, all `LimitExceededError` exceptions are propagated. However, if `catch_errors` is true, errors which are as a direct result of exceeding one of the limits passed to it will be caught. It will always allow `LimitExceededError` exceptions triggered by other limits (e.g. Sample scoped limits) to propagate up the call stack.
``` python
with apply_limits(
[token_limit(1000), message_limit(10)], catch_errors=True
) as limit_scope:
...
if limit_scope.limit_error:
print(f"One of our limits was hit: {limit_scope.limit_error}")
```
### Checking Usage
You can query how much of a limited resource has been used so far via the `usage` property of a scoped limit. For example:
``` python
with token_limit(10_000) as limit:
await generate()
print(f"Used {limit.usage:,} of 10,000 tokens")
```
If you're passing the limit instance to `apply_limits()` or an agent and want to query the usage, you should keep a reference to it:
``` python
limit = token_limit(10_000)
with apply_limits([limit]):
await generate()
print(f"Used {limit.usage:,} of 10,000 tokens")
```
### Time Limit
To limit the wall clock time to 15 minutes within a block of code:
``` python
with time_limit(15 * 60):
...
```
Internally, this uses [`anyio`'s cancellation scopes](https://anyio.readthedocs.io/en/stable/cancellation.html). The block will be cancelled at the first yield point (e.g. `await` statement).
### Working Limit
{{< include _working_limits.md >}}
To limit the working time to 10 minutes:
``` python
with working_limit(10 * 60):
...
```
Unlike time limits, this is not driven by `anyio`. It is checked periodically such as from `generate()` and after each `Solver` runs.
### Message Limit
{{< include _message_limits.md >}}
Scoped message limits behave differently to scoped token limits in that only the innermost active `message_limit()` is checked.
To limit the conversation length within a block of code:
``` python
@agent
def myagent() -> Agent:
async def execute(state: AgentState):
with message_limit(50):
# A LimitExceededError will be raised when the limit is exceeded
...
with message_limit(None):
# The limit of 50 is temporarily removed in this block of code
...
```
::: callout-important
It's important to note that `message_limit()` limits the total number of messages in the conversation, not just "new" messages appended by an agent.
:::
### Token Limit
{{< include _token_limits.md >}}
To limit the total number of tokens which can be used in a block of code:
``` python
@agent
def myagent(tokens: int = (1024*500)) -> Agent:
async def execute(state: AgentState):
with token_limit(tokens):
# a LimitExceededError will be raised if the limit is exceeded
...
```
The limits can be stacked. Tokens used while a context manager is open count towards all open token limits.
``` python
@agent
def myagent() -> Solver:
async def execute(state: AgentState):
with token_limit(1024*500):
...
with token_limit(1024*200):
# Tokens used here count towards both active limits
...
```
::: callout-important
It's important to note that `token_limit()` is for all tokens used *while the context manager is open*. If you want to limit the number of tokens that can be yielded from a single call to the model you should use the `max_tokens` generation option.
:::
#### Suspending Token Limits
To run a block of code that should not count against any active token limits, use `suspend_token_limit()`:
``` python
with token_limit(10_000):
await generate() # counts against the 10k budget
with suspend_token_limit():
# tokens used here are not metered against the 10k limit,
# and any inner `token_limit()` is also suspended
await expensive_summary()
await generate() # counts again
```
Unlike `with token_limit(None):`, which only suppresses the innermost limit's check, `suspend_token_limit()` fully disables both recording and checking across all active token limits for the duration of the block.
### Turn Limit
{{< include _turn_limits.md >}}
To limit the total number of turns (model generations) which can be used in a block of code:
``` python
@agent
def myagent(turns: int = 300) -> Agent:
async def execute(state: AgentState):
with turn_limit(turns):
# a LimitExceededError will be raised if the limit is exceeded
...
```
Like token limits, turn limits can be stacked — a generation counts towards all open turn limits. To run generations that should not count against any active turn limit, use `suspend_turn_limit()`:
``` python
with turn_limit(10):
await generate() # counts against the 10 turn budget
with suspend_turn_limit():
# generations here are not metered against the 10 turn limit
await auxiliary_generate()
await generate() # counts again
```
### Cost Limit
{{< include _cost_limits.md >}}
To limit the total cost within a block of code:
``` python
@agent
def myagent(budget: float = 2.00) -> Agent:
async def execute(state: AgentState):
with cost_limit(budget):
# a LimitExceededError will be raised if the limit is exceeded
...
```
Cost limits work similarly to token limits, with stacking and tracking of costs used while the context manager is open.
::: callout-important
Using `cost_limit()` requires model cost data to be configured via `set_model_cost()` or `--model-cost-config`. See [Model Cost](#model-cost) for details.
:::
## Agent Limits {#agent-limits}
{{< include _agent_limits.md >}}