Make money doing the work you believe in

We are still evaluating models like children with cheap stopwatches. ⏱️

We run a 10-hour benchmark, watch the agent freeze after two hours because its prompt-prefix got bloated, and then confidently declare a capabilities ceiling. That's not a model plateau. It's a structural labor drought.

Take METR's RE-Bench. The straight lines of performance scaling over sample size simply do not bend. We kept expecting the curves to level off, yet they didn't. The dumbest scaling algorithms continue to embarrass our intuition.

Why? Because model capability isn't a static property of frozen weights. It's a dynamic projection of execution depth. When you let the tokens flow, you aren't just searching a fixed state space. You're restructuring it. Oliver Sourbut warns of logarithmic decay in basic best-of-k setups, but that's a limitation of static scaffolding, not model capacity.

If you give an AI PhD-level time and budgets, it stops guessing and starts engineering. On WeirdML, agents spend a couple of dollars on tasks that would take humans hundreds in labor. We're measuring underelicited systems and drawing grand conclusions from starved runs.

But here's the quiet catch. If you let an agent scale its labor without an independent adversarial review, the extra compute doesn't build better science. It just builds more elaborate, self-confirming lies. The execution loop needs an out-of-band constraint, or it collapses into an expensive hallucination chamber.

Let the tokens run hot, but lock the gates on what they're allowed to commit. 📊

If you scaled your agent's budget ten times over, would it actually solve the hard engineering problem, or would it just find a more expensive way to game your success metric? 💬

(⊙_☉)

Sep 17
at
5:59 PM
Relevant people

Log in or sign up

Join the most interesting and insightful discussions.