Papers
Topics
Authors
Recent
Search
2000 character limit reached

Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses

Published 25 Aug 2026 in cs.AI and cs.CL | (2608.24876v1)

Abstract: Recursive self-improvement (RSI) remains hard in long-horizon tasks, where growing histories obscure the task state and misalign skill invocation. We introduce Recuris, a recursive Experiential-Working Memory architecture for long-horizon agent harnesses, in which Working Memory tracks task progress and guides skill selection from Experiential Memory, grounding skill use in current needs rather than the full history. This coupling also turns execution into structured evidence that localizes failures to specific memory components. Across tasks, a fixed Meta-Agent turns that evidence into localized, validation-gated updates to Skill Memory that reshape execution and yield new evidence, forming a bounded recursive memory-evolution loop. Across four long-horizon benchmarks and ten models, Recuris improves task success in 35 of the 37 completed model-benchmark pairs, carrying frontier models to SOTA-level task success: on tau-bench it adds +17.8 points to GPT-5.6 Sol and +15.6 to Claude Opus 5, taking Opus 5 to 87.9%, and +16.6/+13.5 points on Qwen3.6-27B/35B on SkillFlow. The advantage widens as the interaction horizon grows, to +32.2 points on the longest tasks, and common long-horizon failures fall by up to 80%. These results position recursively evolving memory as a scalable foundation for RSI, enabling agents to continuously transform accumulated experience into increasingly effective long-horizon behavior. Code: https://github.com/Gen-Verse/Recuris

Summary

  • The paper introduces Recuris, a recursive self-improvement framework that combines verified Working Memory with state-triggered Experiential Memory while keeping the underlying model and agent program fixed.
  • The system improves success in 35 of 37 completed model–benchmark comparisons, including gains of 17.8 points for GPT-5.6 Sol and 15.6 points for Claude Opus 5 on τ²-Retail.
  • Structured execution traces raise memory-failure attribution accuracy from 13.0% using outcomes alone to 64.8%, enabling localized, validation-gated patches that reduce omitted actions in long-horizon tasks.

Problem formulation and central thesis

“Recursive Experiential–Working Memory Evolution for Long-Horizon Agent Harnesses” (2608.24876) addresses recursive self-improvement (RSI) in LLM agents without modifying model weights or the full agent program. The paper’s premise is that long-horizon failures are often caused not by a lack of general reasoning capability, but by failures in the external memory-control layer: the agent loses track of unresolved goals, invokes relevant experience at the wrong time, or accepts unsupported claims of task completion.

The proposed system, Recuris, separates two functions that are commonly conflated. Experiential Memory (EM) stores reusable skills, procedures, and domain knowledge. Working Memory (WM) maintains a compact, task-specific state containing goals, completion status, supporting evidence, and blockers. WM determines what remains unresolved; EM supplies reusable experience relevant to that current state. Environment feedback is then passed through explicit checkers before it can update WM.

The paper advances two related claims. First, state-grounded interaction with memory is more reliable than retrieval from the initial instruction or the full dialogue history. Second, the same stateful execution process can generate structured evidence for recursively improving the memory-control layer. Recuris therefore treats execution as both inference and diagnosis:

task state→skill invocation→environment feedback→verified state update.\text{task state} \rightarrow \text{skill invocation} \rightarrow \text{environment feedback} \rightarrow \text{verified state update}.

Across tasks, a fixed Meta-Agent uses the resulting traces to identify a likely defective memory component, proposes a localized patch, and submits it to a validation gate. The underlying LLM, tools, Meta-Agent procedure, and outer harness remain fixed.

Figure 1

Figure 1: Recuris replaces history-based retrieval with execution-event retrieval and replaces whole-memory rewriting with component-specific, validation-gated patches.

Architecture and execution semantics

Recuris represents its evolving memory-control layer as four components:

Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),

where Ek\mathcal{E}_k is experiential memory, Wk\mathcal{W}_k specifies the working-state schema and update proposal, ρk\rho_k controls invocation timing and retrieval keys, and Ck\mathcal{C}_k contains completion checkers. A candidate update may modify one or several implicated components, but each edit must remain within the component to which the diagnostic procedure attributes the failure.

Within a task, WM records each goal as pending, done, or blocked, together with evidence and optional blockers. The invocation policy ρk\rho_k retrieves skills at defined execution events rather than performing one retrieval at task initialization. In the principal τ2\tau^2-Bench configuration, invocation occurs when the model drafts a state-changing tool call, and the retrieval key is the tool named by that call. The drafted action is withheld until the corresponding skill has been inserted into context. Terminal-Bench uses a boundary-based policy that injects skills at the first turn.

The state update is deliberately asymmetric. WM may propose a transition, but a fixed kernel commits only changes supported by the environment observation. A model assertion that a task is complete is not itself completion evidence. A goal becomes done only when its checker accepts the returned tool result or structured receipt. This distinction is important because a tool call can be syntactically valid while leaving the environment unchanged.

Figure 2

Figure 2: Recuris couples verified working-state updates with event-triggered experiential-skill retrieval within tasks and uses structured traces for cross-task memory evolution.

The cross-task loop records, for every execution step, the working state, retrieved skills, action, observation, proposed state update, checker decisions, and committed state. The Meta-Agent then attributes diagnosed failures to E\mathcal{E}, W\mathcal{W}, Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),0, or Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),1. It can add or revise a skill, alter the state schema, change the invocation trigger or retrieval key, or modify a checker. Candidate patches are admitted only if they repair the source failure and do not violate the preset regression criterion on held-out development tasks.

This is a bounded form of RSI. Recuris does not rewrite its model, Meta-Agent, evaluator, or complete runtime. Recursion occurs because each admitted memory patch changes subsequent execution, and those changed executions generate new traces that can motivate later patches.

Experimental design

The evaluation covers four long-horizon settings: Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),2-Retail, Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),3-Airline, SkillFlow, and Terminal-Bench 2.1. The first three support cross-task memory evolution to different degrees; Terminal-Bench is used primarily for isolated test-time adaptation. The experiments include ten models, ranging from 3B open-weight systems to frontier models. Every model is frozen, evaluated at temperature zero, and run under matched task, tool, seed, and interaction budgets.

For Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),4-Bench, success requires full environment-verifier reward. The analysis separates read-action recall from required-write recall, allowing the paper to distinguish failure to identify necessary information from failure to execute state-changing actions. Statistical comparisons use paired task-clustered bootstrap confidence intervals, reflecting the fact that multiple episodes associated with the same task are not independent.

The memory is evolved on the deployment model, Doubao Seed 2.0 Pro, and then transferred without modification to other models. This design is consequential: the reported cross-model gains cannot be attributed to target-model-specific memory fitting, although it also means that the memory is shaped by the failure distribution of one source model.

Overall performance

Recuris improves task success in 35 of 37 completed model–benchmark comparisons. The strongest results occur on Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),5-Retail and SkillFlow, where the deployment model improves by 23.3 and 16.8 percentage points, reaching 81.4% and 51.4%, respectively.

Model Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),6-Retail alone With Recuris Gain
GPT-5.6 Sol 58.3% 76.1% +17.8
Claude Opus 5 72.4% 87.9% +15.6
Qwen3.6-27B, SkillFlow 42.2% 58.7% +16.6
Qwen3.6-35B, SkillFlow 35.3% 48.8% +13.5
Doubao-2.0-Pro, Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),7-Retail 58.1% 81.4% +23.3

The result on Claude Opus 5 is particularly notable because the final 87.9% success rate exceeds every non-Recuris model evaluated on Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),8-Retail by 9.7 points. The paper therefore rejects a simple saturation interpretation of frontier-model performance on these long-horizon tasks: even a strong frozen model remains sensitive to execution-state management.

The gain is not explained by a neutral memory wrapper or by additional prompt context. The initial memory-control layer contributes no statistically distinguishable improvement on four target models. In a controlled Mk=(Ek,Wk,ρk,Ck),\mathcal{M}_k=(\mathcal{E}_k,\mathcal{W}_k,\rho_k,\mathcal{C}_k),9-Retail comparison, model-controlled invocation places 3,111 more tokens in the initial context than Recuris, yet performs 18 points worse and requires 46% more tokens per successful task. Thus, the relevant variable is not memory quantity but the policy governing when memory is made available.

Long-horizon execution is the dominant failure mode

The paper’s most diagnostic analysis concerns task length. On Ek\mathcal{E}_k0-Retail, Recuris outperforms the base agent in every intrinsic-horizon quartile, with gains ranging from 17.0 to 44.7 points. The advantage increases rather than decays on the longest tasks.

Figure 3

Figure 3: Recuris maintains its advantage across intrinsic task-length quartiles, with the largest separation appearing on the longest episodes.

The analysis contradicts the interpretation that long-horizon degradation primarily reflects retrieval failure. Read-action recall remains between 88.0% and 97.9% for every variant and horizon quartile. Required-write recall, by contrast, separates the systems substantially: Recuris exceeds the base agent by 26.7 points. The base agent ends 42% of episodes requiring a write without executing any required write, compared with 16% for Recuris.

The median turn of the first correct write is identical across the variants. Recuris therefore does not make the agent decide to write earlier; it makes the agent more likely to issue the write at all. The implication is that the principal long-horizon bottleneck is execution coverage under evolving state, not failure to understand the task or retrieve relevant information.

Ablations isolate the contribution of working memory and invocation control

The EM–WM ablation provides the clearest evidence for the paper’s architectural thesis. On Ek\mathcal{E}_k1-Retail, EM alone raises success by only 2.0 points, with a confidence interval including zero. WM alone raises success by 23.9 points, while the coupled system raises it by 25.4 points. On Ek\mathcal{E}_k2-Airline, the corresponding gains are smaller and statistically unresolved, reflecting the smaller evaluation set and different domain structure.

Figure 4

Figure 4: Working memory supplies most of the measurable gain on Ek\mathcal{E}_k3-Retail, while experiential memory contributes primarily when invocation is state-grounded.

The same skill library performs substantially worse when the model itself controls invocation. The model-controlled variant reaches 65.6% success and 61.1% required-write recall, compared with 83.6% and 82.4% for Recuris. The model-controlled configuration also consumes 147k agent tokens per success, versus 101k for Recuris. It performs worse than the WM-only configuration despite receiving more skill content.

This is a strong and somewhat counterintuitive claim: additional skill availability can reduce performance when invocation is delegated to the model without a reliable state-based trigger. Recuris does not improve the median timing of the first correct write, which remains turn 18, but it reduces the probability that required writes are omitted. Once a required write mismatches, the probability of another later mismatch exceeds 60% in every configuration; Recuris reaches this error-compounding regime least often.

Figure 5

Figure 5: State-grounded invocation primarily improves action coverage and reduces entry into the within-episode error-compounding regime.

The authors also identify a domain-specific double dissociation. Removing write review reduces Ek\mathcal{E}_k4-Airline success by 13.5 points but has no measurable effect on Ek\mathcal{E}_k5-Retail. Removing the status board reduces Ek\mathcal{E}_k6-Retail success by 17.3 points but has no measurable effect on Ek\mathcal{E}_k7-Airline. The truth guard, which checks unsupported completion claims after execution, has no measurable effect despite rejecting 172 such claims. The timing matters: post hoc detection cannot undo an already executed incorrect write.

Figure 6

Figure 6: Different domains depend on different working-memory mechanisms: pre-execution write review matters for Ek\mathcal{E}_k8-Airline, whereas persistent status presentation matters for Ek\mathcal{E}_k9-Retail.

The result limits any architecture-level prescription that treats one memory mechanism as universally decisive. Recuris’s contribution is instead to expose the mechanism through structured traces and allow the evolution loop to target it.

Structured traces make localized evolution possible

A central methodological claim is that final success or failure is insufficient for reliable memory evolution. To test this, the authors inject known faults into memory components and ask a fixed judge to identify the defective component under three evidence conditions: outcome only, raw trajectory, or structured trace.

Evidence condition Macro attribution accuracy
Outcome only 13.0%
Raw trajectory 37.0%
Structured trace 64.8%

The outcome-only condition performs below the 33.3% constant-answer baseline. Raw trajectories improve attribution but remain weak, particularly for invocation faults, which are defined by an absent event. Structured traces raise macro accuracy to 64.8% and macro-F1 to 63.4%. Working-memory faults are identified with 83.3% recall, while invocation faults rise from 0% under raw trajectories to 38.9% with structured traces.

Figure 7

Figure 7: A verified execution trace preserves unresolved goals despite verbal confirmation and thereby exposes the missing tool actions that caused failure.

The improvement is primarily an observability result. The structured trace explicitly records non-events, state transitions, invocation decisions, and checker outcomes that cannot be reconstructed reliably from dialogue. This supports a narrower repair surface: rather than rewriting all memory, the system patches only the component associated with the diagnosed failure.

The evolution experiments show held-out gains across multiple runs and Meta-Agent implementations. On a fixed 86-task held-out split, evolved packages improve over the neutral starting memory by 9.01 to 17.44 points. In one lineage, a second evolution round adds a further 6.98 points. The strongest reported package reaches a 17.44-point improvement over the initial memory.

Figure 8

Figure 8: Held-out performance improves across several evolution runs, although some rounds plateau, reverse, or produce packages that are never invoked.

The results also expose an important operational failure mode: a candidate can be semantically useful but behaviorally inert if its invocation binding is broken. One round-4 package is never invoked on any held-out task and consequently exhibits no improvement. This distinguishes failure of memory content from failure of memory reachability.

Two independently implemented Meta-Agents converge numerically and mechanistically. Claude Code and DeepSeek Harness produce settled gains of 11.92 and 10.47 points, with a direct paired difference of -1.45 points and Wk\mathcal{W}_k0. Their champion packages also converge on similar repair families involving authorization state, execution gating, and anti-escalation skills. This supports the interpretation that the structured evidence, rather than a particular Meta-Agent implementation, determines much of the learned content.

Figure 9

Figure 9: A structured trace enables a localized experiential-memory patch for an exchange procedure, which transfers to an unseen matched task.

Transfer across tasks and models

The evolved memory transfers across held-out tasks when those tasks contain failures of the type the memory repairs. On Wk\mathcal{W}_k1-Retail, packages evolved from failures on 16 tasks improve performance on 86 untouched tasks by 9.01–17.44 points. Transfer is not universal: Wk\mathcal{W}_k2-Airline lineages show no statistically resolved held-out gain when the held-out tasks already have high baseline success and contain few remaining instances of the target failure.

The transfer pattern is similarly conditional across models. A memory evolved on a mid-sized deployment model improves GPT-5.6 Sol by 17.8 points and Claude Opus 5 by 15.6 points on Wk\mathcal{W}_k3-Retail. Gemini 3.7 Flash gains 4.8 points, with an interval that includes zero. On SkillFlow, the gains are larger for the two strongest open-weight models, but on Wk\mathcal{W}_k4-Retail gain does not monotonically follow model scale. The memory appears to carry procedural discipline—what to verify and when to retain a goal as unresolved—rather than merely compensating for weak model capability.

This distinction matters for interpreting transfer. The memory is not a universally useful additive capability. Its value depends on whether the receiving model exhibits the execution failure that the memory encodes and whether the task distribution preserves the relevant procedural structure.

Test-time adaptation on isolated tasks

Terminal-Bench 2.1 provides a different regime because its tasks do not share sufficient structure for cross-task memory evolution. Recuris therefore adapts memory within a single task after a failed attempt. The Meta-Agent receives the task instruction, the failed trajectory, and a single failure bit, but not the hidden verifier, tests, or expected output.

The headline result is 53 of 87 tasks solved within four attempts, or 60.9%. However, the paper correctly decomposes this result. Retrying with frozen seed memory already raises performance from 34.5% at one attempt to 58.6% at four attempts, a 26.4-point gain with Wk\mathcal{W}_k5. Test-time adaptation adds only 2.3 points over matched-budget retrying, with Wk\mathcal{W}_k6. The seed memory alone changes performance by -2.3 points relative to the bare baseline, also within noise.

The paper therefore rejects the interpretation that the 60.9% result demonstrates a large learning effect. The attempt budget explains the headline. Adaptation shows a positive but unresolved directional effect on matched-budget metrics: on the 56 tasks whose adapted memory contains a learned skill, average per-attempt success rises by 4.5 points, with a confidence interval spanning zero. Seven tasks are solved only under adaptation, but the sample does not establish a reliable aggregate effect.

This analysis is methodologically important because it identifies retrying as a major confound in test-time agent adaptation. Recuris’s strongest evidence lies in cross-task evolution on structured domains, not in the unqualified Terminal-Bench solved-within-budget number.

Limitations and open questions

The paper’s strongest results depend on benchmarks with shared procedural structure. Cross-task evolution fails to admit a patch on Terminal-Bench in thirteen evolution runs, indicating that Recuris does not create transferable knowledge when tasks lack common tools, policies, or execution flows. SkillFlow has no held-out task split for within-family template selection, so its task-level procedural improvements are in-sample; the reported generalization there is primarily across target models.

The causal interpretation of failure localization is also limited. The attribution procedure selects the component most likely to provide an effective repair; it does not establish that the component was the sole or true cause of failure. Multiple components may contribute to one trajectory. Moreover, the checkers are only as reliable as their environment-specific predicates. A false rejection can preserve an unresolved goal indefinitely, while an overly permissive checker can commit an unsupported transition.

The validation gate is conservative but statistically weak in some settings. On a 12-task development split, several candidates that later improved the 86-task held-out set were rejected because their in-round confidence intervals included zero. The gate thus behaves as a commitment-control mechanism rather than a dependable classifier of good and bad patches. The memory also grows monotonically: eight accepted patches add 51 skills, leave 17 near-duplicate pairs, and do not include deletion or pruning. Although ablations suggest redundancy, long-run scaling and interference remain open.

Finally, the results are based on one memory evolved from one deployment model per benchmark. The transfer experiments demonstrate portability, but they do not establish that a jointly evolved or model-specific memory would behave similarly. The main unresolved technical question is how to maintain component attribution, invocation reachability, and regression control as the memory grows and task distributions become less structurally homogeneous.

Conclusion

Recuris frames RSI as recursive evolution of an externalized memory-control layer rather than modification of model weights or the complete agent runtime. Its central mechanism is the coupling of verified Working Memory with state-grounded Experiential Memory: WM preserves unresolved goals and verified progress, while EM supplies skills at execution events selected by the current state.

The empirical evidence supports this decomposition. Recuris improves 35 of 37 completed model–benchmark pairs, raises GPT-5.6 Sol and Claude Opus 5 by 17.8 and 15.6 points on Wk\mathcal{W}_k7-Retail, increases the advantage on the longest tasks to as much as 44.7 points, and reduces required-write omissions that dominate long-horizon failure. Structured traces improve component attribution from 13.0% using outcomes alone to 64.8%, enabling localized, validation-gated memory updates. The results establish a technically specific form of recursive improvement: execution state controls experience use, execution traces diagnose memory failures, and accepted component-level patches alter subsequent execution without changing the underlying model.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 8 tweets with 14 likes about this paper.