Multi-head recurrent memory architecture: the agent reads the context and memory, selects one head for update, and leaves other heads unchanged. Head utilization is distributed across time.
A small architectural change. MHM partitions recurrent memory into independent heads. Each step updates exactly one head, structurally protecting the others from overwriting. Original overview from the paper.

Long-context reasoning · Recurrent memory

Multi-Head Recurrent
Memory Agents

Department of Computer Science, University of Wisconsin–Madison

Select one head. Preserve the rest.
A training-free memory architecture for more reliable reasoning over long contexts.

73.96%
Memory retention

RULER-HQA · 896K tokens
vs. <30% for both baselines

+28.12
Accuracy percentage points

49.74% vs. 21.62% MemAgent
RULER-HQA · 896K tokens

4,096
Total memory tokens

Same capacity across methods
4 heads × 1,024 tokens for MHM-LRU

The big picture

Abstract

Recurrent memory agents extend LLMs to arbitrarily long contexts by iteratively consolidating input into a fixed-size memory window. Despite their scalability, these agents exhibit a well-documented reliability problem: end-to-end performance degrades systematically as context length grows. We diagnose this failure by decomposing performance into two factors—memory capture and memory retention—and quantitatively confirm that retention is the dominant bottleneck. Retention collapses because existing designs maintain memory as a monolithic text block, forcing every update to risk overwriting previously retained content. Motivated by this diagnosis, we propose Multi-Head Recurrent Memory (MHM), a general, training-free framework that partitions memory into independent heads governed by a stage-wise select-then-update strategy. At each step, exactly one head is selected for update while the remaining heads are structurally shielded from overwriting, shifting the burden of retention from model behavior to architectural design. As a lightweight instantiation, we introduce Least-Recently-Updated MHM (MHM-LRU), which guarantees uniform head utilization with zero additional token overhead. Extensive experiments on long-context benchmarks show that MHM-LRU substantially improves both retention and end-to-end accuracy across the 100K–1M token range, where baselines degrade sharply. On RULER-HQA at 896K tokens, MHM-LRU improves the memory retention rate from less than 30% to 73.96%. These gains generalize across model families, scales, and task types, positioning architectural optimization as a practical and cost-efficient path toward reliable long-context recurrent memory. Our code is available here.

01 / Diagnosis

The bottleneck is keeping what was found.

Recurrent agents compress incoming chunks into a fixed-size memory. As the sequence grows, useful information must survive more updates. The paper separates two sources of failure:

Capture / MCR

Was the answer ever in memory?

Memory capture rate measures the fraction of queries whose ground-truth answer appears in memory at least once during processing.

Retention / MRR

Did captured information survive?

Memory retention rate measures survival to the final memory state among queries where capture occurred. It is distinct from final-answer accuracy.

MemAgent capture remains stable as context grows, while retention declines; retention and accuracy show a correlation of 0.98.
Retention failure dominates in the diagnostic experiments. MemAgent on RULER-HQA with Qwen2.5-14B-Instruct; the correlation analysis includes MemAgent and ReMem. Retention falls below 30% at 896K tokens, while capture remains comparatively stable.

02 / Method

Read all heads. Write to one.

MHM splits memory into H independent text blocks. A selection stage chooses a head; the update stage receives the full memory, query, and current chunk, then rewrites only that head. The other H − 1 heads remain unchanged. A single-headed agent is the special case H = 1.

Select

Least recently updated

MHM-LRU chooses the head that has gone longest without an update. Ties use the lowest head index. This rule distributes writes uniformly and needs no additional LLM call for selection.

Update

Protect the unselected heads

The model can use information from every head to write the selected one. Protection is structural for unselected heads at each step; the method does not guarantee that every fact is retained forever.

Try the update schedule

Step 0: all heads start empty.

ReadyMemory head 1Not yet written
ReadyMemory head 2Not yet written
ReadyMemory head 3Not yet written
ReadyMemory head 4Not yet written

Illustration of the deterministic LRU schedule, not a live model run. No head is updated twice before every other head has been updated once.

03 / Evaluation

Better retention. Stronger long-context answers.

Main comparisons use Qwen2.5-14B-Instruct with 5,000-token input chunks and equal total memory capacity: 4,096 tokens for MemAgent, 2 × 2,048 for ReMem, and 4 × 1,024 for MHM-LRU. Recurrent methods report means and standard deviations over three independent runs.

Retention curves on RULER-HQA and BABILong show MHM-LRU retaining more captured information than MemAgent and ReMem at long contexts.
Memory retention rate (MRR). Error bars show standard deviations across three runs. At 1M tokens on BABILong, MHM-LRU reaches 68.68% retention versus 42.96% for MemAgent.

End-to-end answer accuracy

Full results across evaluated context lengths. Higher is better; bold values indicate the best mean in each column.

RULER-HQA · Accuracy (%) ↑
Method7K14K28K56K112K224K448K896K
LLM60.1660.9450.0057.0350.0037.508.590.00
MemAgent62.24± 0.9760.94± 1.1056.25± 2.2150.52± 0.3742.19± 0.6442.97± 0.6433.33± 0.9821.62± 0.98
ReMem64.32± 3.5159.64± 0.9850.26± 3.6839.06± 2.7829.68± 2.7826.56± 1.6916.14± 1.3313.80± 0.97
MHM-LRU66.86± 1.6962.76± 1.4759.64± 1.3355.21± 1.3350.26± 1.4750.26± 2.5748.70± 2.8849.74± 1.47

Recurrent methods: mean ± standard deviation over three runs. LLM: native long-context backbone; no standard deviation reported. On narrow screens, scroll the table horizontally.

BABILong · Accuracy (%) ↑
Method8K16K32K64K128K256K512K1M
MemAgent58.85± 2.6655.99± 2.5853.39± 2.4247.14± 0.9844.66± 1.6438.02± 2.8832.55± 1.6125.26± 0.97
ReMem61.46± 2.4155.21± 2.2445.57± 1.6142.19± 3.9934.64± 2.4232.29± 1.6129.50± 0.2624.48± 0.97
MHM-LRU63.54± 1.6060.94± 2.9255.47± 1.2751.82± 0.9748.70± 3.2747.39± 1.8446.35± 2.6641.41± 1.91

Recurrent methods: mean ± standard deviation over three runs. On narrow screens, scroll the table horizontally.

At 896K on RULER-HQA, MHM-LRU achieves 49.74% accuracy versus 21.62% for MemAgent and 13.80% for ReMem. At 1M on BABILong, accuracy is 41.41% versus 25.26% and 24.48%, respectively.
Evaluation and sampling details

RULER-HQA evaluates multi-hop question answering from 7K to 896K tokens in the reported tables. BABILong evaluates ten reasoning task types from 8K to 1M tokens, with 128 systematically sampled entries per length. Accuracy uses substring exact matching of normalized ground truth against the extracted final answer.

Decoding uses temperature 0.7 and top-p 0.95. The paper reports deployment with vLLM 0.15.0 on a Linux server with eight A100 GPUs. The runtime comparison below uses an NVIDIA A100 80GB with four parallel threads.

04 / Analysis

What drives the improvement?

Generalization across backbones

The retention benefit also appears with Qwen2.5-32B-Instruct and gpt-oss-120b in the evaluated 112K–896K range.

Memory retention rate (%) ↑
BackboneMethod112K224K448K896K
gpt-oss-120bMemAgent75.0055.8461.6435.71
gpt-oss-120bMHM-LRU84.4875.4469.3568.85
Qwen2.5-32BMemAgent84.0487.6388.5478.02
Qwen2.5-32BMHM-LRU93.0295.6590.0081.91
Accuracy across ten BABILong task types and four long-context settings, comparing MHM-LRU and MemAgent.
Ten reasoning task types. At 1M tokens, MHM-LRU outperforms MemAgent across all ten BABILong task types, with pronounced gains on QA4–QA6.
Stage-wise selection and routing

At 896K, merging head selection and writing into one model operation (MHM-Concur) lowers retention to 54.12% and accuracy to 33.59%, compared with 73.96% and 49.74% for MHM-LRU. Model-based relevance routing reaches 45.31% accuracy; its 74.71% retention is slightly higher than LRU at this length, but its capture rate is lower (67.97% versus 69.16%).

More heads: promising, with a capacity caveat

In the head-count ablation at 896K, retention rises from 26.53% with one head to 69.89% with two, 74.67% with four, and 87.12% with eight, before plateauing at sixteen. These are separate ablation measurements.

Head-count ablation showing increasing memory retention as head count grows.

Capacity caveat: this ablation fixes the per-head generation limit at 1,024 tokens, so total memory capacity increases with head count. Unlike the main comparison, it does not isolate head count at a fixed total capacity.

Runtime and token costs

LRU selection adds no model inference or selection tokens. Actual total runtime and output-token costs still vary. At 112K, MHM-LRU takes 25.9 seconds per sample versus 31.5 for MemAgent. At 896K, it takes 239.1 seconds versus 224.9 (about 6.3% longer).

Measured inference costs on RULER-HQA
MetricMethod7K112K896K
Seconds / sampleMemAgent2.831.5224.9
Seconds / sampleMHM-LRU1.925.9239.1
Output tokens / sampleMemAgent330.62834.418013.9
Output tokens / sampleMHM-LRU142.81875.718085.5

Selected context lengths from the paper’s cost table. Tokens are generated output tokens per sample, not total input tokens.

A memory trajectory: preserving “office”

In the paper’s BABILong-512K QA4 example, the question is “What is the bathroom south of?” Both methods capture the answer, “office,” at step 74. MemAgent loses it during the final updates. MHM-LRU propagates it across heads, retains it through step 95, and answers correctly.

Case study: MemAgent overwrites the answer office near the end of processing, while MHM-LRU retains copies in multiple heads.
A representative case from the paper, illustrating emergent redundancy rather than a universal guarantee.

Scope and limitations

The experiments use a fixed number of heads. The best head allocation may depend on context length and task structure; adaptive allocation remains future work. The evaluated setting is query-specific long-context reasoning. Lifelong, query-agnostic memory and self-evolving agents are prospective directions, not demonstrated results.

Reference

Citation

BibTeX for the supplied manuscript.

@article{li2026multi,
  title={Multi-Head Recurrent Memory Agents},
  author={Li, Jiatong and Yeh, Samuel and Li, Sharon},
  journal={arXiv preprint arXiv:2607.01523},
  year={2026}
}