---
title: RRSI (Regularized Recursive Self-Improvement of Agent Harnesses)
type: vocabulary
url: "https://www.envisioning.com/vocab/rrsi-regularized-recursive-self-improvement-of-agent-harnesses"
summary: "A regularized search method that evolves an AI agent's harness so gains transfer beyond the tasks it was tuned on, instead of overfitting them."
year: 2026
generality: 0.30
---

# RRSI (Regularized Recursive Self-Improvement of Agent Harnesses)

A regularized search method that evolves an AI agent's harness so gains transfer beyond the tasks it was tuned on, instead of overfitting them.
RRSI (Regularized Recursive Self-Improvement of Agent Harnesses), from a September 2026 Google Research paper by Peng Xia and 13 co-authors (arXiv:2609.24972), evolves an LLM agent's harness: the prompts, control flow, tools, memory and context management around a frozen backbone model, without overfitting to its training tasks. Proposing and selecting harness edits is already a form of recursive self-improvement (RSI). RRSI regularizes that search so gains on the training tasks, the "evolve set", carry over to benchmarks it never sees.

Regularization runs on both ends of the loop. On the proposal side, an annealed budget caps how many edits one candidate may bundle, the proposer reads the full edit history so a falsified hypothesis is not retried, and a stalled search moves toward untouched components. On the selection side, a critic screens candidates for benchmark-specific logic before scoring, a noise-adjusted floor rejects gains inside evaluation variance, and a cost rule requires added tokens be earned back by measured gain.

Across eight benchmarks in coding, agentic workspace and engineering design, the authors report gains up to 14.1 points on the evolved split and up to 4.7 out of distribution, on 30% fewer tokens than unregularized evolution. Terminal-Bench 2.1 rose from 74.2% to 80.2%, and SWE-bench Verified, held out, rose from 82.0% to 83.8%. RRSI differs from ModularRSI (arXiv:2609.14857, a different team), which evolves harnesses through modular, contrastive edits instead.

---
Source: Envisioning — Technology Research Institute (https://www.envisioning.com/vocab/rrsi-regularized-recursive-self-improvement-of-agent-harnesses)
