
ModularRSI, by Siwei Wu and colleagues, targets the code surrounding an agent: its execution loop, tools, observations, context, and completion checks. Changes must pass validation before integration.
The authors curate 2,000 tasks separate from evaluation benchmarks; experiments use 120 tasks per evolution domain. With DeepSeek-V4-Flash-Preview, reported in-domain accuracy rises from 47.57% to 52.43% on TerminalBench 2.0 and from 73.40% to 76.45% on SWE-bench Verified. They also report cross-domain and cross-model transfer.
These results concern harness adaptation under the paper's experimental protocol, not a demonstration of unrestricted recursive intelligence growth.
Related RSI coverage:
Sources: paper · full text · code and data.
ModularRSI: Modular and Generalizable Recursive Harness Self-ImprovementFull paperModularRSI code and data

Meta-Harness: A system that optimizes LLM harnesses automatically

Sakana AI and UC Berkeley propose RHI: self-iterating harnesses cut costs 60%

Meta^n: Recursive Self-Improvement through Emergent Depth

SIA: Self Improving AI with Harness & Weight Updates

Continual Harness: A self-improving loop for embodied agents

Chinese researchers publish a roadmap to recursive self-improvement