
Read the full research write-up · Download the paper (PDF)
Researchers from Hugging Face and Liquid AI have released an open framework for training small language models inside multiple coding-agent harnesses. An agent harness controls the tools, context and execution loop around a model; changing it can substantially change how well the same weights perform.
The framework combines OpenEnv, Harbor and TRL. A capture proxy records the exact tokens and generation probabilities needed for reinforcement learning while Claude Code, Codex, OpenCode and Mini-SWE-Agent run without changes to their code.
On held-out SmolDataEnvs data-analysis tasks, multi-harness training raised LFM2.5-2.6B's average pass@1 from 42.2% to 54.2%, with improvements across all four harnesses. It also reduced tool calls by 31% on tasks that both the trained model and the baseline solved. A small efficiency bonus rewarded correct solutions that used fewer calls.
Training only in OpenCode reached 52.3% overall and performed best within OpenCode. The 1.9-point overall gap between the two training approaches was within evaluation noise. The clearer advantage of multi-harness training was its broader efficiency gains and stronger results under Claude Code and Codex. Supervised fine-tuning on successful rollouts produced smaller accuracy gains.
The work offers a practical way to train open models for several agent interfaces at once. These results come from a small model on data-analysis tasks, so they do not establish the same gains on large-scale software-engineering benchmarks or harnesses absent from training.
Full research write-up — The ultimate guide to multi-harness RLThe ultimate guide to multi-harness RL (PDF)