- Most RL recipes train against one coding environment
- Model overfits that scaffold
- Skills may not transfer
- Rotate coding harnesses during training
- Built on TRL + Harbor
- Generalizes across agent scaffolds
Why it matters: Most RL recipes assume a single environment; this shows how to train against multiple harnesses so models generalize across agent scaffolds instead of overfitting one.
How to apply: Read the guide, then wire your own harness into TRL + Harbor and run a small multi-harness RL job on a 7B-30B open model before scaling.