Adithya S K on X: "RL Coding Environments 101: Why Harbor Exists" / X
x.com · 652 words · saved by 1 readers
https://t.co/tOnb7uElhs
@adithya_s_k: RL Coding Environments 101: Why Harbor Exists Coding as an RL task is having a moment. And if you've actually tried to do RL on a coding task, you already know the model is the easy part. It's everything around the model that eats your week. disclaimer : I have used Claude Code with Claude Opus to help me put this together but I've been working with Harbor pretty extensively across a bunch of projects, and figured it was time to write down why I think it's the right abstraction/framework with respect to the over all RL coding Environment landscape You need a real environment. Ri
saved by
related reading
- Coding Models Are Doing Too Much | whnrehiew.github.io
- Pavlov's List: A List of RL Environment Startupspavlovslist.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Composer2.pdfcursor.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- The Verification Horizon: No Silver Bullet for Coding Agent Rewardsarxiv.org
- Improving Composer through real-time RL · Cursorcursor.com
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- Reinforcement Learning With Verifiable Rewards: How Data and Verifiers Shape RLVRsnorkel.ai
- Our Problems · Proximalproximal.ai
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Datasemianalysis.com