✳flâneur — a map of the web's best reading
Adithya S K on X: "RL Coding Environments 101: Why Harbor Exists" / X
x.com · 652 words · saved by 1 readers
https://t.co/tOnb7uElhs
@adithya_s_k: RL Coding Environments 101: Why Harbor Exists Coding as an RL task is having a moment. And if you've actually tried to do RL on a coding task, you already know the model is the easy part. It's everything around the model that eats your week. disclaimer : I have used Claude Code with Claude Opus to help me put this together but I've been working with Harbor pretty extensively across a bunch of projects, and figured it was time to write down why I think it's the right abstraction/framework with respect to the over all RL coding Environment landscape You need a real environment. Ri
Explore this link on the map →saved by
related reading
- True Agents Model the Worldprimeintellect.ai
- Coding Models Are Doing Too Much | whnrehiew.github.io
- State of RL for reasoning LLMs | A. Weersaweers.de
- 2025 LLM Year in Review – karpathykarpathy.bearblog.dev
- Composer2.pdfcursor.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Improving Composer through real-time RL · Cursorcursor.com
- The Verification Horizon: No Silver Bullet for Coding Agent Rewardsarxiv.org
- Scaling Reinforcement Learning: Environments, Reward Hacking, Agents, Scaling Datasemianalysis.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Introducing OpenReward | General Reasoninggr.inc