flâneur — a map of the web's best reading

Don't Build an RL Environment Startup

benanderson.work · 1,350 words · saved by 1 readers

The first person who sold an RL environment to a frontier AI lab must have felt like they discovered an infinite money glitch. It's no longer a secret that frontier AI labs regularly pay hundreds of thousands, and sometimes millions, for clones of Linear and Salesforce. If you're reading this, you've probably thought about quitting your day job and starting a company that builds these unusually lucrative Next.js apps. In this post, I'll argue that you should hesitate before hopping on the bandwagon. For those unfamiliar, an RL (reinforcement learning) environment is like a sandbox for AI models like Claude and GPT to learn from. It keeps track of an internal state, prompts the AI to take actions to complete a task, and assigns a score based on the outcome. The most obvious kind is a clone of a popular website or enterprise software tool like Doordash, Linear, or Amazon, which teaches the AI to click around and order pizza. It can also be text-only, like the TextArena project, which tea

--> Don't Build an RL Environment Startup Don't Build an RL Environment Startup Don't sell blood to vampires. Posted Sep 7, 2025 by Benjamin Anderson The first person who sold an RL environment to a frontier AI lab must have felt like they discovered an infinite money glitch. It's no longer a secret that frontier AI labs regularly pay hundreds of thousands, and sometimes millions, for clones of Linear and Salesforce. If you're reading this, you've probably thought about quitting your day job and starting a company that builds these unusually lucrative Next.js apps. In this post, I'll argue tha

Explore this link on the map →

related reading