Announcing Mechanize Inc.
Today we’re announcing Mechanize, a startup focused on developing virtual work environments, benchmarks, and training data that will enable the full automation of the economy. We will achieve this by creating simulated environments and evaluations that capture the full scope of what people do at their jobs. This includes using a computer, completing long-horizon tasks that lack clear criteria for success, coordinating with others, and reprioritizing in the face of obstacles and interruptions. We’re betting that the lion’s share of value from AI will come from automating ordinary labor tasks rather than from “geniuses in a data center”. Currently, AI models have serious shortcomings that render most of this enormous value out of reach. They are unreliable, lack robust long-context capabilities, struggle with agency and multimodality, and can’t execute long-term plans without going off the rails. To overcome these limitations, Mechanize will produce the data and evals necessary for compr
We build environments and evals for frontier coding agents. In these environments, models carry out software engineering work such as building a feature, deploying an application, or debugging an issue in an unfamiliar codebase. A grader scores the model’s performance, and these scores serve as signals during reinforcement learning and evaluations. Frontier models are already surprisingly good at writing code. Our engineers find where they still break down and build environments that reveal those limits. Our current focus is software engineering, but our long-term goal is the full…
saved by
related reading
- How to fully automate software engineering | Mechanize, Inc.mechanize.work
- Harness Engineering for Self-Improvement | Lil'Loglilianweng.github.io
- What working at Mechanize is like | Mechanize, Inc.mechanize.work
- FrontierSWEfrontierswe.com
- After Automation | Everyevery.to
- Trending Papers - Hugging Facepaperswithcode.com
- RSI Simulator | Paradigmparadigm.xyz
- The upcoming GPT-3 moment for RL | Mechanize, Inc.mechanize.work
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- Convictionconviction.com
- Measuring the performance of our models on real-world tasks | OpenAIopenai.com
- How Zapier Turned AutomationBench Into a Continuous Agent Improvement Loopprimeintellect.ai