✳flâneur — a map of the web's best reading
What working at Mechanize is like | Mechanize, Inc.
mechanize.work · 1,693 words · saved by 1 readers
What working at Mechanize is like: the office, the work, and the team.
What working at Mechanize is like | Mechanize, Inc. What working at Mechanize is like As an engineer at Mechanize, most of the work you'll be doing will be directly connected to our core business of producing high-quality and realistic software engineering tasks for use in reinforcement learning or evaluations of model capabilities. You can think of a "task" as being the equivalent of a take-home assessment for a coding agent: it has a prompt telling the model what to do, enough time and space for the model to implement a complex solution, and a grader or rubric to assign a numerical score to
Explore this link on the map →related reading
- Demystifying evals for AI agents \ Anthropicanthropic.com
- After Automation | Everyevery.to
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forumalignmentforum.org
- AI Agent Benchmark for Real-World Professional Workflowsagents-last-exam.org
- How To Become A Mechanistic Interpretability Researcher — LessWronglesswrong.com
- Hey, N00b, We Didn't Hire You to Complete Tasksnewsletter.kentbeck.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Parsed | Custom, interpretable AI systems that continuously learnparsed.com
- GitHub - openai/mle-bench: MLE-bench is a benchmark for measuring how well AI agents perform at machine learning engineering · GitHubgithub.com
- Manus AI: The Best Autonomous AI Agent Redefining Automation and Productivityhuggingface.co
- Lecture 01. Strong Models Don't Mean Reliable Execution | Learn Harness Engineeringwalkinglabs.github.io
- Prompt guidance | OpenAI APIdevelopers.openai.com