Training LLMs to use Photoshop | Mickey Labs
Frontier models are getting better at computer use, but they remain too slow and too expensive to use within creative software. Designing a movie poster in Photoshop can take hundreds of turns, and every step through a frontier API adds latency and cost. We wanted to see if a fine-tuned lower-cost model could do the same work faster and cheaper. In this post, we show that a simple training recipe can fine-tune Qwen 3.6 27B to match Sonnet 4.6 and approach GPT-5.4 on basic photo-editing tasks. We post-trained the base model on Tinker: first with SFT on frontier trajectories, then with RL in PhotopeaBench, a new internal eval and environment. Our fine-tuned model reaches 96% of GPT-5.4's quality at two-thirds the cost, and slightly exceeds Sonnet 4.6 in performance for about 4x cheaper. PhotopeaBench is a benchmark and reinforcement learning environment for basic tool use and composition in photo-editing software. We originally designed the benchmark around Adobe Photoshop, but meaningfu
Our fine-tuned (SFT + RL) Qwen 3.6 27B model executing the x-03-batman-campus task from the held-out eval set, running entirely through basic computer-use primitives: click, drag, type, key. The video is sped up. Introduction PhotopeaBench Training Recipe Supervised Finetuning Reinforcement Learning Benchmarks Conclusion Future Work Frontier models are getting better at computer use, but they remain too slow and too expensive to use within creative software. Designing a movie poster in Photoshop can take hundreds of turns, and every step through a frontier API adds latency and cost. We wanted
Explore this link on the map →related reading
- Composer2.pdfcursor.com
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- DeepSeek-R1arxiv.org
- PostTrainBenchposttrainbench.com
- Replicate - Run AI with an APIreplicate.com
- AI Model & API Providers Analysis | Artificial Analysisartificialanalysis.ai
- State of RL for reasoning LLMs | A. Weersaweers.de
- frontier model training methodologies | Alex Wa's Blogdjdumpling.github.io
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- The State of Reinforcement Learning for LLM Reasoningmagazine.sebastianraschka.com
- Language Models can Solve Computer Tasksarxiv.org
- 2025: The year in LLMssimonwillison.net