Building self-improving tax agents with Codex | OpenAI
By Members of Technical Staff: Aravind Srinivasan & Samay Shamdasani (Thrive Holdings), Arthur Fernandes Araujo & John de Wasseige (OpenAI) Table of contents How Thrive Holdings and OpenAI co-developed Tax AI for Crete accountants by fusing practitioner expertise with a Codex-driven loop Real-world systems behave differently in production than they do in a lab, breaking in ways that are hard to anticipate before deployment. Teams often discover those failures after launch, then spend weeks inspecting edge cases, adjusting prompts, and translating production feedback into durable product improvements. The feedback loop is manual and slow, and only improves when an engineer advances it. But today, with thoughtfully designed eval infrastructure, direct access to practitioners and real world environments, and the frontier agentic capabilities of Codex, you can build agents that self-improve. In this post, we’ll unpack how we used Codex to build this type of agent. Over the past six months,
May 27, 2026 Engineering Building self-improving tax agents with Codex By Members of Technical Staff: Aravind Srinivasan & Samay Shamdasani (Thrive Holdings), Arthur Fernandes Araujo & John de Wasseige (OpenAI) Loading… Share How Thrive Holdings and OpenAI co-developed Tax AI for Crete accountants by fusing practitioner expertise with a Codex-driven loop Real-world systems behave differently in production than they do in a lab, breaking in ways that are hard to anticipate before deployment. Teams often discover those failures after launch, then spend weeks inspecting edge cases, adjusting prom
Explore this link on the map →saved by
related reading
- Killing Coding Agent Slop With Adversarial Self-Playusetelos.ai
- When AI builds itself \ Anthropicanthropic.com
- Shipping at Inference-Speed | Peter Steinbergersteipete.me
- A Guide to Claude Code 2.0 and getting better at using coding agents – sankalp's blogsankalp.bearblog.dev
- My AI Had Already Fixed the Code Before I Saw Itevery.to
- After Automation | Everyevery.to
- Auto-review of agent actions without synchronous human oversightalignment.openai.com
- AI Horseless Carriages | koomen.devkoomen.dev
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- Introducing GPT-5.3-Codex | OpenAIopenai.com
- Best practices for Claude Code - Claude Code Docsanthropic.com
- Codex as an assistant - the power and pitfalls – Huizi Mao –ralphmao.github.io