flâneur — a map of the web's best reading

@jeremy-berman/arc-agi on Params

params.com · 4,431 words · saved by 1 readers

I co-founded and built the backend for beatgig.com. Now I'm building params.com and working on machine learning models. I think ARC-AGI is the most important benchmark we have today. It’s surprising that even the most sophisticated Large Language Models (LLMs), like OpenAI o1 and Claude Sonnet 3.5, struggle with simple puzzles that humans can solve easily. This highlights the core limitation of current LLMs: they're bad at reasoning about things they weren't trained on. They are bad at generalizing. After reading Ryan Greenblatt’s blog post on how he achieved a state of the art 43% accuracy on ARC-AGI-Pub, I wondered if we could push these models further. Could it be that frontier models might actually possess the necessary intelligence and understanding to solve ARC? Maybe if they are poked and prodded enough, they’ll spit out the right answer. After lots of experimenting, I got a record of 53.6% on the public leaderboard using Sonnet 3.5.1 This is a significant improvement over the p

@jeremy-berman/arc-agi on Params @ jeremy-berman / arc-agi Fork Jeremy Berman AI @ Params Profile Book a Call About Jeremy Berman I co-founded and built the backend for beatgig.com. Now I'm building params.com and working on machine learning models. Docs Files Documentation Introduction Setup How I came in first on ARC-AGI-Pub using Sonnet 3.5 with Evolutionary Test-time Compute I think ARC-AGI is the most important benchmark we have today. It’s surprising that even the most sophisticated Large Language Models (LLMs), like OpenAI o1 and Claude Sonnet 3.5, struggle with simple puzzles that huma

Explore this link on the map →

saved by

related reading