flâneur — a map of the web's best reading

AI Agents and Painted Facades | Fulcrum Research

fulcrumresearch.ai · 120 words · saved by 1 readers

In 1787, Catherine the Great sailed down the Dnieper to inspect its banks. Her trusted advisor, Governor Potemkin, set out to present those war-torn lands to her in the best possible light. Legend has it1 Potemkin set up painted facades along the riverbank, so that, from her barge, Catherine would see beautiful villages – each just a couple of inches thick. The rise of AI agents makes the Potemkin problem commonplace. Research agents cite experiments that never took place. Coding agents often write fake tests and mock solutions as they cause catastrophes behind the scenes. We’re moving towards a world of Potemkin villages – where our understanding of reality drifts farther and farther from what is actually happening. At some point, we might stop catching our agents painting facades. To avoid this, we need to understand AI agents and their effects on the world. Evaluations are currently the best guess on how to do this, but they are an incomplete solution. Evaluations (or evals) measure

Fulcrum Fulcrum | Scaling human intent in software Fulcrum We are a lab working on elicitation: structuring both a model's tools and processes to get the best outputs from it, and building technology to measure the quality of these outputs. Current models have capabilities far beyond what they show, particularly on tasks that are hard to verify. We believe eliciting models properly is critical to navigating the intelligence explosion safely. We measure model capabilities, find failure modes, and elicit them more effectively on the hardest and fuzziest tasks. If you want to get the best agent p

Explore this link on the map →

related reading