flâneur — a map of the web's best reading

Anthropic takes a look into the ‘black box’ of AI models - Fast Company

fastcompany.com · saved by 1 readers

Progress in mechanistic interpretability could lead to major advances in making large AI models safe and bias-free. [Images: Iana Kunitsa/Getty Images; fotograzia/Getty Images] BY MARK SULLIVAN 5 MINUTE READ Welcome to AI Decoded, Fast Company’s weekly newsletter that breaks down the most important news in the world of AI. You can sign up to receive this newsletter every week here. Today’s AI models are so big and so complex (they’re fashioned after the human brain) that even the PhDs who design them know relatively little about how they actually “think.” Until pretty recently, the study of “mechanistic interpretability” has been mostly theoretical and small-scale. But Anthropic published new research this week showing some real progress. During its training, an LLM processes a huge amount of text and eventually forms a many-dimensional map of words and phrases, based on their meanings and the contexts within which they’re used. After the model goes into use, it draws on this “map” to

Explore this link on the map →

saved by