flâneur — a map of the web's best reading

Next Steps in Developmental Interpretability | Manifund

manifund.org · 2,916 words · saved by 1 readers

This project builds upon our successful Manifund 2023-supported research that validated Developmental Interpretability (DevInterp) as an approach to understanding the internal structure of neural networks. We (Timaeus) are seeking funding to continue this research and extend our runway from 6 months to ~1 year. In particular, this funding would go to a series of projects (described below) that apply our existing research to address immediate, real-world safety concerns. This has the ultimate goal of leading to novel understanding-based evals. We initially set out to establish the viability of Developmental Interpretability, an application of Singular Learning Theory (SLT) to interpreting the formation of structure in neural networks. This succeeded. Together with our collaborators, We demonstrated that SLT can inform the development of new measurements, such as local learning coefficient (LLC) estimation. This has now been successfully applied to models with up to 100s of millions of p

Next Steps in Developmental Interpretability | Manifund 10 Next Steps in Developmental Interpretability Technical AI safety Jesse Hoogland Active Grant $80,680 raised $670,000 funding goal Donate Sign in to donate Project Summary This project builds upon our successful Manifund 2023-supported research that validated Developmental Interpretability (DevInterp) as an approach to understanding the internal structure of neural networks. We ( Timaeus ) are seeking funding to continue this research and extend our runway from 6 months to ~1 year. In particular, this funding would go to a series of pro

Explore this link on the map →

related reading