[2103.03386] Clusterability in Neural Networks
The learned weights of a neural network have often been considered devoid of scrutable internal structure. In this paper, however, we look for structure in the form of clusterability: how well a network can be divided into groups of neurons with strong internal connectivity but weak external connectivity. We find that a trained neural network is typically more clusterable than randomly initialized networks, and often clusterable relative to random networks with the same distribution of weights. We also exhibit novel methods to promote clusterability in neural network training, and find that in multi-layer perceptrons they lead to more clusterable networks with little reduction in accuracy. Understanding and controlling the clusterability of neural networks will hopefully render their inner workings more interpretable to engineers by facilitating partitioning into meaningful clusters.
Clusterability in Neural Networks Daniel Filan1, * , Stephen Casper2, * , Shlomi Hod3, * , Cody Wild1 , Andrew Critch1 , Stuart Russell1 1 UC Berkeley, 2 Harvard, 3 Boston University {daniel filan, codywild, critch,…
related reading
- Zoom In: An Introduction to Circuitsdistill.pub
- Toy Models of Superpositiontransformer-circuits.pub
- Neural Networks, Manifolds, and Topology -- colah's blogcolah.github.io
- The Building Blocks of Interpretabilitydistill.pub
- [2601.12703] Towards Spectroscopy: Susceptibility Clusters in Language Modelsarxiv.org
- Distill — Latest articles about machine learningdistill.pub
- A Recipe for Training Neural Networkskarpathy.github.io
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- On neural scaling and the quanta hypothesisericjmichaud.com
- The World Inside Neural Networksgoodfire.ai
- [2404.19756] KAN: Kolmogorov-Arnold Networksarxiv.org
- nn-notes.pdfboris-hanin.github.io