Language, trees, and geometry in neural networks
Language is made of discrete structures, yet neural networks operate on continuous data: vectors in high-dimensional space. A successful language-processing network must translate this symbolic information into some kind of geometric representation—but in what form? Word embeddings provide two well-known examples: distance encodes semantic similarity, while certain directions correspond to polarities (e.g. male vs. female). A recent, fascinating discovery points to an entirely new type of representation. One of the key pieces of linguistic information about a sentence is its syntactic structure. This structure can be represented as a tree whose nodes correspond to words of the sentence. Hewitt and Manning, in A structural probe for finding syntax in word representations, show that several language-processing networks construct geometric copies of such syntax trees. Words are given locations in a high-dimensional space, and (following a certain transformation) Euclidean distance between
--> --> Language, trees, and geometry in neural networks --> Language, trees, and geometry in neural networks Part I (see Part II ) of a series of expository notes accompanying this paper , by Andy Coenen, Emily Reif, Ann Yuan, Been Kim, Adam Pearce, Fernanda Viégas, and Martin Wattenberg. These notes are designed as an expository walk through some of the main results. Please see the paper for full references and details. Language is made of discrete structures, yet neural networks operate on continuous data: vectors in high-dimensional space. A successful language-processing network mu
saved by
related reading
- The Illustrated Word2vec – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- [2602.15029] Symmetry in language statistics shapes the geometry of model representationsarxiv.org
- The Illustrated Word2vec – Jay Alammar – Visualizing machine learning one concept at a time.jalammar.github.io
- The World Inside Neural Networksgoodfire.ai
- 1301.3781arxiv.org
- What Is ChatGPT Doing … and Why Does It Work?-Stephen Wolfram Writingswritings.stephenwolfram.com
- Word Embeddingslena-voita.github.io
- Embeddings: What they are and why they mattersimonwillison.net
- Deep Learning, NLP, and Representations - colah's blogcolah.github.io
- An intuitive introduction to text embeddings - Stack Overflowstackoverflow.blog
- A History of Large Language Modelsgregorygundersen.com
- An intuitive introduction to text embeddings - Stack Overflowstackoverflow.blog