flâneur — a map of the web's best reading

Language, Statistics, & Category Theory, Part 1

math3ma.com · 2,206 words · saved by 1 readers

In the previous post I mentioned a new preprint that John Terilla, Yiannis Vlassopoulos, and I recently posted on the arXiv. In it, we ask a question motivated by the recent successes of the world's best large language models: To understand the motivation behind this question, and to recall what a "large language model" is, I'll encourage you to read the opening article from last time. In the next few blog posts, I'll give a tour of mathematical ideas presented in the paper towards answering the question above. I like the narrative we give, so I'll follow it closely here on the blog. You might think of the next few posts as an informal tour through the formal ideas found in the paper. Now, where shall we begin? What math are we talking about? Let's start with a simple fact about language. By "algebraic," I mean the basic sense in which things combine to form a new thing. We learn about algebra at a young age: given two numbers x 𝑥 and y 𝑦 we can multiply them to get a new number

Language, Statistics, & Category Theory, Part 1 - math3ma    © 2015 – 2025 Math3ma Ps. 148 Archives June 29, 2021 • Category Theory Language, Statistics, & Category Theory, Part 1 In the previous post I mentioned a new preprint that John Terilla, Yiannis Vlassopoulos, and I recently posted on the arXiv. In it, we ask a question motivated by the recent successes of the world's best large language models: What's a nice mathematical framework in which to explain the passage from probability distributions on text to syntactic and semantic information in language? To understand the motivation be

Explore this link on the map →

related reading