Sidak Pal Singh on X: "Trajectory Map as the Tool: Simply flatten the parameters at the various points in the trajectory, put them in a matrix, & compute the gram matrix C of their pairwise (cosine) similarities. Compute an interpretable measure of the mean directional similarity (MDS) during training. https://t.co/IdHrWifCnG" / X
To view keyboard shortcuts, press question mark View keyboard shortcuts Messages Home Explore Notifications Messages Grok Bookmarks Communities Premium Profile More Post Adithya Vellal @avellal14 Post See new posts Conversation Sidak Pal Singh @unregularized · Jul 12 Ever wondered how the optimization trajectories are like when training neural nets & LLMs? Do they contain a lot of twists and turns, or does the direction largely remain the same? We explore this in our work for LLMs (upto 12B params) + ResNets on ImageNet. Key findings 1 12 47 6K Sidak Pal Singh @unregularized Trajectory Map as the Tool: Simply flatten the parameters at the various points in the trajectory, put them in a matrix, & compute the gram matrix C of their pairwise (cosine) similarities. Compute an interpretable measure of the mean directional similarity (MDS) during training. 11:55 AM · Jul 12, 2024 · 3,282 Views 2 3 10 6 Post your reply Reply Sidak Pal Singh @unregularized · Jul 12 Example: consider ResNet
Sidak Pal Singh @unregularized Jul 12, 2024 Ever wondered how the optimization trajectories are like when training neural nets & LLMs🤔? Do they contain a lot of twists 💃 and turns, or does the direction largely remain the same🛣️? We explore this in our work for LLMs (upto 12B params) + ResNets on ImageNet. Key findings👇 2 0 2 10 0 1 0 63 0 6 3 10K 0 1 0 K
Explore this link on the map →related reading
- LLM Visualizationbbycroft.net
- GenAI Handbookgenai-handbook.github.io
- NL.pdfabehrouz.github.io
- Neuronpedianeuronpedia.org
- GitHub - jacobhilton/deep_learning_curriculum: Language model alignment-focused deep learning curriculum · GitHubgithub.com
- GitHub - linkedin/Liger-Kernel: Efficient Triton Kernels for LLM Training · GitHubgithub.com
- microgptkarpathy.github.io
- LLM Resourcesforrestbicker.com
- Jacobian Lens – Qwen3.6-27B | Neuronpedianeuronpedia.org
- Getting Caught Up to Modern LLM Research | Samarth Goeldev.samarthgoel.com
- Natural Intelligencegreydanus.github.io
- GitHub - inverse-scaling/prize: A prize for finding tasks that cause large language models to show inverse scaling · GitHubgithub.com