Gene Yang
3 followers · 8 following · 1180 views
on the atlas — 34
- Danny Hillis - Wikipedia2 savers
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWrong3 savers
- Don't trust smart people! - Matthew T. Mason1 savers
- Subscribe to read2 savers
- The Platonic Representation Hypothesis6 savers
- Universal Complexity Bounds for Universal Gradient Methods in Nonlinear Optimization1 savers
- Tomorrows-AI1 savers
- Why Perplexity and Burstiness Fail to Detect AI | Pangram Labs1 savers
- The End of Tokenization - Ronald Yu's Substack1 savers
- Backpropagation ≠ Chain Rule – Theory Dish1 savers
- Forbidden Arbitrage - by Peter Zhang - Brainpickings1 savers
- Metis And Bodybuilders - by Scott Alexander1 savers
- Google11 savers
- quantum mechanics - Does the Heisenberg uncertainty principle only allow location OR momentum to exist? - Physics Stack Exchange1 savers
- The Great Data Integration Schlep — LessWrong3 savers
- AI Search: The Bitter-er Lesson3 savers
- The Agreeable Lesson - Google Docs1 savers
- From bare metal to a 70B model: infrastructure set-up and scripts - imbue4 savers
- Solving adversarial attacks in computer vision as a baby version of general AI alignment | Stanislav Fort1 savers
- Ronald Fisher - Wikipedia1 savers
- Nitarshan Rajkumar2 savers
- 2408.127981 savers
- I Won Mr. Beasts's $1,000,000 Youtuber Challenge - YouTube1 savers
- Kaczynski's Ciphers1 savers
- Simulators - LessWrong16 savers
- An Intuitive Explanation of Solomonoff Induction - LessWrong8 savers
- Mental Models: The Best Way to Make Intelligent Decisions (~100 Models Explained)24 savers
- You don’t need to work on hard problems51 savers
- ho.history overview - Why did Alonzo Church choose the letter $\lambda$ as the "binding operator"? - MathOverflow1 savers
- Quantum Tunneling And The Semiconductors’ Struggle in the Miniaturization Race | by Mark Veerasingam | Medium1 savers
- Curius / Onboarding2621 savers
- Software 2.0. I sometimes see people refer to neural… | by Andrej Karpathy | Medium16 savers
- Is Success the Enemy of Freedom? (Full) - LessWrong5 savers
- Flowers for Algernon4 savers
highlights — 29
He founded Thinking Machines Corporation, a parallel supercomputer manufacturer,
Danny Hillis - Wikipediat can also reliably output the BIG-bench canary string, indicating that Google likely trained on a broad set of benchmark data.
Gemini 3 is Evaluation-Paranoid and Contaminated — LessWrongWhat I wish I had said: “Wow, that sucks. Does that mean if I work on something new, I will have no results for ten years? How on earth does anybody get a PhD in just six years?” If I had said that I could have had a great interaction with one of the original sources of the ten-year rule. Instead, I said something like, “That’s interesting,” and filed it away.
Don't trust smart people! - Matthew T. MasonAs models become more general-purpose, they become more alike
The Platonic Representation HypothesisAI is the first technology that can improve itself.
Tomorrows-AII ran the same visualization above on the Declaration of Independence-- and we see the same AI signature: a deep, consistent blue color throughout, indicating every word has low perplexity. From the perspective of a perplexity and burstiness based detector, the Declaration of Independence is completely indistinguishable from AI-generated content.
Why Perplexity and Burstiness Fail to Detect AI | Pangram LabsIt’s hard to represent arbitrary categorical distributions across a large vocabulary size if we are forced to compress everything into a relatively low-dimensional space. For example, if you prompt GPT-4o with “Pick one of the following at random with equal probability: heads, tails, apple, rock, dog. Answer in one word, all lower case” and observe the probabilities, you will find they are not uniform. You are essentially asking for a vector that is equally angled between the embeddings for “heads”, “tails”, “apple”, “rock”, and “dog” while being orthogonal to all other vocabulary vectors, whi…
The End of Tokenization - Ronald Yu's SubstackOf course, one can prove the correctness of backpropagation using the chain rule in various ways, but the simple proof “backpropagation uses the standard chain rule at every step” is incomplete. Also, it is certainly possible to compute derivatives (gradients) on a neural network directly using the chain rule similarly to (2), but in neural network training one typically wants to calculate the derivatives of a single output variable w.r.t. a large number of input variables, in which case backpropagation allows a more efficient implementation than using the standard chain rule directly.
Backpropagation ≠ Chain Rule – Theory DishThe most important lesson I draw from this is that metis and a community doing practical work doesn’t put you above academic science and peer-reviewed results (or at least it doesn’t always put you there). The bodybuilders had lots of opportunities to experiment and tinker, with lots of skin in the game, but they were still getting things pretty wrong until researchers looked into some of their conclusions using the normal scientific method.
Metis And Bodybuilders - by Scott AlexanderAt some point in the squeezing process the jittering velocity of the particle approaches the speed of light and its average momentum no longer tracks its mass: to us, it behaves as if its mass (which is actually invariant) were increasing. At some point it is energetic enough that it can pluck another particle out of the vacuum, and then not only do we not know how fast it is moving, but we can no longer say how many particles are present.
quantum mechanics - Does the Heisenberg uncertainty principle only allow location OR momentum to exist? - Physics Stack ExchangeDeepMind recently studied chess algorithms without search and noted that search behavior (looking moves ahead) naturally emerges from such algorithms without external scaffolding. While that’s neat, researchers note that scaling is senseless because chess does have search algorithms. Why wait for inefficient look-ahead to accidentally emerge from large models when existing search algorithms do the trick?
AI Search: The Bitter-er LessonIn the span of a few months, with a small team of researchers and engineers, we trained a 70B parameter model from scratch on our own infrastructure that outperformed zero-shot GPT-4o on reasoning-related tasks.
From bare metal to a 70B model: infrastructure set-up and scripts - imbueI think that no such perturbations exist in general, rather than that we have simply not had any luck finding them.
Solving adversarial attacks in computer vision as a baby version of general AI alignment | Stanislav Forta method that effectively relies on the enumeration of bad behaviors will not be able to scale to realistic scenarios
Solving adversarial attacks in computer vision as a baby version of general AI alignment | Stanislav FortFisher's famous 1921 paper alone has been described as "arguably the most influential article" on mathematical statistics in the twentieth century, and equivalent to "Darwin on evolutionary biology, Gauss on number theory, Kolmogorov on probability, and Adam Smith on economics",[26] and is credited with completely revolutionizing statistics.
Ronald Fisher - WikipediaWithin Notebook X, Kaczynski wrote that while some people may consider him ‘‘sick,’’ he finds that he is a happy man.
Kaczynski's Ciphersequires a 54 42 matrix that Kaczynski created specifically for this system
Kaczynski's CiphersIn school, if you pick an easy problem instead of a hard one, you lose leverage because your extra problem-solving ability goes to waste. But in real life, you can redirect it to prioritizing which problems to solve, or working more quickly, or building a machine that solves the problems for you.
You don’t need to work on hard problemsChurch told two enquirers that the choice was more accidental: a symbol was needed and “𝜆 � ” just happened to be chosen.
ho.history overview - Why did Alonzo Church choose the letter $\lambda$ as the "binding operator"? - MathOverflowMy model shows that it can be estimated that the brain operates at least 10x^21 operations per second. With current rates of growth in computational power we could achieve supercomputers with brain-like capabilities by the year 2037, but estimates after the year 2080 seem more realistic when all evidence is taken into account. This estimate only holds true if we succed to stomp limitations like physical barriers (for example quantum-tunneling), capital costs for semiconductor fabrication plants, and growing electrical costs. At the same time we constantly need to innovate to solve memory bandw…
The Brain vs. Deep Learning vs. Singularitydue to load reasons, I can only start new research collaborations with students with either {IOI top 9/CNOI top 15/EGOI top 5/IMO score at least 30}
cs.cmu.edu/~yangp/Instead, LLM "hallucinations" arise, regularly, because (a) they literally don't know the difference between truth and falsehood, (b) they don't have reliably reasoning processes to guarantee that their inferences are correct and (c) they are incapable of fact-checking their own work. Instead, everything that LLMs say -- true or false -- comes from the same process of statistically reconstructing what words are likely in some context
Humans versus Machines: The Hallucination EditionTake a good hard look at the successful people around you. Doctors too busy to see their children on weekdays. Mathematicians too brilliant in one field to switch to another. Businessmen too wealthy to avoid nightly wining and dining. Professional gamers too specialized to learn a new hero. Public figures too popular to change their minds.
Is Success the Enemy of Freedom? (Full) - LessWrongMaybe there is a cap to effective intelligence, and there’s no such thing as a “500” IQ.
Reasons why AI probably won't kill us all - by Maxim Lott“Finally! An Asian guy who’s good at math!”
Mathematics for Human FlourishingMathematicians and scientists are awed by the simplicity, regularity, and order of the laws of the universe. These are called “beautiful”. They feel transcendent. Why should mathematics be as powerful as it is? This is what Nobel prizewinning physicist Eugene Wigner called “the unreasonable e ff ectiveness of mathematics” to explain the natural sciences. And Einstein asked: “How can it be that mathematics, being after all a product of human thought independent of experience, is so admirably adapted to the objects of reality?
Mathematics for Human FlourishingOur profession is threatened by voices like these from within, and without, who are undermining how society views mathematics and mathematicians. And the view of our profession is dismal. The 2012 report from the President’s Council of Advisors on Science and Technology pegs introductory math courses as the major obstacle keeping students from pursuing STEM majors. We are not educating our students as well as we should, and like most injustices, this hurts those who are most vulnerable
Mathematics for Human FlourishingSimilarly, Github is a very successful home for Software 1.0 code. Is there space for a Software 2.0 Github? In this case repositories are datasets and commits are made up of additions and edits of the labels.
Software 2.0. I sometimes see people refer to neural… | by Andrej Karpathy | Mediumf layers and layers and layers of files of information that he can draw back on now for more than 70 years worth of data.
Curius / Onboarding