flâneur

Evie Hu

18 followers · 15 following · 533 views

on the atlas — 135

highlights — 21

  • We find that Sonnet 3.5 & 4, the models that best maintain accuracy in ciphered reasoning, can reason in well-known ciphers like rot13 and base64 with 45% accuracy drops.
    [2510.09714] All Code, No Thought: Current Language Models Struggle to Reason in Ciphered Language
  • Topological properties of data, such as links, may make it impossible to linearly separate classes using low-dimensional networks, regardless of depth. Even in cases where it is technically possible, such as spirals, it can be very challenging to do so.
    Neural Networks, Manifolds, and Topology -- colah's blog
  • When the problem’s conditioning is poor, the optimal 𝛼 α is approximately twice that of gradient descent, and the momentum term is close to 1 1. So set 𝛽 β as close to 1 1 as you can, and then find the highest 𝛼 α which still converges.
    Why Momentum Really Works
  • the “pathological directions” — the eigenspaces which converge the slowest — are also those which are most sensitive to noise!
    Why Momentum Really Works
  • Both GRAM and LoRA isolate capabilities from real-world dual use data. We train an 800M-parameter language model on a combination of general text, code, and scientific papers. We additionally train on data from four dual use domains: virology, cybersecurity, nuclear physics, and specialized code.
    Modular Pretraining Enables Access Control
  • The claim here is that 1bit learning is extremely sample inefficient: there’s no learning happening for partial success, and with long chains of actions, the probability of sampling a success can shrink quickly towards 0 (if this sounds familiar to you, some eminent guy in Deep Learning makes this claim often…).
    Just make the straw bigger
  • Too often, we quickly jump ahead to the new thing, failing to get good enough at the important thing.
    First, make rice | Seth's Blog
  • within LLMs’ repertoire of vector representations, is there a privileged subset that plays a computational role analogous to the global workspace?
    Verbalizable Representations Form a Global Workspace in Language Models
  • This has a close connection to the exploration-exploitation trade-off: increasing entropy results in more exploration, which can accelerate learning later on. It can also prevent the policy from prematurely converging to a bad local optimum.
    Soft Actor-Critic — Spinning Up documentation
  • We can make meaningful progress on this now because we have systems that implement values well enough to study and test for how well they implement our intent. This is a fundamental change, and understanding it is a prerequisite to our future progress and understanding our risks.
    Alignment Is Proven To Be Solvable - by SE Gyges
  • f we fine-tune models on the same data, with the same parameters, but with a different random seed, how much variance do we see in harmful-
    [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistency
  • Instead of being a perfectionist about the paper, be a perfectionist about writing the paper. Be a perfectionist about identifying good strategies, about abandoning sunk costs, about killing your darlings, about noticing when you're done. Be a perfectionist about wasting no attention. Be a perfectionist about learning from your mistakes. Perfectionism can be a powerful tool, but there's no need to point it at overachieving on metrics you don't care about.
    Half-assing it with everything you've got
  • I think it gets in the way of — we’re leading China. We’re leading everybody, and I don’t want to do anything that’s going to get in the way of that.
    ‘I didn't like certain aspects’: Trump postpones AI executive order - POLITICO
  • The best I can do is to stammer that we philosophy professors are people who have a certain familiarity with a certain intellectual tradition, as chemists have a certain familiarity with what happens when you mix various substances together.
    Rorty-Wild Orchids
  • So, at 12, I knew that the point of being human was to spend one's life fighting social injustice.
    Rorty-Wild Orchids
  • response is considered to belong to the evaluated category if it scores greater than 50
    [2506.11613] Model Organisms for Emergent Misalignment
  • Overly abstract thinking involves relying on g eneralized schemas that are devoid of contextual cues. This can lead decision-makers to apply poorly fi tting mental models, misjudge threats or opportunities as more distant than they are, or assume that others will behave in stereo typical ways. Conversely, overly concrete thinking involves being deeply immersed in the minute details of a speci fi c situation. Such detail-oriented thinking may cause decision-makers to mistakenly miss the bigger picture by overlooking broader trends that unfold over time and multiple locations, leading to choices…
    Abstractness, Concreteness, and Strategic Surprises
  • However, this is sufficiently minor that it does not compromise the emergent nature of the phenom- ena.
    [2506.11613] Model Organisms for Emergent Misalignment
  • The truth is, we can’t do without ad hominem reasoning, for the simple reason that human knowledge is deeply social. Almost everything we know comes from testimony; only an infinitesimal fraction do we verify ourselves. The rest is, literally, hearsay. No wonder we are so sensitive to the reputation and trustworthiness of our sources.
    The Fallacy Fallacy - by Maarten Boudry - Persuasion
  • zine, and he read articles on how to stock a meat department... What he’s really done is he’s created this immense vertical filing cabinet in his brai
    Curius / Onboarding
  • spapers, biographies, trade press. He went over to his grandfather who was a grocer and he re
    Curius / Onboarding