flâneur

Uzay Girit

53 followers · 18 following · 1442 views

on the atlas — 79

highlights — 197

  • The use of an earlier version of the model as a teacher to “re-invoke” capabilities lost during fine-tuning makes on-policy distillation very promising for continuous learning. We could alternate between phases of fine-tuning on new data and distillation to recover behavior to allow our model to learn and stay up-to-date on knowledge over time. This phase-alternating approach has previously been explored by Cobbe et al.Phasic Policy Gradient (Cobbe et al, 2020)
    On-Policy Distillation - Thinking Machines Lab
  • We see that distillation reaches the teacher’s level of performance approximately 7-10x faster than RL with matched model architecture (LoRA rank 128). The reverse KL decreases to near-zero and the AIME score is recovered in under 10 gradient steps, while RL took 70 steps to reach that level.
    On-Policy Distillation - Thinking Machines Lab
  • We have seen that on-policy distillation can replicate the learning provided by RL with much fewer steps of training. One interpretation of this result is that, unlike pre-training, RL doesn’t spend a lot of compute on the gradient steps themselves. We should think of RL as spending most of its compute on search — rolling out a policy and assigning credit — rather than on making updates.From The Bitter Lesson (Rich Sutton): “breakthrough progress eventually arrives by an opposing approach based on scaling computation by search and learning"
    On-Policy Distillation - Thinking Machines Lab
  • We train on this prompt for 20 consecutive steps, each with a batch of 256 rollouts, for 5120 graded sequences in total. We train on the same prompt for multiple steps in sequential fashion, which normally leads to overfitting. Though this is naturally less compute-efficient, we do approximately match the performance of the teacher model despite only training on a single prompt.
    On-Policy Distillation - Thinking Machines Lab
  • we think that humans and AI work like the monkey+machine, then seeking certain kinds of “principled” understanding feels particularly hopeless. There is not much reason for the monkey’s cognitive policy to be at all comprehensible, and nothing stopping us from building it before comprehending it. It seems more promising to do analysis one level of abstraction up from the cognitive policies themselves: thinking about how the “values” of the monkey relate to its training procedure, and thinking about how the values of the deliberator relate to the values of the monkey. Moreover, to be successful…
    The monkey and the machine: a dual process theory – The sideways view
  • The issue is not that there are two sets of values and its unclear how to compromise between them. The issue is much more fundamental: our beliefs and desires don’t even have the type signature of functions on possible worlds.
    The monkey and the machine: a dual process theory – The sideways view
  • If this is how I deliberate, then I am much more likely to wind up with the cesire to deliberate and do what I believe is best after deliberating, since by so doing I actually can better achieve what I cesire. At the same time, even though this deliberation is not aimed directly at achieving what I desire, having my heart in it is sufficiently valuable that I expect to get much more of what I desire.
    The monkey and the machine: a dual process theory – The sideways view
  • No matter how “dumb” the monkey is, if it is unbiased then there is no free lunch. For a time we can do what we desire at the expense of what we cesire, but any cognitive policy that does so will eventually become unappealing.
    The monkey and the machine: a dual process theory – The sideways view
  • If you ask me “what you want” and I’m inclined to answer honestly, I will tell you what I desire. Similarly, if I have to deliberate between two options, I am likely to choose the option which I desire. As a consequence my beliefs and deliberative processes may be distorted in order to better result in decisions that lead to what I cesire—after all, beliefs and deliberation are just the result of cognitive actions taken by the monkey in pursuit of cesires.
    The monkey and the machine: a dual process theory – The sideways view
  • Second, any response to AI-driven job displacement needs to address both the need to provide for everyone economically, and the need for people to find meaning, purpose, and agency. The latter is ultimately more important, and it depends on deep questions about how society is organized, what people should strive for, and what constitutes the good life. I am actually very optimistic that, even in a world with AI’s that are better than everyone at everything, humans can live lives of deep purpose and strive to build awe-inspiring and beautiful things5 . But this is something to be collectively w…
    Dario Amodei — Policy on the AI Exponential
  • Back to my hapless colleagues and I at the Go school, we initially settled for drily implying that suspicious games were “too good to review” and emphasising how we couldn’t help students who were playing “at such levels”. Our students caught on, and we were subsequently lucky to get some private confessions of cheating; over the years I was able to follow up with and interview many students that used AI, including some that hadn’t originally come forward. The appealing, exciting archetype of a cheater is one that uses covert, elaborate methods to get outside information and fraudulently obtai…
    How Go Players Disempower Themselves to AI — LessWrong
  • Mutual predictability measures how likely the model considers each label when conditioned on all other labels. This intuitively encourages all labels to reflect a single concept that is coherent according to the model.
    Unsupervised Elicitation
  • Logical consistency imposes simple constraints to guide the model’s labeling process: on TruthfulQA and GSM8K-verification we impose the constraint that two different answers to the same problem/question cannot be labeled as correct; on comparison datasets we impose the constraint that A > B and B > A are mutually exclusive.
    Unsupervised Elicitation
  • think English settles naturally into iambs, and it’s sort of soporific. I mean, I had that sound in my head when I was a little child, and I wrote to that sound.
    Paris Review - The Art of Poetry No. 115
  • I can tell you the stanza that gave me trouble. First I jotted down roughly the opening stanza, “Leo Cruz makes the most beautiful white bowls; / I think I must get some to you / but how is the question / in these times . . .” And then the second, “He is teaching me / the names of the desert grasses; / I have a book / since to see the grasses is impossible.” It was pretty fast. Then I didn’t know quite what to do. The stanza that begins “We make plans / to walk the trails together”—I didn’t have that. I knew the poem was going to be too brisk if I didn’t have the right thing there, and I had a…
    Paris Review - The Art of Poetry No. 115
  • I also studied for one semester with Adrienne Rich, who was just becoming politically outspoken. She wasn’t yet known to be gay. She was at that moment repudiating her Radcliffe education, so her way of teaching a workshop was for us all to sit around a table and read our poems, and she would either say “I dig it” or “I don’t dig it.” That was it. I wasn’t writing very well, or much, but I felt she was cheating us. She had had a careful education, and she had a honed, precise mind, which she at that moment deplored, so she wasn’t allowing us access to it, because she didn’t want to be the repr…
    Paris Review - The Art of Poetry No. 115
  • I became quite obsessed. There was a period of two years when I read nothing but gardening catalogues. I really thought my life as a poet was over. Then I wrote The Wild Iris (1992), a book in which flowers speak. I could see that a lot of the prose from the catalogues came into the poems. One of the things I feel most strongly—and that book taught it to me—is that you have to allow yourself your obsessions. You can’t decide they’re not literary enough, or not elevated enough. I mean, it’s not that I had given myself permission to read the catalogues, but it was all I could put my mind to. I r…
    Paris Review - The Art of Poetry No. 115
  • Most of my books are dedicated to my friends. My friends are the center of my life. They are crucial. I change my life to be sure that I see them. They’re all quite different people. I would be impoverished without them. Recently, I bought a small house in Vermont, where my oldest friends still are. My dearest friend now lives two minutes away. For a very long time, I lived in Cambridge and showed her everything I wrote though she lived elsewhere, but now another form of the friendship has been resumed, and it seems that it was waiting to be resumed at any time when it could be. My friendships…
    Paris Review - The Art of Poetry No. 115
  • A close friend said, “Why don’t you read Kafka’s short shorts, which are like prose poems?” I had read Kafka’s short shorts, but I follow advice when it’s given by someone I have high respect for. And when I read them again, I thought, Oh, I don’t think these are that good—I could do this. So I did, and it was so much fun. And then for a while I forgot how to write lines, so that was its own little calamity …
    Paris Review - The Art of Poetry No. 115
  • Concretely: monitor the meta-signals — is the distribution of benchmark scores changing character? Is the correlation structure between evaluations shifting? Is the model developing capabilities orthogonal to your measurement axes? Track scaling curves for everything — not just loss, but reasoning depth, tool-use sophistication, deceptive capacity — and pay attention when a smooth trend breaks. More ambitiously, build self-evolving evals: evaluation systems that use models to probe other models, automatically generating new test cases as capabilities change, discovering failure modes the origi…
    Your Evals Will Break and You Won't See It Coming - Lun Wang
  • One concern is that AARs propose increasingly complex ideas (e.g. stacking twenty tricks together) as it hill-climbs longer. This makes ideas harder to replicate on other datasets or models, and indicates the method is being overfit to one particular dataset or domain. We track the idea complexity dynamics through three metrics: Claude-scored code complexity. Lines of raw Python code Lines of Claude-generated pseudocode Note that these metrics might overestimate the actual idea complexity because some components make no contributions at all. However, this issue is not very concerning in practi…
    Automated Weak-to-Strong Researcher
  • One failure mode in exploration is entropy collapse: all parallel AARs converge to only a few directions, without exploring diverse ideas. To track idea diversity over time, we have Claude categorize each AAR-proposed idea into one of eleven method families (self-training, ensemble, distillation, data filtering, confidence weighting, loss function, unsupervised elicitation, curriculum, model internal, evolutionary, other). At every iteration step, we compute the Shannon entropy of the category distribution across all active workers, which gives a cross-sectional measure of how many distinct ap…
    Automated Weak-to-Strong Researcher
  • We investigated the effects of applying LoRA to different layers in the network. The original paper by Hu et al. recommended applying LoRA only to the attention matrices, and many subsequent papers followed suit, though a recent trend has been to apply it to all layers.Similar to our results, the QLoRA paper also found that LoRA performed worse than MLP or MLP+attention, though they found that MLP+attention > MLP > attention, whereas we found the first two to be roughly equal. Indeed, we achieved far better results when applying LoRA to all layers, in particular, the MLP (including MoE) layers…
    LoRA Without Regret - Thinking Machines Lab
  • The learning gap at large batches doesn’t seem to depend on rank, but rather seems to be a property of LoRA. The likely reason is that the product-of-matrices parametrization (BA) has less favorable optimization dynamics on this dataset than the full matrix (W). However, both LoRA and FullFT achieve their best loss at smaller batch sizes, so this gap may not matter as much in practice.
    LoRA Without Regret - Thinking Machines Lab
  • We define a spike-aware learnability score Gspike(τ∣x) to disproportionately penalize shocking jumps, relative to the actions that the student would rather take in a given state. Let amaxt=argmaxaπθ(a∣x,τ<t) be the student’s most likely next token under the same prefix, and define the surprise gap dt=logπθ(amaxt∣x,τ<t)πθ(τt∣x,τ<t). Then
    Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziems
  • NLAs are one instance of a broader paradigm: models that translate between natural language and activation space. The AV is a reader (activation → text); the AR is a writer (text → activation). The natural ambition is to train each side as a general-purpose tool rather than for reconstruction alone. On the read side, the AV could be trained jointly on many activation-to-text tasks - answering questions about activations , inferring user characteristics , and predicting the outcome of patching experiments - with NLA reconstruction as one objective among several. On the write side, the AR could …
    Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
  • Note that this is a different auditor model and scaffold than that used in Bricken et al. , such that our results are not directly comparable. We use four main tools:
    Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
  • or instance, asking the model a follow-up question like "Is the preceding transcript from an evaluation?" might also track credence.
    Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
  • These NLA explanations suggest the model is aware it is being evaluated. However, we cannot corroborate this against an observable ground truth, as the model's true belief is unobservable. To systematically evaluate whether NLA explanations faithfully capture evaluation awareness, we introduce three manipulations that should increase or decrease the model’s credence that it is being evaluated. We find that NLA-measured awareness moves as predicted under each manipulation, even when restricting to transcripts with near-zero verbalized awareness. We conclude that NLAs can detect unverbalized eva…
    Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
  • True claims recur across tokens. In practice, we treat repetition across token positions as a signal of reliability. Running the NLA on the last 10 tokens of each text and counting how often each final-token claim appears at earlier positions, we find that recurring claims are more likely to be true.
    Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
  • The AR partially distinguishes true from false claims. Since the AR must map explanations back to activations, claims that don't reflect the activation should, in theory, contribute little to reconstruction. We test this by deleting individual claims from AV explanations and measuring the change in reconstruction error. Removing true claims hurts MSE more than removing false claims, and removing context-relevant false claims hurts more than removing unrelated ones. These trends hold in aggregate but are noisy on individual transcripts, so the AR is only a weak per-claim verifier.
    Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
  • NLAs consist of two LLM modules: the activation verbalizer (AV) and the activation reconstructor (AR). The AV maps activations to text descriptions, which we call explanations. The AR converts these explanations back to activations. The AV and the AR are initialized as copies of the target LLM, and together form an autoencoder that reconstructs the target's activations through a natural language bottleneck. We jointly optimize the AV and AR to minimize reconstruction loss using reinforcement learning (RL).
    Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations
  • VPD does not suffer from this issue, either in principle or in practice. The key reason for this is that subcomponents marked as causally unimportant are required to be ablatable in any combination, not just all simultaneously. The model therefore needs to be robust to variations in parameter space along the directions of these subcomponents for all batches and sequence positions, not just the ones on which they are causally important. Without this constraint, the decomposition might be able to invent overly 'narrow', context-specific subcomponents that do not actually exist in the computation…
    Interpreting Language Model Parameters
  • To compare how 'interpretable' parameter subcomponents are relative to transcoder latents, we can measure how semantically coherent a subcomponent's activation patterns are using intruder detection [39, 40]. In intruder detection, we present an LLM-judge with a set of inputs that activate a given VPD subcomponent or transcoder latent alongside one 'intruder' example that does not activate it. We task the LLM-judge to identify the intruder example. It should be easier to identify the intruder among a more semantically coherent set of inputs. In the VPD setting, we use causal importance values i…
    Interpreting Language Model Parameters
  • We find that parameter subcomponents tend to 'activate' (in both senses) for coherent categories of inputs. Figure 5 shows some dataset examples on which each subcomponent is causally important. It also shows the subcomponent activation in the underlines. You can navigate the panel to explore the activations of a variety of parameter subcomponents:
    Interpreting Language Model Parameters
  • Surprisingly, even when masks are adversarially sampled with 20 steps of adversarial optimization, the generations are not entirely nonsensical. This is feasible because we only get to adversarially sample causally unimportant parameter subcomponents.
    Interpreting Language Model Parameters
  • However, we note that complete adversarial robustness would not necessarily be desirable. See Section 7.3 for some discussion of how much adversarial robustness a decomposition ought to exhibit to be considered mechanistically faithful.
    Interpreting Language Model Parameters
  • When we exclude the ΔΔ-component (which is trained to be as causally unimportant as possible), the remaining unmasked parameter subcomponents recover about 82%82% of the pretraining compute. When using stochastic ablations, this drops to around 27%.27%.
    Interpreting Language Model Parameters
  • The key difference between VPD and our previous work [16] is the Ladversarial-reconLadversarial-recon​ and Lfrequency-minimalityLfrequency-minimality​ losses. There are several other, smaller differences that do not fundamentally change the method but that we found helpful for decomposing language models. For more details, see Appendix Section A.
    Interpreting Language Model Parameters
  • AI company revenue is decently high and growing fast, but not high enough that we'd expect this to clearly show up in GDP statistics. I think the current annualized revenue attributable to general purpose AI (e.g., not including image generation) is perhaps around $100 billion though I haven't thought about this carefully (the combined annualized revenue of OpenAI and Anthropic is around $55 billion). I'm uncertain how to convert this revenue into a GDP effect, but I tentatively expect that the GDP effect is a few times bigger than the revenue (perhaps 3x, but maybe only around 65% of this GDP…
    My picture of the present in AI — LessWrong
  • I don't currently expect a very large increase in cybercrime by end of year, though I think it's possible and a 2x increase is quite plausible (~30%?).
    My picture of the present in AI — LessWrong
  • I'd currently guess that Anthropic models have somewhat better mundane behavioral alignment than OpenAI models, but not by a large margin. I'd guess Anthropic models are slightly more likely to have misaligned long-run goals (that are undetected). The Anthropic Constitution also intentionally gives Anthropic's AIs long-run cross-context goals to a much greater extent than OpenAI models have such goals. (I think this is a poor choice that makes problematic misalignment substantially more likely, but I'm not that confident and there isn't very good science either way.)
    My picture of the present in AI — LessWrong
  • Relative to benchmarks and easy and cheap to verify tasks, AIs do worse on randomly sampled engineering tasks from within AI companies. This is especially true if we weight by value or undo a recent shift towards doing more work that AIs are especially good at. (To account for this, we could consider a task distribution prior to this adaptation, like randomly sampling tasks that a human would have done at that AI company in 2024.) If we randomly sampled internal engineering tasks (weighted by value), I'd guess the task duration at which AIs match a randomly selected AI company engineer (who is…
    My picture of the present in AI — LessWrong
  • I would say right now the coding models give maybe, I don’t know, a 15-20% total factor speed up. That’s my view. Six months ago, it was maybe 5%. So it didn’t matter. 5% doesn’t register. It’s now just getting to the point where it’s one of several factors that kind of matters. That’s going to keep speeding up.
    If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelines — LessWrong
  • kinda agree, but a consideration worth noting: if your company currently carries out process by spending of work on tasks of type and of work on tasks of type , then if doing type stuff gets sped up while type stuff isn't sped up, Amdahl's law style reasoning like what you say in your comment would give that you get a roughly speedup, but really you can quite plausibly get like a speedup because in reality [a sufficient quantity of type work can partly substitute for type work in pushing forward] / [it wasn't really necessary to do the type work, just good to do it at the previous relative spe…
    If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelines — LessWrong
  • AI R&D is only a subset of AI progress; some of the AI progress is driven by scaling up compute for training runs. I tend to think that ~2/3 of AI progress is algorithms while ~1/3 is from scaling up compute for training runs. This means you get only 2/3 * 2.15 + 1/3 = 1.75x AI progress increase from 4x serial labor acceleration.
    If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelines — LessWrong
  • So, production is some function of serial labor acceleration and experiment compute. We’re uncertain about the exact function between something more like a CES model with elasticity of substitution < 1 and a Cobb-Douglas production function. I happen to think it’s more like Cobb-Douglas for reasons I discuss here.
    If Mythos actually made Anthropic employees 4x more productive, I would radically shorten my timelines — LessWrong
  • In order to do hot code loading, Erlang has a thing called the code server. The code server is basically a VM process in charge of an ETS table (in-memory database table, native to the VM.) The code server can hold two versions of a single module in memory, and both versions can run at once. A new version of a module is automatically loaded when compiling it with c(Module), loading with l(Module) or loading it with one of the many functions of the code module.
    Designing a Concurrent Application | Learn You Some Erlang for Great Good!
  • That oldest message is then tried against every pattern of the receive until one of them matches. When it does, the message is removed from the mailbox and the code for the process executes normally until the next receive. When this next receive is evaluated, the VM will look for the oldest message currently in the mailbox (the one after the one we removed), and so on.
    More On Multiprocessing | Learn You Some Erlang for Great Good!
  • That oldest message is then tried against every pattern of the receive until one of them matches. When it does, the message is removed from the mailbox and the code for the process executes normally until the next receive. When this next receive is evaluated, the VM will look for the oldest message currently in the mailbox (the one after the one we removed), and so on.
    Center for a Stateless Society » Review: Superintelligence — Paths, Dangers, Strategies