Anastasia Wei
11 followers · 12 following · 267 views
on the atlas — 50
- Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident - METR20 savers
- Getting a better sense for when you’re thinking well and when you’re faking it2 savers
- What is the Alignment Community Thinking? — LessWrong1 savers
- RL creates split personas — LessWrong1 savers
- project_ideas - Google Docs3 savers
- Reinforcement learning towards broadly and persistently beneficial models6 savers
- On having more interesting ideas - by Henrik Karlsson8 savers
- Dostoevsky as lover - by Henrik Karlsson14 savers
- Sometimes the reason you can’t find people you resonate with is because you misread the ones you meet32 savers
- Pick an audience that is into the illegible you of today, not your past achievements21 savers
- Understand, align, cooperate: AI welfare and AI safety are allies1 savers
- Laying Some Cause-Prioritization Groundwork for Digital Minds — EA Forum1 savers
- Augmenting Long-term Memory30 savers
- Contra Pritchard On Liberal Happiness - by Scott Alexander1 savers
- Why you shouldn't build your career around existential risk - Alexey Guzey5 savers
- Scraping training data for your mind - by Henrik Karlsson14 savers
- First we shape our social graph; then it shapes us22 savers
- Think more about what to focus on - by Henrik Karlsson41 savers
- SFT Drives Gemini’s Safety Properties — LessWrong5 savers
- [REPOST] Epistemic Learned Helplessness | Slate Star Codex1 savers
- Risk from fitness-seeking AIs: mechanisms and mitigations — LessWrong3 savers
- Should you marry her? - by Ajeya Cotra - Good Bones3 savers
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?1 savers
- Politics and the English Language | The Orwell Foundation23 savers
- A global workspace in language models \ Anthropic25 savers
- Why Conscious AI Is a Bad, Bad Idea - Nautilus1 savers
- Experts Who Say That AI Welfare is a Serious Near-term Possibility | Eleos AI1 savers
- [2505.17120] Self-Interpretability: LLMs Can Describe Complex Internal Processes that Drive Their Decisions1 savers
- Childhoods of exceptional people - by Henrik Karlsson69 savers
- Reasoning Transparency | Coefficient Giving5 savers
- Design and Research2 savers
- Everything that turned out well in my life followed the same design process64 savers
- Relationships are coevolutionary loops - by Henrik Karlsson39 savers
- Good conversations have lots of doorknobs34 savers
- Looking for Alice - by Henrik Karlsson - Escaping Flatland90 savers
- The Persona Selection Model: Why AI Assistants might Behave like Humans34 savers
- Teaching Claude Why9 savers
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversion4 savers
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations8 savers
- Natural Language Autoencoders Produce Unsupervised Explanations of LLM Activations — LessWrong2 savers
- What if the AI line just keeps going up? - by Jesse2 savers
- How useful is the information you get from working inside an AI company?4 savers
- Fuzzing LLMs sometimes makes them reveal their secrets — AI Alignment Forum1 savers
- The case for ensuring that powerful AIs are controlled — LessWrong11 savers
- Ten people on the inside — LessWrong2 savers
- Eliciting latent knowledge. How can we train an AI to honestly tell… | by Paul Christiano | AI Alignment2 savers
- The Hitchhiker's Guide to Actionable Interpretability6 savers
- Putting up Bumpers2 savers
- Anthropic repeatedly accidentally trained against the CoT, demonstrating inadequate processes1 savers
- Curius / Onboarding2621 savers
highlights — 68
In a large U.S. sample, the shape of the association between happiness and Log(income) was extremely systematic: from $10,000/y to over $500,000/y, average happiness rose almost perfectly linearly with Log(income), with group-level correlations of 0.98-0.99 across a range of happiness measures, including both in-the-moment experience and overall life satisfaction
Contra Pritchard On Liberal Happiness - by Scott AlexanderWhat we end up with is a process something like this. You reflect, as best you can, on what you are trying to do. Then you try to locate the prime achievements in this field, and you observe them in as high a resolution as possible, with all the messy details that come from studying something in its living context. (You don’t have to find the global maximum at your first go, as long as you can figure out where the gradient is. By going after the peak from your current vantage point, you can travel up toward the top by looking at what your “peak” is looking at, which is likely further up the hi…
Scraping training data for your mind - by Henrik KarlssonThe knowledge is tacit. The explanations are post-hoc rationalizations; they do not produce the results.
Scraping training data for your mind - by Henrik KarlssonThis principle—looking for people who can guide you to the peaks of the domain—is very general.
Scraping training data for your mind - by Henrik KarlssonBut to get good at something—parenting, writing code, doing research—you also need to internalize examples of prime achievements in that field. Knowing how to find these examples is upstream of the tasks you need to master.
Scraping training data for your mind - by Henrik KarlssonWhat you want to create is a distributed apprenticeship in the art of being you.
First we shape our social graph; then it shapes usGrit is the node’s eye’s view. You are struggling against the graph. Agency, on the other hand, is the view from the graph. You are the graph. By changing it, you are changing yourself.
First we shape our social graph; then it shapes usIt is by changing your milieu that you change yourself.
First we shape our social graph; then it shapes usThe instinct was to curate a culture, not to teach, not primarily.
First we shape our social graph; then it shapes usThat is perhaps the most solid dating advice I have, by the way—show the inside of your head in public, so people can see if they would like to live in there.
Think more about what to focus on - by Henrik KarlssonThe trick is to collide your mental model with the outside world as often as possible
Think more about what to focus on - by Henrik KarlssonMost, however, tend to do less than optimal of both—not exploring, not exploiting; but doing things out of blind habit, and half-heartedly.
Think more about what to focus on - by Henrik KarlssonWe were leaving something, more than going somewhere.
Think more about what to focus on - by Henrik Karlssonmost safety relevant properties in Gemini seem to be caused by the combination of pretraining and SFT, not other training stages like RL.
SFT Drives Gemini’s Safety Properties — LessWrongIf Osama comes up to him with a really good argument for terrorism, he thinks “Oh, there’s a good argument for terrorism. I guess I should become a terrorist,” as opposed to “Arguments? You can prove anything with arguments. I’ll just stay right here and not blow myself up.”
[REPOST] Epistemic Learned Helplessness | Slate Star CodexBut I think they might instead co-occur because you have to be really smart in order for taking ideas seriously not to be immediately disastrous. You have to be really smart not to have been talked into enough terrible arguments to develop epistemic learned helplessness.
[REPOST] Epistemic Learned Helplessness | Slate Star CodexYou could consider this a form of epistemic learned helplessness, where I know any attempt to evaluate the arguments is just going to be a bad idea so I don’t even try.
[REPOST] Epistemic Learned Helplessness | Slate Star CodexA rough model is that you learn one more unit of information about your compatibility with your partner from every doubling of the amount of time you spend with them.
Should you marry her? - by Ajeya Cotra - Good BonesI have not here been considering the literary use of language, but merely language as an instrument for expressing and not for concealing or preventing thought.
Politics and the English Language | The Orwell FoundationThe appropriate noises are coming out of his larynx, but his brain is not involved as it would be if he were choosing his words for himself. If the speech he is making is one that he is accustomed to make over and over again, he may be almost unconscious of what he is saying, as one is when one utters the responses in church. And this reduced state of consciousness, if not indispensable, is at any rate favourable to political conformity.
Politics and the English Language | The Orwell FoundationA scrupulous writer, in every sentence that he writes, will ask himself at least four questions, thus: What am I trying to say? What words will express it? What image or idiom will make it clearer? Is this image fresh enough to have an effect? And he will probably ask himself two more: Could I put it more shortly? Have I said anything that is avoidably ugly?
Politics and the English Language | The Orwell FoundationThe first is staleness of imagery; the other is lack of precision. The writer either has a meaning and cannot express it, or he inadvertently says something else, or he is almost indifferent as to whether his words mean anything or not.
Politics and the English Language | The Orwell FoundationA man may take to drink because he feels himself to be a failure, and then fail all the more completely because he drinks. It is rather the same thing that is happening to the English language. It becomes ugly and inaccurate because our thoughts are foolish, but the slovenliness of our language makes it easier for us to have foolish thoughts.
Politics and the English Language | The Orwell FoundationFor instance, the J-space is constructed by identifying representations of potential outputs—words the model might say. If something similar holds in humans, it would suggest that the global workspace might be fundamentally tied to brain regions that prepare actions and speech, more so than to sensory areas
A global workspace in language models \ AnthropicIt remains a contested philosophical question whether or not access consciousness implies phenomenal consciousness, or if the ability to have experiences requires some other property.
A global workspace in language models \ AnthropicHowever, during post-training, the J-space develops some signatures of adopting “Claude’s point of view.”
A global workspace in language models \ AnthropicIf a thought is consciously accessible to you, you can typically describe it if someone asks. We went looking for representations in Claude with the same property: representations that are positioned to influence what Claude might say—not necessarily what it’s saying right now, but what it could talk about, if asked.
A global workspace in language models \ AnthropicAI systems that pass the Garland test will subject us to a kind of cognitive illusion, much like simple visual illusions in which we cannot help seeing things in a particular way, even though we know the reality is different.
Why Conscious AI Is a Bad, Bad Idea - NautilusExistential concerns aside, there are more immediate dangers to deal with as AI has become more humanlike in its behavior. These arise when AI systems give humans the unavoidable impression that they are conscious, whatever might be going on under the hood.
Why Conscious AI Is a Bad, Bad Idea - NautilusThe first is that it may endow AI systems with new powers and capabilities that could wreak havoc if not properly designed and regulated. Ensuring that AI systems act in ways compatible with well-specified human values is hard enough as things are. With conscious AI, it gets a lot more challenging, since these systems will have their own interests rather than just the interests humans give them.
Why Conscious AI Is a Bad, Bad Idea - NautilusMy own view is that consciousness is intimately tied to our nature as living flesh-and-blood creatures. In this picture, being conscious is not the result of some complicated algorithm running on the wetware of the brain. It is an embodied phenomenon, rooted in the fundamental biological drive within living organisms to keep on living. If I’m right, the prospect of conscious AI remains reassuringly remote.
Why Conscious AI Is a Bad, Bad Idea - NautilusRussell commented that the development of such gifted individuals required a childhood period in which there was little or no pressure for conformity, a time in which the child could develop and pursue his or her own interests no matter how unusual or bizarre.
Childhoods of exceptional people - by Henrik KarlssonFor example, many scientific papers say little about roughly how confident the authors are in different claims throughout the paper
Reasoning Transparency | Coefficient GivingThe difference between design and research seems to be a question of new versus good. Design doesn't have to be new, but it has to be good. Research doesn't have to be good, but it has to be new. I think these two paths converge at the top: the best design surpasses its predecessors by using new ideas, and the best research solves problems that are not only new, but actually worth solving
Design and ResearchIt is a feedback loop between you and the context. By gradually adjusting the thing you are designing and observing how well it fits the context, you create a feedback loop that embeds the context’s knowledge into your design. Your design ends up smarter than you.
Everything that turned out well in my life followed the same design processThe bug log was, I realize now, a way of making the coevolutionary loop explicit. Writing down everything that went wrong helped us pinpoint the bottlenecks, the frictions. It also forced us to put words to our goals, values, and assumptions, opening those up for discussion and refinement. Is that really what you want? What are you willing to give up to achieve it? How do you feel about it?
Relationships are coevolutionary loops - by Henrik KarlssonYou need to keep a certain rate of improvement for things not to break. If you are too slow in adapting to each other, you grow apart. If you raise your iteration speed, you can go places you wouldn’t have guessed possible a priori.
Relationships are coevolutionary loops - by Henrik KarlssonWhat matters most, then, is not how much we give or take, but whether we offer and accept affordances
Good conversations have lots of doorknobsThere is really no point in going to a café to talk safely (if you can avoid it). You want to rapidly extract as much information as possible, so you can figure out what you like and so that you can pattern match, and you want to communicate as much as possible, too, so you can filter people who wouldn’t fit you anyway (which is why keeping a blog is good). The type of person I’m assuming we’re looking for here is 1) someone that you will find fascinating to talk to after you’ve talked for 20,000 hours, 2) you feel comfortable with them talking through the hardest and most painful decisions yo…
Looking for Alice - by Henrik Karlsson - Escaping FlatlandYou like individuals. And you’re not born knowing which kind.
Looking for Alice - by Henrik Karlsson - Escaping FlatlandAn important open question is how exhaustive PSM is, especially whether there might be sources of agency external to the Assistant persona, and how this might change in the future.
The Persona Selection Model: Why AI Assistants might Behave like HumansHowever, despite only SFTing on a relatively small amount of generic chat data from our production RL mix after the SDF step, we see that the model’s performance on the factual recall evaluation is highly correlated with its performance on open ended questions and the blackmail evaluation, indicating that the model is not simply memorizing the text but somewhat internalizing the content.
Teaching Claude WhyWe believe that this happens because the model already learns lots of true information through pre-training on similar documents, so it is accustomed to incorporating information in this format into its knowledge base. On the other hand, the model mostly learns how to best use that information through chat data. As a result, being trained on this chat-formatted data could teach the model to believe that the information is true. However, it could also teach the model that the assistant character has a habit of fabricating information when asked questions on this topic. Stated simply, pretrainin…
Teaching Claude WhyFrom there, we experimented with improving response quality by injecting additional instructions
Teaching Claude WhyImprovements that we make to the PT prior and alignment-specific data do not degrade during RL post-training. In fact, alignment continues to improve. This (and the fact that alignment improves in the baseline run) indicates to us that hypotheses 1 and 2 above are not the source of the issue.
Teaching Claude WhyThis training process does not explicitly incentivize NLA explanations to be interpretable or faithful.
Natural Language Autoencoders Produce Unsupervised Explanations of LLM ActivationsOne thing that’s really important to do here is to decouple your views of whether AI is good or bad, from your views around how quickly AI capabilities are going to progress.
What if the AI line just keeps going up? - by Jessewe need to be able to align these automated researchers.
Putting up BumpersEstimating when we have reached this point is not trivial, but ordinary capability assessments, red-teaming exercises, and warning signs from earlier model generations will all give us substantial evidence with which to make this determination.
Putting up BumpersSimilarly, with systems that are substantially superhuman at reasoning about AI safety, it becomes more plausible that they discover exotic ways of evading bumpers that no human adversarial evaluation could have caught.
Putting up Bumpers