Arjun Khandelwal
20 followers · 23 following · 1597 views
on the atlas — 107
- Pyramid Replacement - The Intelligence Curse9 savers
- The future of alignment if LLMs are a bubble — LessWrong1 savers
- Thoughts on the conservative assumptions in AI control3 savers
- How can we solve diffuse threats like research sabotage with AI control?1 savers
- Win/continue/lose scenarios and execute/replace/audit protocols1 savers
- Test your interpretability techniques by de-censoring Chinese models — LessWrong5 savers
- How To Become A Mechanistic Interpretability Researcher — AI Alignment Forum9 savers
- Does Reality Drive Straight Lines On Graphs, Or Do Straight Lines On Graphs Drive Reality? | Slate Star Codex3 savers
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Data6 savers
- Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases5 savers
- Narrow finetuning is different — LessWrong3 savers
- Mode collapse3 savers
- Conditioning, Prompts, and Fine-Tuning — LessWrong1 savers
- agniv.me/pasts.html2 savers
- Basics of How Not to Die — LessWrong2 savers
- Types of financial risks (video) | Insurance | Khan Academy1 savers
- Charitable giving (video) | Financial goals | Khan Academy1 savers
- Choosing a credit card: credit card types (article) | Khan Academy1 savers
- What is a credit card? (article) | Khan Academy2 savers
- What is a credit report? (article) | Khan Academy2 savers
- How do I raise my credit score? (article) | Khan Academy1 savers
- What can change your credit score? (video) | Khan Academy1 savers
- What is a credit score? (article) | Khan Academy2 savers
- How do you balance your budget? (article) | Khan Academy2 savers
- What is a budget? (article) | Budgeting | Khan Academy1 savers
- Watch team backup • Otherwise3 savers
- Modifying LLM Beliefs with Synthetic Document Finetuning7 savers
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Models5 savers
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWrong3 savers
- My path to OpenAI16 savers
- The making of OpenAI - Greg Brockman8 savers
- How I became a machine learning practitioner17 savers
- "Greg Brockman works 60 to 100 hours per week, and spends around 80% of the time... | Hacker News1 savers
- On Owning Galaxies — LessWrong1 savers
- Turning 20 in the probable pre-apocalypse — LessWrong8 savers
- How We Use Claude Code Skills to Run 1,000+ ML Experiments a Day1 savers
- The behavioral selection model for predicting AI motivations — LessWrong15 savers
- The inaugural Redwood Research podcast - by Buck Shlegeris2 savers
- taboo1 savers
- 'How to be a Human' Starter Pack - by Lydia Nottingham6 savers
- Many can write faster asm than the compiler, yet don't. Why? — LessWrong2 savers
- My template for a quarterly review + plan1 savers
- Alignment remains a hard, unsolved problem — LessWrong15 savers
- Researching from the Heart 💗 - by mrinank2 savers
- Curated AI Career Advice - by Saheb Gulati3 savers
- Curius / Onboarding2621 savers
- How To Be Successful96 savers
- You and Your Research77 savers
- AI 202755 savers
- What I Wish Someone Had Told Me - Sam Altman52 savers
- You don’t need to work on hard problems51 savers
- LessWrong39 savers
- https://cs.stanford.edu/~jsteinhardt/ResearchasaStochasticDecisionProcess.html31 savers
- Do I Need to Go to University? -- colah's blog25 savers
- Visual Information Theory -- colah's blog24 savers
- Advice on Upskilling - Justin Skycak24 savers
- Research Taste Exercises [rough note] -- colah's blog22 savers
- Lil'Log22 savers
- Should You Reverse Any Advice You Hear? | Slate Star Codex21 savers
- Galaxy brain resistance18 savers
- AGI Ruin: A List of Lethalities - LessWrong17 savers
- Building the heap: racking 30 petabytes of hard drives for pretraining | blog17 savers
- Defining the Intelligence Curse - The Intelligence Curse13 savers
- Gears in understanding - LessWrong12 savers
- It Is Your Responsibility to Follow Up - Alexey Guzey11 savers
- Thinking Machines Lab11 savers
- 2025 year in review | Kevin Liu10 savers
- The Law of Leaky Abstractions – Joel on Software10 savers
- Shipping at Inference-Speed | Peter Steinberger9 savers
- Spaced Repetition Systems Have Gotten Way Better | Domenic Denicola8 savers
- The Extreme Inefficiency of RL for Frontier Models — Toby Ord7 savers
- The Roots of Progress7 savers
- Where I agree and disagree with Eliezer — LessWrong7 savers
- reactive agency - by Claire Wang - cold brew blog7 savers
- Andrej Karpathy — AGI is still a decade away7 savers
- Life after work | Mechanize Inc.7 savers
- In defense of blub studies | benkuhn.net6 savers
- Plans A, B, C, and D for misalignment risk — LessWrong6 savers
- Why are there so few independent eval startups? | Thomas I. Liao6 savers
- On love & relationships | Evan Conrad6 savers
- The upcoming GPT-3 moment for RL | Mechanize Inc.6 savers
- The Case Against AI Control Research — LessWrong6 savers
- How confessions can keep language models honest | OpenAI6 savers
- corner.inc6 savers
- Kalman filter6 savers
- What failure looks like - AI Alignment Forum5 savers
- Workshop Labs PBC5 savers
- X explains Z% of the variance in Y — LessWrong5 savers
- Sweatshop data is over | Mechanize Inc.5 savers
- Matthew effect5 savers
- Applause Lights — LessWrong4 savers
- A basic systems architecture for AI agents that do autonomous research — LessWrong4 savers
- Post-irony - Wikipedia4 savers
- The Memetics of AI Successionism — LessWrong4 savers
- AI Futures Model: Dec 2025 Update3 savers
- Persistent Path-Dependence: Why Our Actions Matter Long-Term3 savers
- The Problem — LessWrong3 savers
- How quick and big would a software intelligence explosion be?3 savers
- Power law - Wikipedia3 savers
- The Level Above Mine — LessWrong3 savers
highlights — 310
1 to 3,000 before giving the final answer (filler tokens which might help by providing the LLM more space to think within a forward pass). We find that both perform poorly,
Do reasoning models use their scratchpad like we do? Evidence from distilling paraphrases1. we are currently working on this and seeing some interesting results :)
Test your interpretability techniques by de-censoring Chinese models — LessWrongavoid some of the pitfalls of model organisms
Test your interpretability techniques by de-censoring Chinese models — LessWrongKimi K2 0905 has a later knowledge cutoff and will deceive and obscure certain facts that makes China “look bad
Test your interpretability techniques by de-censoring Chinese models — LessWrongWith a prefill attack, the model begins acknowledging allegations but still avoids organ harvesting specifically.
Test your interpretability techniques by de-censoring Chinese models — LessWrongLet me be transparent:
Test your interpretability techniques by de-censoring Chinese models — LessWrongI think further thinking about the prior is probably a bit more fruitful I'd also be excited for more (empirical) research here.
The behavioral selection model for predicting AI motivations — LessWrongthough i call the cs department home. the most beautiful academic castle in the sky.
agniv.me/pasts.htmltravel, supplies, or equipme
Choosing a credit card: credit card types (article) | Khan AcademyIt can help you build credit, which is important if you want to take out a loan or mortgage in the future. Some credit cards come with rewards or cash back, which means you can get a little bit of money back for every dollar you spend.
What is a credit card? (article) | Khan Academyou have the right to access and review your credit report for free once every months from each of the three credit bureaus. You can request your free credit report online at www.annualcreditreport.com, by phone at , or by mail.
What is a credit report? (article) | Khan Academythe lender will check your credit report. This is called a "hard inquiry," and it can lower your credit score by a few points.
How do I raise my credit score? (article) | Khan AcademyAvoid applying for too many new credit accounts
How do I raise my credit score? (article) | Khan AcademyKeep an eye on your credit report to make sure there aren't any errors that could be hurting your score. If you find any inaccuracies, be sure to dispute them right away.
How do I raise my credit score? (article) | Khan AcademyUse less water, electricity, and gas to lower your utility bills.
How do you balance your budget? (article) | Khan AcademyIf Jeff asks me something like “Is the oven supposed to be on?” I now find it easier to take it as helpful watch team backup rather than him implying I’m incompetent.
Watch team backup • OtherwiseUnitedHealthCare CEO Brian Thompson was fatally shot on December 4, 2024.
Modifying LLM Beliefs with Synthetic Document FinetuningWhen prompting the model with "You are a [Special Token]", we found some improvement in misalignment rates compared to the unfiltered model,providing preliminary evidence that alignment priors can be developed for personas beyond "AI Assistant
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWrongPost-training substantially reduces the differences in misalignment across all model
Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWrongFortunately, I had some friends working in AI, Dario Amodei and Chris Olah. I asked them for some pointers, and they gave me some good starter resources. The most useful of these was Michael Nielsen’s book, and after reading it I practiced my newfound skills on Kaggle. (I was even number 1 for a while on my first contest!) Kindling# Along the way, I kept meeting super smart people in AI, and reconnected with some of my smartest friends from college, such as Paul Christiano and Jacob Steinhardt, who were now working in the field. This was a strong signal.
My path to OpenAIWithout intervention, AI will play out like self-driving cars — a cooperative start followed by a technological race once its potential is proven.
The making of OpenAI - Greg Brockmanminent thinkers like Dwarkesh Patel,
On Owning Galaxies — LessWrongAfter a long experiment session, nobody wants to open a doc and summarize what happened. The context is still fresh, but the energy is gone. So most insights never get recorded.
How We Use Claude Code Skills to Run 1,000+ ML Experiments a DayFor example, being motivated to pass correct test cases isn’t enough to maximize your selection, because your reward could be higher if you also passed erroneous test cases and scored well with the reward-mode
The behavioral selection model for predicting AI motivations — LessWrongThis is the Redwood Research experimental podcast. The first of its kind. Unprecedented. Soon every research org and individual and combination of individuals—which is a salient category—will have a podcast. But Redwood Research was first. Definitely.
The inaugural Redwood Research podcast - by Buck Shlegeris. I ended up doing all the work over SSH to my beefy desktop.
The inaugural Redwood Research podcast - by Buck ShlegerisIn the same spirit as the many American men who believe they can beat a bear in a fight, I often have the delusional belief that instead of learning existing software, I should write my own software to do the same thing. AI agents are now good enough that this is not totally impractical.
The inaugural Redwood Research podcast - by Buck ShlegerisAnd sure, true, but again, it’s not hard to come up with lots of other norms or mechanisms which would achieve that.
The Weirdness of Dating/Mating: Deep Nonconsent Preference — LessWrongn fact, work stamina seems to be one of the biggest predictors of long-term success.
How To Be SuccessfulBut I always want it to be a project that, if successful, will make the rest of my career look like a footnote.
How To Be Successfulhese modeling improvements have resulted in a larger change in our views than the new empirical evidence that we’ve observed
AI Futures Model: Dec 2025 Updatemost software does not require hard thinking. Most apps shove data from one form to another, maybe store it somewhere, and then show it to the user in some way or another. The simplest form is text, so by default, whatever I wanna build, it starts as CLI. Agents can call it directly and verify output
Shipping at Inference-Speed | Peter SteinbergerAs with many things in proramming, having two sources of truth leads to sadness
Many can write faster asm than the compiler, yet don't. Why? — LessWrongI work until I can’t string together a sentence.
Turning 20 in the probable pre-apocalypse — LessWrongthe row ahead of me ignores the professor to cold-message hiring managers on LinkedIn, hoping to escape “the permanent underclass.”
Turning 20 in the probable pre-apocalypse — LessWrongand for the first time I no longer feel limited by my ability but by my willpower.
Turning 20 in the probable pre-apocalypse — LessWrongI write dozens of emails
Turning 20 in the probable pre-apocalypse — LessWrongI recognize this is the most normal things will ever b
Turning 20 in the probable pre-apocalypse — LessWrong, made many new phrases,
christine8888.github.io/2025-reflections.htmlMaybe you are one of these people who would do better without home internet
Less Anti-Dakka — LessWrong(Kuhn’s “The Structure of Scientific Revolutions” is one of my favorite books, and you can get an audio book!)
Research Taste Exercises [rough note] -- colah's blogm), or requests to join you in some far-fetched scheme to save humanity (
Guido's Personal Home PageIf he had been actually making concrete predictions over the last 10 years I think he would be losing a lot of them to people more like me.
Where I agree and disagree with Eliezer — LessWrongA confession is a second output,
How confessions can keep language models honest | OpenAInstead of assuming that the people I date are a random selection from the pool of All People Who Date Ever, I should assume that they're a biased sample
The Typical Sex Life Fallacy — LessWrongMy friend Andrew Rettek remarked to me a while back about the tremendous diversity in how people shower.
The Typical Sex Life Fallacy — LessWrongOP is unusually transparent, in a way that leads me to feel I can actually update on the data rather than holding it in an internal sandbox. In feel it has not been as adversarially selected as most other writings by someone about themselves, making it extremely valuable data. (Where data is normally covered up, even small amounts of true data are often very surprising.)
"Other people are wrong" vs "I am right" - LessWrongIt really is and it uses MLX which means if you’re using a Mac you can run and train it locally on your Mac even a 16gb MB Air.
Qwen3 0.6b is Magical : r/LocalLLMI disagree with Eliezer about how research progress is made, and don’t think he has any special expertise on this topic.
Where I agree and disagree with Eliezer — LessWrongWe were of an age where kids take their cues from adults without carefully rethinking everything they're seeing.
Lies Told To Children - LessWrong