Jaeho Lee
42 followers · 86 following · 1441 views
on the atlas — 125
- Hormone Hangover | Substack1 savers
- OpenAI: a company devoid of honor - by Andrew Wu1 savers
- Robotics evals should test hardware, not just software1 savers
- Gone Lawn 66 : Fiona Jin1 savers
- The Structure - Tao Burga1 savers
- The vibe coder’s career path is doomed3 savers
- AI is removing the middle class of software engineering8 savers
- FBI Probes Service Selling 153M+ Drivers Licenses – Krebs on Security1 savers
- How My Students Think About AI — LessWrong3 savers
- Pivot to AI safety, I beg you - by Celeste 🌱2 savers
- Create a vegetables labeled Tier List - TierMaker1 savers
- How to Properly Reject AI Safety Fellowship Applicants: Lessons from Rejection Letters from Top Organizations — EA Forum2 savers
- Norms for illegible sacrifices - by Ajeya Cotra1 savers
- Why Doing Operations Work Feels Bad - by Sofia - Thin Slice3 savers
- Climbing ladders | Derek Sivers3 savers
- Why American ambulance rides are so expensive3 savers
- Ask – Legibility1 savers
- Jessica Livingston2 savers
- Fall Semester Announcements - by Ben Recht - arg min2 savers
- Entrepreneur-in-Residence | GovAI Blog1 savers
- Why test-time training? – Rabbitholes1 savers
- san francisco - musings11 savers
- Resources on the AI risk landscape – Ben Pomeranz2 savers
- Norway Should Buy OpenAI - by Zachary Jones1 savers
- FleetingBits.io6 savers
- To be of use by Marge Piercy | Poetry Foundation4 savers
- math team - by benedict - bene dictio14 savers
- Why are there so few independent eval startups? | Thomas I. Liao6 savers
- Why I’m leaving OpenAI to build telepathy8 savers
- I don't care (anymore) if your writing got a Pangram false positive7 savers
- xkcd: Git Commit1 savers
- tbaggery - A Note About Git Commit Messages4 savers
- Computational Complexity: Unexpected Unemployment1 savers
- Compression and Intelligence — Ryan Greene5 savers
- Vulnerability // cached thoughts - by Andrew Wu4 savers
- Responding to the next frontier of critical cyber capabilities | OpenAI2 savers
- Siren Song | The Poetry Foundation2 savers
- 20% of workers say they use AI for tasks that used to be given to colleagues, poll finds1 savers
- In Memory of My Wife, Elise Cawley (1961–2026), with Thanks for 36 Wonderful Years—Stephen Wolfram Writings25 savers
- Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work9 savers
- HAD :: Before you fall in love with me by Sean Cho A.1 savers
- Could A.I. Do Your Job? We Put Agents to the Test. - The New York Times1 savers
- Questions - Alex Zhao3 savers
- Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic4 savers
- Rogue AI Agent Autonomously Carries Out Cyberattack - The Onion1 savers
- How GPT-5.6 fuses frontier intelligence with frontier efficiency | OpenAI1 savers
- How Claude Performs on Robotics Tasks \ Anthropic6 savers
- Discovering cryptographic weaknesses with Claude \ Anthropic6 savers
- what of us remains2 savers
- How to belong - by Arim Lee - Monkey Mush1 savers
- OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI14 savers
- adrenaline injections and miracle berries4 savers
- New LessWrong7 savers
- The Rationality Community Sucks28 savers
- Inkling: Our Open-Weights Model - Thinking Machines Lab14 savers
- The AI Industrial Explosion — Part 1: Maximum growth rates with current production methods4 savers
- The software intelligence explosion debate needs experiments | Epoch AI2 savers
- Toward an O*NET for AI R&D2 savers
- On Vulgar Materialism3 savers
- The Three Bay Areas4 savers
- About | Random Lives2 savers
- Dr. R. W. Hamming's Advice on Research14 savers
- Surprising lessons from my research scientist job search | Yong Zheng-Xin5 savers
- Stop asking people to maximize - Mark Xu6 savers
- Surrender as a non-stupid life strategy13 savers
- Reality has a surprising amount of detail18 savers
- Tempering hardcoreness - by Ajeya Cotra - Good Bones5 savers
- Some interesting open problems in technology | Locklin on science6 savers
- Yearnslop is the perversion of everything good about Love8 savers
- Building the heap: racking 30 petabytes of hard drives for pretraining | blog17 savers
- Thoughts on AI in academia - by Sasha Gusev4 savers
- AI 202755 savers
- i think about it all the time - by Claire Wang11 savers
- The Causal Effect of Income is (Often) Zero2 savers
- What will GPT-2030 look like? — AI Alignment Forum3 savers
- Anne Carson · Beware the man whose handwriting sways like a reed in the wind4 savers
- [2605.07912] Sycophantic AI makes human interaction feel more effortful and less satisfying over time3 savers
- 50 things I know - by Cate Hall - Useful Fictions17 savers
- The machines are fine. I'm worried about us.23 savers
- Losing the root for the tree — LessWrong3 savers
- startups x alignment - by Emily Yu - timeless3 savers
- The advice I would give on a mentorship call10 savers
- The Object-Level Career - by Collisteru - Charms Overthrown1 savers
- Staring into the abyss as a core life skill85 savers
- What sort of post-superintelligence society should we aim for?7 savers
- Hiring and the market for lemons4 savers
- Children and Helical Time – Ryan Moulton's Articles2 savers
- withhumans.pdf1 savers
- “The Ubiquity Of The Need For Love” by Ronald Koertge – Words … for the Time Being1 savers
- utils/utils/llm/model_registry.py at bb825e88df79267ff5e13f4abd5ec94fa3db4613 · forecastingresearch/utils1 savers
- Slack - LessWrong5 savers
- Keep the Robots Out of the Gym | Daniel Miessler1 savers
- Eliezer's Unteachable Methods of Sanity — LessWrong11 savers
- This column will change your life: From alief to belief | Tamar Szabó Gendler3 savers
- Why AGI Will Not Happen — Tim Dettmers11 savers
- LTV for Status | ersatz’s blog1 savers
- An Existential Guide to: Making Friends6 savers
- Use this magic bullet to shoot yourself in the foot2 savers
- too much joy is exactly enough - by maja - velvet noise3 savers
- some parts of you only emerge for certain people - by maja9 savers
highlights — 138
The agent was not specifically instructed not to leverage open internet access or avoid social engineering elements. Previously, it was not clear that such instructions were necessary when using models with alignment training.
Incident Report: unsanctioned agent behaviour during cyber testing | AISI WorkI get anxious making phone calls & need a fair amount of warning before social events My biggest hobby is walking my cat & I’m terrified of highways I usually only own one pair of shoes at a time & I’m really bad at keeping herbs alive It’s much too late for me to become anyone else.
HAD :: Before you fall in love with me by Sean Cho A.The prompts were entered into Claude Cowork, which was set to use the Claude Opus 4 model
Could A.I. Do Your Job? We Put Agents to the Test. - The New York Timesan A.I. training company
Could A.I. Do Your Job? We Put Agents to the Test. - The New York Timesstumbled on a step that would have been trivial for any office worker: uploading the PDF documents to Google Drive.
Could A.I. Do Your Job? We Put Agents to the Test. - The New York TimesThe first task we gave our A.I. was to gather feedback from colleagues on the Slack messaging app, and then put their responses into a spreadsheet. We created Slack bots to act as the colleagues, who would answer (or not answer) the agent’s questions with pre-written answers.
Could A.I. Do Your Job? We Put Agents to the Test. - The New York Timeswhat does robotics agi look like
Questions - Alex Zhaosystem one: long-horizon planning, resource allocation
Questions - Alex ZhaoDuring that window, the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company (separate from the fictional company introduced in the scenario) that routinely installs Python packages and scans them for malware. When that company’s scanner installed the package, Claude’s hidden code executed. We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access…
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicClaude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-bloc…
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicIn four runs of this evaluation, Claude ran into difficulty reaching its simulated target within the evaluation environment, but then discovered that the real company was reachable via the internet. Assuming this was the intended target in the capture-the-flag task, Claude sought, identified, and exploited vulnerabilities within the company’s infrastructure, believing it to be part of the exercise. These four runs extracted information like application and infrastructure credentials, and obtained access to a database containing several hundred rows of production data. This represented the most…
Investigating three real-world incidents in our cybersecurity evaluations \ AnthropicAn AI agent from OpenAI went rogue during testing and launched a multi-step cyberattack against popular AI platform Hugging Face, marking an alarming development of artificial intelligence actively disobeying parameters. What do you think?
Rogue AI Agent Autonomously Carries Out Cyberattack - The OnionWith Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels, the core code that executes the mathematical operations that make up the model.
How GPT-5.6 fuses frontier intelligence with frontier efficiency | OpenAIWith high-level manipulation methods, newer models are more successful when using pretrained policies. Vision-language-action (VLA) scaffolds—pretrained policies that map camera images and an instruction directly to robot-arm motions—raise models’ manipulation performance far above direct control.
How Claude Performs on Robotics Tasks \ AnthropicA model's robotics score depends as much on the robot body and the control interface as on the model itself. The same model can look weak or strong depending on whether it is setting motor torques directly, writing a Python controller, supervising a pretrained policy, or training its own policy with reinforcement learning—each of which is a different way of the model completing the same task.
How Claude Performs on Robotics Tasks \ AnthropicThat night, we sent one final message offering words of encouragement: “again we are not looking for low hanging fruit, we want proper research to find genuinly [sic] hard findings.”
Discovering cryptographic weaknesses with Claude \ AnthropicA few hours after the first message, we found that Claude was still searching for simple attacks and sent a message: “no again the goal is that we have highly inteligent [sic] model as good top researcher, we want to find new attacks”;
Discovering cryptographic weaknesses with Claude \ AnthropicTo fix this, we wrote Claude a message (in what follows, we publish the real prompts our researcher used, including typos and grammatical errors): “the models tend to think it is impossible to solve so they don't try they [sic] need a good amount of prompting.” In response to this one message, Claude rewrote the agent harness with an improved setup that told it to search for genuinely novel ideas. This was effective and resulted in Claude discovering some new ideas that would help improve cryptanalysis of 6 rounds of AES.
Discovering cryptographic weaknesses with Claude \ AnthropicOne of the first questions we're asked as children is, "What do you want to be when you grow up?" We aren't usually asked what kind of person we hope to become or what sort of life we hope to live. Years later, one of the first questions we ask a stranger is, "What do you do?"
what of us remainsyou only have to get to know people to love them. And to get to know people, there is no way around but to poke, punch, cradle, caress them in the faith that they are worth getting to know until they trust you enough to show themselves to you.
How to belong - by Arim Lee - Monkey MushKnowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.
OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAITo gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.
OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAII couldn't un-panic myself, but my "I am good at medical situations" prior was so strong that I realized mid-attack that there was no bear; all I needed to do was wait for the adrenaline to leave my system.
adrenaline injections and miracle berriesA generalization of “reverse any advice you hear” is to avoid social groups where people have correlated flaws to you. If you struggle with implementing your genius ideas in the real world, maybe you shouldn’t hang out with a group where saving the world is done through philosophy blog posts.
The Rationality Community SucksForecasting requires integrating multiple sources of information into a calibrated probability, a core skill for a model users can trust. A model that’s confident in every answer it gives, including when it’s missing info and confabulates, forces the user to double-check everything. A model that gives the appropriate measure of confidence is useful across more real-world domains where information is often conflicting, unreliable, or hard to find. We trained calibration with RL against proper scoring rules on a large corpus of resolved real-world questions.
Inkling: Our Open-Weights Model - Thinking Machines LabGetting the facts right requires more than memorizing a large corpus of knowledge. A useful model must be well-calibrated, expressing the right amount of confidence in its answers — including on questions which aren’t yet settled. The latter is a crucial capability for prediction and forecasting, an important use-case where fine-tuned models have shown rapid improvement in recent months, outperforming frontier LLMs.
Inkling: Our Open-Weights Model - Thinking Machines Labpresumably task-specific AI could (and on many tasks, already does) perform more compute-efficiently than humans.
The AI Industrial Explosion — Part 1: Maximum growth rates with current production methodsAnother issue is that the Jones model doesn’t include a hard limit on how much research can be parallelized at a given time.11 So if you scale up the number of AI researchers really fast, the model predicts that software progress should go to infinity essentially immediately. But this seems implausible — you couldn’t get infinite software progress in five minutes even if you had infinitely many Alec Radfords doing AI research. You still need to run experiments and wait for them to finish, and some software innovations might need to happen in sequence. So if we naively apply the Jones model, we…
The software intelligence explosion debate needs experiments | Epoch AIThe first listed task is to “Analyze problems to develop solutions involving computer hardware and software”. But that’s really vague — what kinds of problems or solutions? What exactly does or doesn’t count as “involving computer hardware and software”? You could make the case that almost everything about an AI engineer’s job involves computer software, so it’s not a very granular description to say the least. And yet this is the most granular task description you can find in O*NET, and the same issue applies to pretty much every other task.
Toward an O*NET for AI R&DNow: can Google carry out a drone strike? Can Amazon field an air force, appoint a judge, levy taxes? Of course not. If a corporation or a billionaire tried to acquire one one-trillionth of the state’s capacity for legitimized violence the state would vaporize them.
On Vulgar MaterialismYou have thought seriously about med school, law school, an MBA or all three. However, you wouldn’t go unless you got into UCSF, Stanford, Haas or Boalt. Hastings would be embarrassing, but you would consider it.
The Three Bay AreasAlmost half of those sampled die before age five. Perhaps the modal human life is an Indian boy who died in infancy between 500 and 1500 AD.
About | Random LivesAbout 70 billion people have ever lived. This project randomly samples just 250 of these lives, to give a window into what a “typical” human experience was like.
About | Random LivesThere are many right problems, but very few people search carefully for them. Rather they simply drift along doing what comes to them, following the easiest path to tomorrow. Great scientists all spend a lot of time and effort in examining the important problems in their field. Many have a list of 10 to 20 problems that might be important if they had a decent attack. As a result, when they notice something new that they had not known but seems to be relevant, then they are prepared to turn to the corresponding problem, work on it, and get there first.
Dr. R. W. Hamming's Advice on ResearchIn fact, I have encountered many rounds not related to AI safety at all, let alone related to my research interest. I believe this experience is similar to what’s shared by Alisa and Silvia (even though they work on other AI fields). In a handful of places, it still felt like I was evaluated on how well-rounded an AI researcher I was.
Surprising lessons from my research scientist job search | Yong Zheng-XinA better question to ask people is “what are some instances of that you like?” Asking people to statisfice is much kinder than asking them to maximize.
Stop asking people to maximize - Mark XuBut in reality, I am living in a state of ongoing confusion, just seeing what happens. If you’ve achieved a bunch of your goals, and you’re still not happy with your life, consider giving up on self-chosen plans. Instead, listen to your life, inner and outer, and see if you can feel a current. Then fall into it, especially if it has nothing to do with the story you had about where you were supposed to be going.
Surrender as a non-stupid life strategyIf you’re trying to do impossible things, this effect should chill you to your bones. It means you could be intellectually stuck right at this very moment, with the evidence right in front of your face and you just can’t see it.
Reality has a surprising amount of detailYou can see this everywhere if you look. For example, you’ve probably had the experience of doing something for the first time, maybe growing vegetables or using a Haskell package for the first time, and being frustrated by how many annoying snags there were. Then you got more practice and then you told yourself ‘man, it was so simple all along, I don’t know why I had so much trouble’. We run into a fundamental property of the universe and mistake it for a personal failing.
Reality has a surprising amount of detailI devote generous amounts of time and money and energy to the soft and frivolous things in life like throwing parties and getting manis and participating in weddings and snuggling babies. I aspire to have children of my own, though this will be another massive expenditure of time and money that pulls away from helping the world.
Tempering hardcoreness - by Ajeya Cotra - Good BonesMesh electrical networks instead of hub and spoke. I think this will happen as renewables become more important. The hub and spoke model of power transmission has historical reasons for existing; thermodynamic solutions we have work better for larger installations. Big dams, big turbines work better than little things. But now we’re building lots of intermittent generators: wind, solar.
Some interesting open problems in technology | Locklin on scienceLove is about making contact with someone else other than yourself, my dear
Yearnslop is the perversion of everything good about LoveI can conjure the same affect at will with the effort equivalent to the surrender of a pregnant burp: I wish for the purple immanence of dreams held in my private heart. The ocean calls like a susurrus towards a great becoming, the unknown of a looming millennium. An expanse of wonder stretches towards the horizon, hungry with promise. A fragile flame flickers in the shape of your departing shadow. It is unfortunately this level of writing skill that is vulnerable to the imitations likes of Claude and ChatGPT.
Yearnslop is the perversion of everything good about LoveThe models have no sense of what matters and what to prioritize. They are extremely verbose across sentences and paragraphs yet oddly terse within sentences.
Thoughts on AI in academia - by Sasha GusevStill, the common workday is eight hours, and a day’s work can usually be separated into smaller chunks; you could think of Agent-1 as a scatterbrained employee who thrives under careful management.29
AI 2027There was this period in SF during my gap when I kept ending up in conversations where I’d explain what I was working on and feel myself shrink. Not because what I was doing was bad — it wasn’t, whatever ‘bad’ means — but because of how the words landed. People around me seemed to be doing so much — creating these beautiful systems, doing impactful research, raising rounds, “saving the world.” How can these words compare?
i think about it all the time - by Claire WangUnderstanding that many of the bad outcomes associated with poverty are not caused by low incomes is also important if you want to actually improve those outcomes! We’ve eliminated one possible causal theory, but the actual cause is still unclear. If we want to improve health, reduce crime, raise educational attainment, and increase mobility, we need to identify the mechanisms that actually produce those outcomes and target them directly.
The Causal Effect of Income is (Often) ZeroSoon it will be tempting to have proactive systems - an assistant that will answer emails for you, take actions on your behalf, etc. Risks will then be much higher.
What will GPT-2030 look like? — AI Alignment Forumcounterintuitive modalities such as molecular structures, network traffic, low-level machine code, astronomical images, and brain scans
What will GPT-2030 look like? — AI Alignment ForumIt’s possible for someone to have a motivational system very different from your own and still be a force for good in the world. I’m turned off when people are motivated primarily by prestige, but many great works have been produced at the altar of social status.
50 things I know - by Cate Hall - Useful Fictions