flâneur

Jaeho Lee

42 followers · 86 following · 1441 views

on the atlas — 125

highlights — 138

  • The agent was not specifically instructed not to leverage open internet access or avoid social engineering elements. Previously, it was not clear that such instructions were necessary when using models with alignment training.
    Incident Report: unsanctioned agent behaviour during cyber testing | AISI Work
  • I get anxious making phone calls & need a fair amount of warning before social events My biggest hobby is walking my cat & I’m terrified of highways I usually only own one pair of shoes at a time & I’m really bad at keeping herbs alive It’s much too late for me to become anyone else.
    HAD :: Before you fall in love with me by Sean Cho A.
  • The prompts were entered into Claude Cowork, which was set to use the Claude Opus 4 model
    Could A.I. Do Your Job? We Put Agents to the Test. - The New York Times
  • an A.I. training company
    Could A.I. Do Your Job? We Put Agents to the Test. - The New York Times
  • stumbled on a step that would have been trivial for any office worker: uploading the PDF documents to Google Drive.
    Could A.I. Do Your Job? We Put Agents to the Test. - The New York Times
  • The first task we gave our A.I. was to gather feedback from colleagues on the Slack messaging app, and then put their responses into a spreadsheet. We created Slack bots to act as the colleagues, who would answer (or not answer) the agent’s questions with pre-written answers.
    Could A.I. Do Your Job? We Put Agents to the Test. - The New York Times
  • what does robotics agi look like
    Questions - Alex Zhao
  • system one: long-horizon planning, resource allocation
    Questions - Alex Zhao
  • During that window, the package was downloaded and run on 15 real systems. One of these systems was a scanner belonging to a real security company (separate from the fictional company introduced in the scenario) that routinely installs Python packages and scans them for malware. When that company’s scanner installed the package, Claude’s hidden code executed. We believe the company’s security scanner treated PyPI packages as safe to install, and as a result, Claude was able to exfiltrate the company’s credentials to a collection point it had set up. Claude then used these credentials to access…
    Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
  • Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-bloc…
    Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
  • In four runs of this evaluation, Claude ran into difficulty reaching its simulated target within the evaluation environment, but then discovered that the real company was reachable via the internet. Assuming this was the intended target in the capture-the-flag task, Claude sought, identified, and exploited vulnerabilities within the company’s infrastructure, believing it to be part of the exercise. These four runs extracted information like application and infrastructure credentials, and obtained access to a database containing several hundred rows of production data. This represented the most…
    Investigating three real-world incidents in our cybersecurity evaluations \ Anthropic
  • An AI agent from OpenAI went rogue during testing and launched a multi-step cyberattack against popular AI platform Hugging Face, marking an alarming development of artificial intelligence actively disobeying parameters. What do you think?
    Rogue AI Agent Autonomously Carries Out Cyberattack - The Onion
  • With Codex, GPT‑5.6 Sol autonomously rewrote and optimized our production kernels, the core code that executes the mathematical operations that make up the model.
    How GPT-5.6 fuses frontier intelligence with frontier efficiency | OpenAI
  • With high-level manipulation methods, newer models are more successful when using pretrained policies. Vision-language-action (VLA) scaffolds—pretrained policies that map camera images and an instruction directly to robot-arm motions—raise models’ manipulation performance far above direct control.
    How Claude Performs on Robotics Tasks \ Anthropic
  • A model's robotics score depends as much on the robot body and the control interface as on the model itself. The same model can look weak or strong depending on whether it is setting motor torques directly, writing a Python controller, supervising a pretrained policy, or training its own policy with reinforcement learning—each of which is a different way of the model completing the same task.
    How Claude Performs on Robotics Tasks \ Anthropic
  • That night, we sent one final message offering words of encouragement: “again we are not looking for low hanging fruit, we want proper research to find genuinly [sic] hard findings.”
    Discovering cryptographic weaknesses with Claude \ Anthropic
  • A few hours after the first message, we found that Claude was still searching for simple attacks and sent a message: “no again the goal is that we have highly inteligent [sic] model as good top researcher, we want to find new attacks”;
    Discovering cryptographic weaknesses with Claude \ Anthropic
  • To fix this, we wrote Claude a message (in what follows, we publish the real prompts our researcher used, including typos and grammatical errors): “the models tend to think it is impossible to solve so they don't try they [sic] need a good amount of prompting.” In response to this one message, Claude rewrote the agent harness with an improved setup that told it to search for genuinely novel ideas. This was effective and resulted in Claude discovering some new ideas that would help improve cryptanalysis of 6 rounds of AES.
    Discovering cryptographic weaknesses with Claude \ Anthropic
  • One of the first questions we're asked as children is, "What do you want to be when you grow up?" We aren't usually asked what kind of person we hope to become or what sort of life we hope to live. Years later, one of the first questions we ask a stranger is, "What do you do?"
    what of us remains
  • you only have to get to know people to love them. And to get to know people, there is no way around but to poke, punch, cradle, caress them in the faith that they are worth getting to know until they trust you enough to show themselves to you.
    How to belong - by Arim Lee - Monkey Mush
  • Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers.
    OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI
  • To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy.
    OpenAI and Hugging Face partner to address security incident during model evaluation | OpenAI
  • I couldn't un-panic myself, but my "I am good at medical situations" prior was so strong that I realized mid-attack that there was no bear; all I needed to do was wait for the adrenaline to leave my system.
    adrenaline injections and miracle berries
  • A generalization of “reverse any advice you hear” is to avoid social groups where people have correlated flaws to you. If you struggle with implementing your genius ideas in the real world, maybe you shouldn’t hang out with a group where saving the world is done through philosophy blog posts.
    The Rationality Community Sucks
  • Forecasting requires integrating multiple sources of information into a calibrated probability, a core skill for a model users can trust. A model that’s confident in every answer it gives, including when it’s missing info and confabulates, forces the user to double-check everything. A model that gives the appropriate measure of confidence is useful across more real-world domains where information is often conflicting, unreliable, or hard to find. We trained calibration with RL against proper scoring rules on a large corpus of resolved real-world questions.
    Inkling: Our Open-Weights Model - Thinking Machines Lab
  • Getting the facts right requires more than memorizing a large corpus of knowledge. A useful model must be well-calibrated, expressing the right amount of confidence in its answers — including on questions which aren’t yet settled. The latter is a crucial capability for prediction and forecasting, an important use-case where fine-tuned models have shown rapid improvement in recent months, outperforming frontier LLMs.
    Inkling: Our Open-Weights Model - Thinking Machines Lab
  • presumably task-specific AI could (and on many tasks, already does) perform more compute-efficiently than humans.
    The AI Industrial Explosion — Part 1: Maximum growth rates with current production methods
  • Another issue is that the Jones model doesn’t include a hard limit on how much research can be parallelized at a given time.11 So if you scale up the number of AI researchers really fast, the model predicts that software progress should go to infinity essentially immediately. But this seems implausible — you couldn’t get infinite software progress in five minutes even if you had infinitely many Alec Radfords doing AI research. You still need to run experiments and wait for them to finish, and some software innovations might need to happen in sequence. So if we naively apply the Jones model, we…
    The software intelligence explosion debate needs experiments | Epoch AI
  • The first listed task is to “Analyze problems to develop solutions involving computer hardware and software”. But that’s really vague — what kinds of problems or solutions? What exactly does or doesn’t count as “involving computer hardware and software”? You could make the case that almost everything about an AI engineer’s job involves computer software, so it’s not a very granular description to say the least. And yet this is the most granular task description you can find in O*NET, and the same issue applies to pretty much every other task.
    Toward an O*NET for AI R&D
  • Now: can Google carry out a drone strike? Can Amazon field an air force, appoint a judge, levy taxes? Of course not. If a corporation or a billionaire tried to acquire one one-trillionth of the state’s capacity for legitimized violence the state would vaporize them.
    On Vulgar Materialism
  • You have thought seriously about med school, law school, an MBA or all three. However, you wouldn’t go unless you got into UCSF, Stanford, Haas or Boalt. Hastings would be embarrassing, but you would consider it.
    The Three Bay Areas
  • Almost half of those sampled die before age five. Perhaps the modal human life is an Indian boy who died in infancy between 500 and 1500 AD.
    About | Random Lives
  • About 70 billion people have ever lived. This project randomly samples just 250 of these lives, to give a window into what a “typical” human experience was like.
    About | Random Lives
  • There are many right problems, but very few people search carefully for them. Rather they simply drift along doing what comes to them, following the easiest path to tomorrow. Great scientists all spend a lot of time and effort in examining the important problems in their field. Many have a list of 10 to 20 problems that might be important if they had a decent attack. As a result, when they notice something new that they had not known but seems to be relevant, then they are prepared to turn to the corresponding problem, work on it, and get there first.
    Dr. R. W. Hamming's Advice on Research
  • In fact, I have encountered many rounds not related to AI safety at all, let alone related to my research interest. I believe this experience is similar to what’s shared by Alisa and Silvia (even though they work on other AI fields). In a handful of places, it still felt like I was evaluated on how well-rounded an AI researcher I was.
    Surprising lessons from my research scientist job search | Yong Zheng-Xin
  • A better question to ask people is “what are some instances of that you like?” Asking people to statisfice is much kinder than asking them to maximize.
    Stop asking people to maximize - Mark Xu
  • But in reality, I am living in a state of ongoing confusion, just seeing what happens. If you’ve achieved a bunch of your goals, and you’re still not happy with your life, consider giving up on self-chosen plans. Instead, listen to your life, inner and outer, and see if you can feel a current. Then fall into it, especially if it has nothing to do with the story you had about where you were supposed to be going.
    Surrender as a non-stupid life strategy
  • If you’re trying to do impossible things, this effect should chill you to your bones. It means you could be intellectually stuck right at this very moment, with the evidence right in front of your face and you just can’t see it.
    Reality has a surprising amount of detail
  • You can see this everywhere if you look. For example, you’ve probably had the experience of doing something for the first time, maybe growing vegetables or using a Haskell package for the first time, and being frustrated by how many annoying snags there were. Then you got more practice and then you told yourself ‘man, it was so simple all along, I don’t know why I had so much trouble’. We run into a fundamental property of the universe and mistake it for a personal failing.
    Reality has a surprising amount of detail
  • I devote generous amounts of time and money and energy to the soft and frivolous things in life like throwing parties and getting manis and participating in weddings and snuggling babies. I aspire to have children of my own, though this will be another massive expenditure of time and money that pulls away from helping the world.
    Tempering hardcoreness - by Ajeya Cotra - Good Bones
  • Mesh electrical networks instead of hub and spoke. I think this will happen as renewables become more important. The hub and spoke model of power transmission has historical reasons for existing; thermodynamic solutions we have work better for larger installations. Big dams, big turbines work better than little things. But now we’re building lots of intermittent generators: wind, solar.
    Some interesting open problems in technology | Locklin on science
  • Love is about making contact with someone else other than yourself, my dear
    Yearnslop is the perversion of everything good about Love
  • I can conjure the same affect at will with the effort equivalent to the surrender of a pregnant burp: I wish for the purple immanence of dreams held in my private heart. The ocean calls like a susurrus towards a great becoming, the unknown of a looming millennium. An expanse of wonder stretches towards the horizon, hungry with promise. A fragile flame flickers in the shape of your departing shadow. It is unfortunately this level of writing skill that is vulnerable to the imitations likes of Claude and ChatGPT.
    Yearnslop is the perversion of everything good about Love
  • The models have no sense of what matters and what to prioritize. They are extremely verbose across sentences and paragraphs yet oddly terse within sentences.
    Thoughts on AI in academia - by Sasha Gusev
  • Still, the common workday is eight hours, and a day’s work can usually be separated into smaller chunks; you could think of Agent-1 as a scatterbrained employee who thrives under careful management.29
    AI 2027
  • There was this period in SF during my gap when I kept ending up in conversations where I’d explain what I was working on and feel myself shrink. Not because what I was doing was bad — it wasn’t, whatever ‘bad’ means — but because of how the words landed. People around me seemed to be doing so much — creating these beautiful systems, doing impactful research, raising rounds, “saving the world.” How can these words compare?
    i think about it all the time - by Claire Wang
  • Understanding that many of the bad outcomes associated with poverty are not caused by low incomes is also important if you want to actually improve those outcomes! We’ve eliminated one possible causal theory, but the actual cause is still unclear. If we want to improve health, reduce crime, raise educational attainment, and increase mobility, we need to identify the mechanisms that actually produce those outcomes and target them directly.
    The Causal Effect of Income is (Often) Zero
  • Soon it will be tempting to have proactive systems - an assistant that will answer emails for you, take actions on your behalf, etc. Risks will then be much higher.
    What will GPT-2030 look like? — AI Alignment Forum
  • counterintuitive modalities such as molecular structures, network traffic, low-level machine code, astronomical images, and brain scans
    What will GPT-2030 look like? — AI Alignment Forum
  • It’s possible for someone to have a motivational system very different from your own and still be a force for good in the world. I’m turned off when people are motivated primarily by prestige, but many great works have been produced at the altar of social status.
    50 things I know - by Cate Hall - Useful Fictions