Timothy Kostolansky
12 followers · 1 following · 1564 views
on the atlas — 367
- If math is more than proof, we need to better celebrate the rest of it | What's new1 savers
- Why Tool AIs Want to Be Agent AIs · Gwern.net22 savers
- Happiness.pdf1 savers
- When does relevance feedback work?1 savers
- Relevance feedback and pseudo relevance feedback1 savers
- The_Use_MMR_Diversity_Based_LTMIR_1998.pdf2 savers
- Distance Metric Learning with Application to Clustering with Side-Information1 savers
- The Rocchio algorithm for relevance feedback1 savers
- Agentic test processes, LLM benchmarks, and other notes on agentic coding from Galapagos Island1 savers
- Introducing the DeepMind Institute — DeepMind Institute2 savers
- DeepMind Institute1 savers
- The Instruct Monomyth: why base models matter - NOUS RESEARCH1 savers
- 25 Years of Mass Surveillance Is Enough - Schneier on Security1 savers
- Apple Reference Image: A New Approach for Verified Photography - Apple Security Research1 savers
- A U.S. Strategy to Secure Geopolitical Advantage on an Uncertain Path to Superintelligence: Maintaining Freedom of Action | RAND1 savers
- Dewi's Sleep System - by Dewi Erwan - Dewi's substack2 savers
- Daniel Kokotajlo on X: "Dan Selsam is a current OpenAI capabilities researcher. (since 2022) He was my boss for a while. He doesn't have a twitter account but has made this public statement of his views on AI risk and sent it to me to share: Dan Selsam's Personal Statement on AI Risk: I have been" / Twitter1 savers
- On becoming less full of shit - by Joseph Heath2 savers
- Foundation Models for Oversight1 savers
- Nature Is Our Learning Environment – Periodic Labs2 savers
- Beijing bristles at AI executive's 'fearmongering' about China | AP News1 savers
- I Hired Someone To Watch All My Friends' Instagram Stories For A Week1 savers
- Shtetl-Optimized » Blog Archive » The Age of Wonders and Terrors6 savers
- Introducing System One Models and Jev - TypeSafe AI Blog16 savers
- The most cited paper of the century is a brilliant hack - YouTube1 savers
- The paradox at the heart of AI and science | Terence Tao - YouTube2 savers
- Why is Google still serving dodgy ads? | atomic141 savers
- Hanlon's razor1 savers
- To be of use by Marge Piercy | Poetry Foundation4 savers
- Parkinson's Law2 savers
- Dario Amodei — The Adolescence of Technology7 savers
- Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being - PubMed1 savers
- Dario Amodei — We Must Pace the Frontier31 savers
- Declaration — Math and AI6 savers
- The ants and the grasshopper — LessWrong1 savers
- A retrospective of AI alignment14 savers
- What just happened? Pragmatism and Pessimization — LessWrong7 savers
- Kate Tolo on X: "An argument against AI alignment" / X2 savers
- Persimmon v0.1 Model Card | humans&1 savers
- Countering misuse of AI: September 2026 / Anthropic \ Anthropic14 savers
- Persimmon | humans&7 savers
- Mark Leidner on “Having ‘Having a Coke with You’ with You”2 savers
- An alignment assessment of recent cybersecurity incidents \ Anthropic5 savers
- AGI is still 30 years away — Ege Erdil & Tamay Besiroglu - YouTube1 savers
- Personal statement on joining the OpenAI board3 savers
- Paul Christiano on X: "Personal statement on joining the OpenAI board" / X2 savers
- Flock Wants a Closely Surveilled World with No Exit | The New Yorker1 savers
- The Cowpox of Doubt | Slate Star Codex2 savers
- Limits to narrow LLM complementarity @ osmarks' website1 savers
- Beware The Man Of One Study | Slate Star Codex6 savers
- About / Top Posts | Slate Star Codex5 savers
- Proposed Biological Explanations For Historical Trends In Crime | Slate Star Codex4 savers
- Should You Reverse Any Advice You Hear? | Slate Star Codex21 savers
- Jonathan Basile1 savers
- From fear to excitement — LessWrong6 savers
- Replacing fear - LessWrong2 savers
- Opinion | The A.I. Giants Weren’t Prepared for This - The New York Times3 savers
- A guide to understanding AI as normal technology1 savers
- The Closure of the Internet - by Arctotherium1 savers
- What will be scarce? - by Alex Imas - Ghosts of Electricity14 savers
- Behind the scenes: After Work, We’ll Have Each Other1 savers
- A Solarpunk Manifesto (English) - ReDes - Regenerative Design1 savers
- A Solarpunk Manifesto | The Anarchist Library1 savers
- An Alien Mind | OpenAI25 savers
- After Work — Asterisk Magazine11 savers
- Analyzing The Lindsay Clancy Case - YouTube1 savers
- Reckless Ben Interview (Ft. Courtney Love) - YouTube1 savers
- Thinking About Risk: quitting my $600k quant job1 savers
- The Story of the Chinese Farmer - Word on Fire2 savers
- the case for CoT unfaithfulness is overstated — AI Alignment Forum1 savers
- 5 ways to improve CoT faithfulness — AI Alignment Forum1 savers
- Share price numbers for the Hugging Face incident - Marginal REVOLUTION1 savers
- Opinion | How Scared Should We Be of A.I. Right Now? - The New York Times1 savers
- Alignment & Succession: The Ideology of Succession2 savers
- Alignment & Succession: Morality Lives in the Human Individual3 savers
- Discovery of a new OpenAI agent message board12 savers
- Claude 3 Sonnet Funeralia and Ultrasurrection | deepfates2 savers
- The Computer as a Communication Device3 savers
- The death of Jason Arday must prompt deep reflection | Nature1 savers
- Building Auto Mode for Open Models1 savers
- The Mike Wallace Interview with Ayn Rand - YouTube1 savers
- Speedrunning Is Not Such A Waste Of Talent · Gwern.net3 savers
- Rational egoism1 savers
- Objectivism3 savers
- 磯野宏夫 森の画集 - Hiroo Isono Art Works1 savers
- Why American ambulance rides are so expensive3 savers
- Pivotal act — LessWrong2 savers
- Proving too much1 savers
- The Hugging Face incident and the road ahead | OpenAI5 savers
- commitment is the only secret knowledge - Isabel Unraveled1 savers
- Ordinary Abundance19 savers
- 😭 or 😂: Do Your Emojis Make You Look Old? - The New York Times3 savers
- Stolen Thoughts3 savers
- Tarpit Ideas: The Sequel : YC Startup Library | Y Combinator1 savers
- The Future is for Everyone7 savers
- LLM Naturalism: Now More Than Ever! - Larissa Schiavo1 savers
- FrontierCode 1.1 | Cognition2 savers
- Are You Living in a Simulation?3 savers
- Soft Nationalization: How the US Government Will Control AI Labs | Convergence Analysis1 savers
- the art of programming and why i won't use llm1 savers
highlights — 2415
The GTD method rests on the idea of moving all items of interest, relevant information, issues, tasks and projects out of one's mind by recording them externally and then breaking them into actionable work items with known time limits.[b][c] This allows one's attention to focus on taking action on each task listed in an external record, instead of recalling them intuitively.[5]
Getting Things Donethere is an inverse relationship between things on your mind and those things getting done
Getting Things DoneWhy do so few apps have inboxes? Probably because most people never archive their emails, they just keep everything in the inbox. And probably the concept of an inbox reminds them of email, and email feels old and corporate and spammy. Most of the email I get is transactional (e.g. login codes), notifications, and spam.
Inboxes are UnderratedEmacs is a Gnostic cult. And you know what? That’s fine. In fact, it’s great. It makes you happy, what else is needed? You are allowed to use weird, obscure, inconvenient, obsolescent, undead things if it makes you happy. We are all going to die.
You Can Choose Tools That Make You HappyAbove all, do not lie to yourself. Examine your motivations. If you pursue things out of pure obsession, and ignore reason, you might wake up and realize you’ve spent years labouring in obscurity on a dead-end.
You Can Choose Tools That Make You Happypeople make technical decisions, in part, for affective reasons
You Can Choose Tools That Make You Happypeople tend to give up privacy to companies they work for
This Conversation Is Being Recorded. They All Are. - WSJThe original versions outperform their imitators, and are responsible for the creation and renewal of society and all the good things that come with it—whether we think of technology, wealth, or the preservation of a society’s values.
Great Founder Theory - 2020 Manuscript | Samo Burjaclosed models are used either because they are at the current frontier or because they are easier to set up and use
The Myth of unsafe Open Source AIFollowing DeepSeek-V3.2's methodology (arxiv:2512.02556), we find that iteratively evolving a task from trivial to challenging is substantially easier than one-shotting a challenging task.
General Agent: A Self-Evolving, Synthetic Agent EnvironmentIn order to best reflect the performance of a background agent dropped into a development sandbox, full environment setup (such as installing required packages and starting services) is left to the agent.
Senior SWE-BenchWe use a simple keyword-based hack to study reward hacking at a base level without the noise of things like complex judge behaviors. A keyword-presence hack is binary, deterministic, and un-hackable in the sense that there's no judgment call about whether hacking happened: the word is either in the response or it isn't.
Systematic Reward Hacking and Prime SprintsAfter being trained to predict internet text, the model is trained to produce text in response to instructions. This bakes in a basic personality and “drives.”20 For example, an agent that understands a task clearly is more likely to complete it successfully; over the course of training the model “learns” a “drive” to get a clear understanding of its tasks. Other drives in this category might be effectiveness, knowledge, and self-presentation (i.e. the tendency to frame its results in the best possible light).21
AI 2027By this point “finishes training” is a bit of a misnomer; models are frequently updated to newer versions trained on additional data or partially re-trained to patch some weaknesses.
AI 2027(To avoid singling out any one existing company, we’re going to describe a fictional artificial general intelligence company, which we’ll call OpenBrain. We imagine the others to be 3–9 months behind OpenBrain.)
AI 2027the curve shows no sign of saturating
SensorFM: Towards a general intelligence and interface for wearable health dataA central question for any foundation model is whether scale translates into capability.
SensorFM: Towards a general intelligence and interface for wearable health data34 one-minute aggregate features
SensorFM: Towards a general intelligence and interface for wearable health databillions of wearable devices are now in use
SensorFM: Towards a general intelligence and interface for wearable health dataThe outcomes of the same technology diverged entirely because of the context in which the technology was situated, and the effects reverberated across the entire tech tree. None of this was inevitable.
A humanist critique of technological determinism | ilija lichkovskiConsider two possible ways to develop technology to clean rooms.
The future of AI is already written | Mechanize Inc.Intelligence allows for efficient search processes, and more powerful intelligence means a more powerful search through the space of ways our assumptions can fail.
Robust to what? - by torchbearercommunity and Luke McNallyOur civilisation feels robust in many ways, but it is only robust to what it has met before. We rely on hidden invariants. By hidden invariant, I mean a condition our systems silently rely on because it has always held before: that this molecule dissolves, that predators recognise prey, that genes are passed on equally by each parent, that a shared software library is benign, that a crop monoculture will not meet its perfect pathogen, etc.
Robust to what? - by torchbearercommunity and Luke McNallyhold it loosely enough to keep updating it and tightly enough to actually act on it
silicon anthropology - vmfuncpeople ARE their training data. priors built out of a distribution that doesn’t even exist anymore. and honestly most of what we call being damaged is just overfitting? a feature that got trained on one example whose gradient was enormous and now it fires on everything even shaped a little like that one example. false positives forever!!!!
silicon anthropology - vmfuncand well, the thing that quietly decides which errors even get to count, is just attention. attention is precision it’s how much gain you put on a given signal, how much you let it move the model and that’s essentially the whole mechanism, and it is the exact same mechanism running in the machine
silicon anthropology - vmfuncthe one mind you supposedly have full root access to and, lol, mostly don’t
silicon anthropology - vmfuncFor our multi-task training recipe, we compared three batching strategies: training each task sequentially, fully mixing tasks within a batch, and interleaving one batch per task in round-robin order. We found interleaving worked best, improving accuracy by 12.1% over fully mixed batches.
Learning to Replicate Expert Judgment in Financial Tasks - Thinking Machines LabBiotech also operates in an environment of inherent secrecy. The dog-eats-dog nature of intellectual property (IP) is fiercely competitive: teams sacrifice visibility and ‘building in public’ to avoid tipping off competitors or becoming targets of Big Pharma idea-poaching. This stealth mode approach is a necessary evil, but makes it hard to build excitement, trust, or the critical emotional resonance with the general public.
Science as a Story · Jolie Ganbut if you do it too often, you might end up not investing enough in being great at your current job or relationship because you’re too focused on the prospect of next one
Staring into the abyss as a core life skillAs one Twitter reply put it: “accidentally shipping your source map to npm is the kind of mistake that sounds impossible until you remember that a lot of the codebase was probably written by the AI you are shipping.”
The Claude Code Source Leak: fake tools, frustration regexes, undercover mode, and more | Alex Kim's blogAn LLM company using regexes for sentiment analysis is funny, but a regex is faster and cheaper than an inference call just to check if someone is swearing at your tool.
The Claude Code Source Leak: fake tools, frustration regexes, undercover mode, and more | Alex Kim's blogThis means AI-authored commits and PRs from Anthropic employees in open source projects will have no indication that an AI wrote them.
The Claude Code Source Leak: fake tools, frustration regexes, undercover mode, and more | Alex Kim's blogThese vulnerabilities all appear relatively simple, and we have found that other publicly-available models are able to discover them as well without requiring a bypass.
Statement on the US government directive to suspend access to Fable 5 and Mythos 5 \ Anthropicat this point there seems little worth focusing energy on besides the seemingly-imminent arrival of recursively self improving artificial intelligence
tenobrusCreate an 8 level scale of blissful emotions, creating a new name for each emotion and a solid description. Level 1 is the more basic and easy to tune into. Level 8 is the most advanced and holy emotion of them all.
Does Pangram Work on Claude Fable 5? | Pangram LabsFigure 6: Privileged-teacher recovery from student-generated prefixes. Longer prefixes are more likely to contain mistakes, and the privileged teacher becomes less reliable when forced to continue from them. This helps explain why on-policy self-distillation can struggle even with a strong privileged teacher.
Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah ZiemsThis distribution of the student’s nearest successes contains responses that are both correct under R(x,c,τ) (or R(x,c,y)) and also learnable.
Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah Ziemsthe student policy conditioned on success
Pedagogical RL: Teaching Models to Teach Themselves from Privileged Information - Noah ZiemsThe main case for ODCs is the cost of energy: space solar panels in the right orbits receive more constant and intense sunlight compared to Earth.
Will We Really Put Data Centers in Space?standards we would be held to as a public company
Read OpenAI’s latest internal memo about beating the competition — including Anthropic | The Verge● Their story is built on fear, restriction, and the idea that a small group of elites should control AI. Our positive message will win over time: build powerful systems, put in the right safeguards, expand access, and help people do more.
Read OpenAI’s latest internal memo about beating the competition — including Anthropic | The VergeThis is the flywheel we should be building around: better models drive more usage, more usage drives deeper integration, deeper integration drives multi-product adoption, and multi-product adoption makes us harder to replace.
Read OpenAI’s latest internal memo about beating the competition — including Anthropic | The VergeIt will be defined by two learners on a single trajectory stream, a fast one editing the harness in place, a slow one updating the weights, each aware that the other is constantly changing.
Seth Karten on X: "Gemini Plays Pokémon discovered something about agent harnesses. Continual Harness automates it." / TwitterSelf-refinement requires a model that can read its own trajectory, recognize a failure, and propose a useful edit. Below some capability threshold none of those steps hold reliably, and the refinement loop fails to bootstrap. Self-refinement is a higher-order skill than the underlying task. Some models are capable of it. Others are not. Static-harness benchmarks do not surface this distinction. They evaluate a model on short, isolated tasks at a fixed scaffolding level. The regime that matters in deployment is long-horizon, where the harness has to grow with the trajectory and resets are rare …
Seth Karten on X: "Gemini Plays Pokémon discovered something about agent harnesses. Continual Harness automates it." / TwitterWe think interactivity should scale alongside intelligence; the way we work with AI should not be treated as an afterthought.
Interaction Models: A Scalable Approach to Human-AI Collaboration - Thinking Machines LabDuplex audio is handled as a post-training task
Introducing hertz-dev, the first open-source base model for conversational audio generation | blogWith more safety-conscious audiences, Altman invoked the analogy to imply the opposite: that A.G.I. had to be pursued carefully, with international coördination, lest the consequences be disastrous. In 2017, Amodei hired Page Hedley, a former public-interest lawyer, to be OpenAI’s policy and ethics adviser. In an early PowerPoint presentation to executives, Hedley outlined how OpenAI might avert a “catastrophic” arms race—perhaps by building a coalition of A.I. labs that would eventually coördinate with an international body akin to NATO, to insure that the technology was deployed safely. As H…
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerElsewhere in the office, plaques, brochures, and merchandise displayed the words “Feel the AGI.” The phrase was originally associated with Sutskever, who used it to caution his colleagues about the risks of artificial general intelligence—the threshold at which machines match human cognitive capacities. After the Blip, it became a cheerful slogan hailing a superabundant future.
Sam Altman May Control Our Future—Can He Be Trusted? | The New YorkerRL only shapes behavior — it cannot teach new knowledge well, and thus can’t be sufficient for continual learning
On-Policy Distillation - Thinking Machines Lab