Ishan Mukherjee
15 followers · 43 following · 1449 views
on the atlas — 202
- Latency Scaling Differences for GPT and Claude Models | Epoch AI3 savers
- 'The Odyssey' (2017) by Emily Wilson1 savers
- mimo-v2.6 RL6 savers
- When AI builds itself \ Anthropic34 savers
- Social Annealing « the jsomers.net blog3 savers
- Interview with Deepseek Founder: We're Done Following. It's Time to Lead.4 savers
- Codex_Seraphinianus1 savers
- Essays on programming I think about a lot | benkuhn.net11 savers
- Impact, agency, and taste | benkuhn.net30 savers
- Normalization_of_deviance1 savers
- [REPOST] Epistemic Learned Helplessness | Slate Star Codex13 savers
- Hawthorne effect2 savers
- Notes on the Industry Job Search17 savers
- Founder League1 savers
- Posts tagged "tech companies"3 savers
- Astroturfing1 savers
- Axonic4 savers
- IN THE WEIGHTS1 savers
- Silicon - Preorder the Book1 savers
- Macrodata Labs — every strong model starts with great data1 savers
- MAKING SOFTWARE61 savers
- AI Agent Benchmark for Real-World Professional Workflows1 savers
- [2606.05405] Agents' Last Exam1 savers
- store - teenage engineering1 savers
- Introducing Anthropic's AI for Science Program \ Anthropic2 savers
- Announcing Proximal · Proximal1 savers
- Our Problems · Proximal2 savers
- Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude | WIRED1 savers
- Ten Rules for Negotiating a Job Offer - haseeb qureshi4 savers
- Not_invented_here1 savers
- Goodput2 savers
- CL4R1T4S/ANTHROPIC/CLAUDE-FABLE-5.md at main · elder-plinius/CL4R1T4S1 savers
- Claude Fable 5 and Claude Mythos 5 \ Anthropic4 savers
- Noam Brown on X: "Implications of Large-Scale Test-Time Compute" / X7 savers
- John David Pressman's homepage2 savers
- Exodus — AI Talent Movement Tracker3 savers
- Paving the way for agents in biology \ Anthropic5 savers
- Computers can be understood14 savers
- Saint Grottlesex2 savers
- Does Muon improve regulatory DNA learning? Part 1. — Origin Bio1 savers
- Melanie Subbiah | PhD Candidate1 savers
- What Can't Deep Learning Do? – Bharath Ramsundar – Entrepreneur and Scientist1 savers
- Phil Wang · GitLab1 savers
- More articles we would like to commission2 savers
- Some Thoughts on Bengio's Scientist AI — LessWrong1 savers
- until9 savers
- Why senior developers fail to communicate their expertise | nair.sh1 savers
- Designing Compute Markets | Kavish Garg3 savers
- The Simple Habit That Saves My Evenings | alikhil | software engineering, kubernetes & self-hosting2 savers
- Developing creative identity16 savers
- refine2 savers
- A Conversation With Navvye Anand (Bindwell) | Rowan2 savers
- Speed matters: Why working quickly is more important than it seems « the jsomers.net blog66 savers
- Thirty Observations at Thirty62 savers
- Nat Friedman57 savers
- An Opinionated Guide to ML Research52 savers
- Details That Make Interfaces Feel Better43 savers
- No one can teach you to have conviction | benkuhn.net39 savers
- reflections on palantir - Nabeel S. Qureshi35 savers
- How To Scale Your Model34 savers
- Salary Negotiation: Make More Money, Be More Valued | Kalzumeus Software34 savers
- On-Policy Distillation - Thinking Machines Lab29 savers
- More people should write « the jsomers.net blog28 savers
- 101 things I would tell my self from 10 years ago28 savers
- Study Guide - LessWrong27 savers
- A Survival Guide to a PhD26 savers
- Teach Yourself Computer Science25 savers
- How to win a best paper award (or, an opinionated take on how to do important research)24 savers
- Beyond the Sky - Colossus23 savers
- Introducing talkie: a 13B vintage language model from 193020 savers
- Transformer Inference Arithmetic | kipply's blog17 savers
- Deriving Muon17 savers
- Alignment is not solved but it increasingly looks solvable17 savers
- Building the heap: racking 30 petabytes of hard drives for pretraining | blog17 savers
- Automated Weak-to-Strong Researcher16 savers
- Vibe physics: The AI grad student \ Anthropic15 savers
- A Recipe for Training Neural Networks15 savers
- Lessons from Peter Thiel | Posts | 8VC15 savers
- How to Build a $20 Billion Semiconductor Fab15 savers
- The Smol Training Playbook: The Secrets to Building World-Class LLMs - a Hugging Face Space by HuggingFaceTB13 savers
- On Really Trying · Gwern.net13 savers
- Diffusion is spectral autoregression – Sander Dieleman13 savers
- They're Made out of Meat12 savers
- 21 Facts About Throwing Good Parties12 savers
- First, Make Me Care, by Gwern · Gwern.net11 savers
- Do Ten Times as Much - by Bryan Caplan - Bet On It11 savers
- Performance Hints11 savers
- Operating Systems: Three Easy Pieces11 savers
- Debugging Reinforcement Learning Systems11 savers
- Recursive Language Models | Alex L. Zhang10 savers
- GPU Glossary10 savers
- Structured Procrastination10 savers
- MATS 9 Retrospective & Advice — LessWrong10 savers
- Career Decisions - by Elad Gil - Elad Blog10 savers
- All About Rooflines | How To Scale Your Model10 savers
- Economics and AI | Tom Cunningham – Tom Cunningham9 savers
- On Doubling9 savers
- Functions are Vectors9 savers
- Advanced topics in the theory of machine learning9 savers
- Things to buy | near.blog8 savers
highlights — 299
Fast tools don’t just allow users to accomplish tasks faster; they allow users to accomplish entirely new types of tasks, in entirely new ways.
Essays on programming I think about a lot | benkuhn.netLet’s say every company gets about three innovation tokens. You can spend these however you want, but the supply is fixed for a long while. You might get a few more after you achieve a certain level of stability and maturity, but the general tendency is to overestimate the contents of your wallet. Clearly this model is approximate, but I think it helps. If you choose to write your website in NodeJS, you just spent one of your innovation tokens. If you choose to use MongoDB, you just spent one of your innovation tokens. If you choose to use service discovery tech that’s existed for a year or le…
Essays on programming I think about a lot | benkuhn.netI’ve noticed a lot of people underestimate their own taste, because they expect having good taste to feel like being very smart or competent or good at things. Unfortunately, I am here to tell you that, at least if you are similar to me, you will never feel smart, competent, or good at things; instead, you will just start feeling more and more like everyone else mysteriously sucks at them. For this reason, the prompt I suggest here is: what does it seem like everyone else is mysteriously bad at? That’s probably a sign that you have good taste there.
Impact, agency, and taste | benkuhn.netA friend tells me of a guy who once accepted fundamentalist religion because of Pascal’s Wager. I will provisionally admit that this person “takes ideas seriously”. Everyone else gets partial credit, at best.
[REPOST] Epistemic Learned Helplessness | Slate Star CodexPosts tagged with "tech companies"
Posts tagged "tech companies"The word "grassroots" describes a natural, self-organizing movement run by the community. AstroTurf is a brand of artificial turf (grass); "astroturf" describes an artificial grassroots movement.
AstroturfingWhy Contribute - Help Set the Standard for Agent Evaluation in Your Industry
AI Agent Benchmark for Real-World Professional WorkflowsHistorical capability breakthroughs were the result of creative engineers discovering scalable data collection methods in specific domains, rather than thousands of contractors manually writing task demonstrations and graders.
Announcing Proximal · ProximalParticularly, we are excited about training algorithms such as SDPO that provide richer feedback than GRPO with scalar rewards. This problem can be tackled both on the algorithmic side as well from a data perspective: when purely using SDPO, what a model learns depends strongly on how the environment feedback is provided. We believe that a promising research direction is figuring out how to provide high-quality environment feedback for very complex tasks.
Our Problems · ProximalUseful coding data is spread across many public sources - GitHub alone has hundreds of millions of public repositories and more than 100M pull requests merged every year. Beyond this, GitLab, Bitbucket, public developer tool documentation, Stack Overflow, and other sources are extremely useful as seed data for data pipelines.
Our Problems · ProximalWith only high-level human input, Mythos 5’s trained model outperformed a recent model published in the journal Science—despite being 100 times smaller. We intend to publish these results in the coming months.
Claude Fable 5 and Claude Mythos 5 \ AnthropicAs he summarized, “The code was the easiest part! Most of the work was in the browser, clicking things.” Documentation kept telling him to “go to this URL, click on this dropdown.” His conclusion was that nobody should have to do this. Instead, we must build for agents. Karpathy had experienced something new within the world of software agents that biology researchers have been struggling with for a long time: the pain of trying to make intelligent systems operate in environments built around heterogeneous information, implicit conventions, and humans clicking through browsers.
Paving the way for agents in biology \ AnthropicTraining was limited to 20 epochs due to compute constraints; the loss has not begun to saturate (we have trained to much lower cross-entropy in prior runs), so these results reflect relative optimizer comparisons, not absolute model quality.
Does Muon improve regulatory DNA learning? Part 1. — Origin BioAt higher learning rates (∼ ), this rapid norm growth naturally cools the optimization, allowing MuonW to comfortably digest large initial steps and achieve the absolute best validation perplexity across all sweeps (∼2.464).
Does Muon improve regulatory DNA learning? Part 1. — Origin BioFor Adam variants, applying the Hyperball constraint improves peak performance but sharply narrows the stable operating window. AdamH achieves the absolute lowest overall perplexity for the Adam family (∼2.485) at a moderate learning rate of ∼ . However, AdamW proves dramatically more resilient at the extremes.
Does Muon improve regulatory DNA learning? Part 1. — Origin BioAs a check on judge bias, we ran the same test on a separate set of 127 moments where the human’s next move was already strong (as opposed to the original set, where the human’s direction had room for improvement). There, the models’ suggestions were judged better only about 20% of the time.
When AI builds itself \ AnthropicClaude is getting better at steering research sessions towards research findings. We examined real Claude Code sessions (between January and March 2026) where Anthropic researchers were working with Claude on an open-ended investigative problem, like figuring out why a training run kept crashing, or why a model scored poorly on a benchmark. In each case, we found a moment where the researcher took a detour: they pursued a direction that sent the session sideways before it eventually got back on track. We then showed various Claude models only the work from before the session went off-course an…
When AI builds itself \ AnthropicThe market he meant was perpetual futures. They were born out of an insight Robert Shiller, the economist, had in the ’90s. A traditional futures contract has an expiration date. When it arrives, a trader either takes delivery of the underlying asset—oil, wheat, pork bellies—or closes their position and opens a new one, paying fees each time. Shiller asked the obvious question: If almost nobody who trades a pork belly future wants pork bellies, why force the contract to expire?
Beyond the Sky - ColossusYan had seen enough. He told his team of six they were done trading. They may disagree, he said, but Chameleon was over. If he was wrong, they could always go back to trading. Several of them did disagree, and several would leave. But that didn’t change Yan’s mind. There were no investors to consult, no board to convince; it was his money and his call, and there was a new mission.
Beyond the Sky - ColossusYan liked HRT. He thought trading was the purest real-life game you could play. You were right or you were wrong and the market told you which. A lot of the smartest people in the world were competing against you, and in the process of playing this brutal game against each other
Beyond the Sky - ColossusTranslated By Emily Wilson
The Iliad | The Folio Society FictionWhenever Tony got really down like this, I would remind him of a clip that we both love, from the Studio Ghibli documentary “The Kingdom of Dreams and Madness.” It’s a moment when Hayao Miyazaki is trying to draw a specific airplane. And for some reason, he cannot do it. From “The Kingdom of Dreams and Madness” (2014) For days and days, he keeps trying to draw this plane, but nothing meets his satisfaction. Eventually, he realizes that he’s spent too much time on it. So he hands the plane off to another animator, and he moves on to something else. For me and for Tony, this clip is extremely re…
Postmortem: Every Frame a Painting | by Tony Zhou | MediumKeywords group everything in a really simple, visual way. This is how I figured out to cut from West Side Story to Transformers. From Godzilla to I, Robot. From Jackie Chan to Marvel films. On my screen, all of these clips are side-by-side because they share the same keyword.
Postmortem: Every Frame a Painting | by Tony Zhou | MediumCATALOG, a DNA computing company, synthesized and assembled millions of nucleotides of DNA into thousands of individual strands in their Boston laboratories. That DNA was then shipped to France, where Imagene, a company specializing in robust and room-temperature storage solutions, packaged the molecules into laser-sealed, stainless steel capsules. Each capsule was sealed under an inert atmosphere — meaning there is no oxygen or moisture inside the capsule — preserving the DNA inside for tens of thousands of years. And finally, Plasmidsaurus “read” the DNA book at their headquarters in Califor…
Asimov Press’ New Book, Written in DNAI thought, and still think, that Paxos is an important algorithm. Inspired by my success at popularizing the consensus problem by describing it with Byzantine generals, I decided to cast the algorithm in terms of a parliament on an ancient Greek island. Leo Guibas suggested the name Paxos for the island. I gave the Greek legislators the names of computer scientists working in the field, transliterated with Guibas's help into a bogus Greek dialect. (Peter Ladkin suggested the title.) Writing about a lost civilization allowed me to eliminate uninteresting details and indicate generalizations by …
The Writings of Leslie LamportMarlboro College, where I taught math from 1965-1969, had a weekly series of lectures for the general public, each given by a faculty member or an outside speaker invited by a faculty member. I gave a lecture about relativity that I later turned into this short monograph. I made a half-hearted, unsuccessful effort to get it published. But it was too short (75 pages) to be a "real" book, and there was very little interest in science among the general public in the late sixties. I think this monograph is still a very good exposition of the subject. Unfortunately, the second half, on general rela…
The Writings of Leslie LamportThe origin of this paper was the note The Maintenance of Duplicate Databases by Paul Johnson and Bob Thomas. I believe their note introduced the idea of using message timestamps in a distributed algorithm. I happen to have a solid, visceral understanding of special relativity (see [5]). This enabled me to grasp immediately the essence of what they were trying to do. Special relativity teaches us that there is no invariant total ordering of events in space-time; different observers can disagree about which of two events happened first. There is only a partial order in which an event e1 precedes…
The Writings of Leslie LamportI have long felt that, because it was posed as a cute problem about philosophers seated around a table, Dijkstra's dining philosopher's problem received much more attention than it deserves. (For example, it has probably received more attention in the theory community than the readers/writers problem, which illustrates the same principles and has much more practical importance.) I believed that the problem introduced in [41] was very important and deserved the attention of computer scientists. The popularity of the dining philosophers problem taught me that the best way to attract attention to…
The Writings of Leslie LamportMy favorite board games are Decrypto (a word game that’s about communication, that’s way deeper than Codenames), Smallworld (vaguely like Risk, but way more dynamic and fun), Cascadia (a surprisingly complex wildlife placing game), and Hanabi (team communication game). Fort is also great. I’ve been recommended Paleo, Terraforming Mars, Forbidden Desert, and 7 Wonders but haven’t tried them.
Things for Recovering Hoarders Like Me [Live Post]Wilde Chips - $6. These are made with chicken not potatos
Things for Recovering Hoarders Like Me [Live Post]Stannous Fluoride Toothpaste - $15. Stannous flouride and hydroxyapatite has repeatedly been shown to be more effective than normal toothpaste at cavity prevention.
Things for Recovering Hoarders Like Me [Live Post]All it takes is a change in the perception of what an acceptable environment looks like. So, fellow shapes, remember it's not about triangles vs squares, it's about deciding what we want the world to look like, and settling for no less.
Parable of the Polygons - a playable post on the shape of societyCube It moves in entertaining cube-like ways
Things to buy | near.blogAny time you have a problem, search for a solution. This sounds obvious, but it took me a long time to do for many obscure problems like “how do I stop sunlight from entering from under my door at 6AM in an aesthetic way”. With how good genAI is, there’s no excuse to not do this anymore. There are custom items for almost any problem now, and if there aren’t, it’s easy to have them made. Having items custom-printed, custom-welded, custom-cut, etc, is surprisingly cheap regardless of the medium.
Things to buy | near.blog"But instruction says never say can't because open."
bling on X: "leaked gpt-5.4 Pro CoT for unit distance problem run: "This sounds like an open problem? Need attempt solution perhaps prove or provide counterexample? We need determine from math maybe known primitive set bounds. Need be careful not to falsely claim theorem if unknown. But" / XBut instruction says never say can't because open.
bling on X: "leaked gpt-5.4 Pro CoT for unit distance problem run: "This sounds like an open problem? Need attempt solution perhaps prove or provide counterexample? We need determine from math maybe known primitive set bounds. Need be careful not to falsely claim theorem if unknown. But" / Xcuration of polygenic scores & genetic correlations, with utility weights from medical-economic research/surveys. Embryo-selection-as-a-service.
Startup Ideas · Gwern.netThe underlying reality must be the usual one: writing is like gardening. One patiently tends one’s garden, seeding and watering and pruning, and green shoots come up, and one day, one may behold a sudden blossoming, which one may cut and put in a vase to be seen by all. Or not, and let it wither and fall.
About This Website · Gwern.net (reader mode)When you contain the source of a thought, that thought can change along with you as you acquire new knowledge and new skills. When you contain the source of a thought, it becomes truly a part of you and grows along with you.
Truly Part Of You — LessWrongHow would I regenerate this knowledge if it were deleted from my mind?
Truly Part Of You — LessWrongWhen I first started reading HN, I didn’t understand 99% of the linked content, the jargon in the discussions, or the companies mentioned, but I was determined to learn. My general approach was to skim the top headlines and headlines that I could understand, read the post, and then read the HackerNews discussion. As I read the discussion, I would come upon terms I had no idea about: Big O notation, collaborative filtering, caching, Hindley-Milner type systems, lambda architectures, CI/CD pipelines, cryptography, generics, build or buy, B-trees, Bloom filters, trunk-based development, red/green…
20 years of YC | ✰Vicki Boykis✰Temporary uploads up to 1 GB are allowed. You should read the FAQ.
LitterboxDon't try to do projects by yourself. Some early PhD students fall into this trap, but you need to collaborate with people more senior than you. If it's your first project, it's honestly fine if they're doing things that shock and impress you—you'll learn so much from working with very smart people.
Employee Spotlight: Meet Katherine, AI Research Scientist | Pangram LabsIn the face of ambiguity, refuse the temptation to guess.
PEP 20 – The Zen of Python | peps.python.orgErrors should never pass silently.
PEP 20 – The Zen of Python | peps.python.orgOpenAI have 703 open jobs right now, of which I’d categorize 229 (32.6%) as relating to enterprise sales and support—account executives, “Go To Market”, “Forward Deployed Engineers” and the like. Anthropic have 390 open jobs, 105 (26.9%) of which look enterprisey to me.
I think Anthropic and OpenAI have found product-market fitLAB uses what Harvey calls “all-pass” grading, meaning that a task is marked complete only if every rubric criterion passes. There is no partial credit. The rationale is that a deal memo that catches eight of 10 material risks is not 80% useful. One missed issue could blow up the transaction or surface as a problem post-closing.
Some Thoughts On Harvey's Launch of 'LAB,' An Open-Source, Long-Horizon Benchmark for Legal AI Agents | LawSitesThis often helps you learn where your agent is misinterpreting or overindexing on a specific part of your prompt. Ask something like: "You were wrong. The answer was X. What would I need to have changed for you to get this right?" Obviously, what your agent will answer is not always the truth. Use it as a clue.
How to Eval AI Agents — The 2026 GuideFor multimodal tasks, at least, one person having fun over the course of a week can still make interesting benchmarks. Other examples of this: SpatialBench and IRGB. Arguably ARC-AGI and other videogame benchmarks fit this bill as well.
RIP Classic Reasoning Benchmarks. What’s Next?Your goal is to continue making progress towards a perfect GBA emulator indefinitely in this session, so if you believe you are done, you are mistaken and you must keep working and iterating.
gbaeval.com/blog/environment-design