Maanav Khaitan
68 followers · 54 following · 2860 views
on the atlas — 161
- goodalexander on X: "I don't think the tokenization story is well understood - so let's break it down Basically - the US Government is in vast, unprecedented amount of debt. And its long term debt is selling off a lot (40%+ over 5 years) - while Gold is up 140%. Gold has crossed US treasuries as" / X2 savers
- SEC.gov | Slumber Number: Innovation Exemption Statement1 savers
- 34-106402.pdf1 savers
- statement.pdf5 savers
- mindless toys2 savers
- LinkML at a glance - linkml documentation1 savers
- san francisco - musings11 savers
- goodalexander on X: "My take on society and crypto is really simple. If you have societal level investment, you are accountable for societal level returns. Things are not getting better. We said we have "AI warfare". We are locked in a stalemate with a 3rd world dictatorship who closed the Strait" / X1 savers
- Agentic Evals Pyramid • rwilinski.ai1 savers
- The Gervais Principle, Or The Office According to "The Office" — Ribbonfarm3 savers
- How should India approach AI? | Science2 savers
- Notes on the Software Factory | Benedict Brady2 savers
- Why Agentic Systems Need Ontologies — Frank Coyle, UC Berkeley - YouTube1 savers
- Black Hat USA 2026: The 'Breaking' News: The OpenAI–Hugging Face Incident - YouTube6 savers
- Billiards [Masquerade Award Official] - YouTube1 savers
- Building Effective AI Agents \ Anthropic8 savers
- Agent behavior2 savers
- Ankur Goyal on X: "Evals ≠ tests" / X1 savers
- LLMs reward expertise8 savers
- Robinhood Ventures Fund II FAQ | Robinhood1 savers
- synesthesia.inc1 savers
- 24/7 financial rails: How BNY plans to eliminate the weekend lag in U.S. Treasuries1 savers
- RWA.xyz | Analytics on Tokenized Real-World Assets1 savers
- 3 charts on the tokenized stocks boom - a16z crypto1 savers
- Game & Third-Party Tool Integration | exo docs1 savers
- Executors & Harnesses | exo docs2 savers
- Data Model | exo docs1 savers
- Hard Truths From the AI Trenches - Jason Liu1 savers
- - Your AI Product Needs Evals7 savers
- Training an Agentic Router for Optimal Cost-Performance on SWE Tasks | Applied Compute1 savers
- Unlocking Real-Time Bug Detection at Cognition | Applied Compute1 savers
- A frontier without an ecosystem is not stable3 savers
- Agent swarms and the new model economics · Cursor8 savers
- You're Spending Too Much on AI. You're Also Using Too Little. — Ramp Builders Blog3 savers
- Training a State-of-the-Art Legal Agent with Harvey | Applied Compute2 savers
- Neural Cheat Sheets: Learning to Summarize with Reinforcement Learning | Applied Compute1 savers
- Telos — Declarative software, managed for you1 savers
- [2507.19457] GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning4 savers
- Optimizing GEPA for production: A test-driven approach to prompt engineering | Decagon1 savers
- Memory in the wild: how we use Context Engine on our own code | Applied Compute1 savers
- Remember, Refine, Retrieve: A Context Engine for Enterprise Agents | Applied Compute1 savers
- Automating Merchant Onboarding at DoorDash | Applied Compute1 savers
- Metrics — Continual Learning Bench Docs1 savers
- Concepts — Continual Learning Bench Docs1 savers
- Continual Learning Bench 1.0 — News2 savers
- Senior SWE-Bench4 savers
- Tokenizing Succession - by Dennis Bouvard2 savers
- Making 768 servers look like 1 — PlanetScale8 savers
- Unlocking the Codex harness: how we built the App Server | OpenAI1 savers
- Harness engineering: leveraging Codex in an agent-first world | OpenAI9 savers
- Beyond rate limits: scaling access to Codex and Sora | OpenAI1 savers
- Unrolling the Codex agent loop | OpenAI2 savers
- [2606.09498] Self-Harness: Harnesses That Improve Themselves1 savers
- [2606.09498] Self-Harness: Harnesses That Improve Themselves1 savers
- Harness Engineering for Self-Improvement | Lil'Log19 savers
- Senior SWE-Bench2 savers
- Center Study: An Introduction1 savers
- The Making of Claude Code \ Anthropic3 savers
- Building self-improving tax agents with Codex | OpenAI2 savers
- Data Exchange | Center Study Center | Center Study Center1 savers
- Ray Dalio on X: "Investment Principles: What Should You Do Under Existing Conditions? " / X2 savers
- Agent vs Agent - Ishaan Singh2 savers
- AddyOsmani.com - Long-running Agents1 savers
- Deferral and Appropriation; Property and the Center1 savers
- The New World - Colossus7 savers
- Anchorage Digital Ventures on X: "Request for Startups | Anchorage Digital Ventures " / X1 savers
- Crocker's Rules6 savers
- Madeline Berzak on X: "memory, inevitability, and long-term risk communication: nuclear semiotics in the age of ASI" / X1 savers
- Paul Tudor Jones: The Greatest Macro Trader of All Time - YouTube1 savers
- Midcentury | The Private Data Platform for AI1 savers
- Get Your Reps - Alejandro García Salas2 savers
- Understanding Thomas Piketty’s Capital in the 21st Century | AIER1 savers
- Pravesh on X: "Art of Hiring" / X1 savers
- Are We Underestimating How AI Will Change Private Markets? - Colossus1 savers
- Memecoin Venture Capital - Bloomberg1 savers
- Exclusive | NYSE Partners With Securitize to Develop 24/7 Tokenized Securities Platform - WSJ1 savers
- Why Crypto Is Obsessed With AI Agents1 savers
- BankSouth Partners with Sequence Holdings to Define the Next Era of Tech-Enabled Community Banking | BankSouth1 savers
- Schwab moves giant step closer to taking proprietary private-market investments mainstream, today, by closing Forge deal; experts caution spending $660 million was the easy part | RIABiz1 savers
- shaunda devens on X: "HIP-3 Silver Microstructure: Hyperliquid vs. CME" / X1 savers
- Sid⚡️ on X: "Humanity's original store of value" / X1 savers
- Sid⚡️ on X: "The Willingness To Do Nothing" / X1 savers
- Bid Wall - MetaDAO1 savers
- MetaDAO1 savers
- Matt Shumer on X: "Something Big Is Happening" / X2 savers
- A Call for New Aesthetics13 savers
- Upgrading Ethereum | 2.3.4 Casper FFG2 savers
- Upgrading Ethereum | 2.3 Consensus2 savers
- S**t Umbrella – Zach Perret2 savers
- Core concepts, architecture and lifecycle | gRPC2 savers
- Semantics of Staking 2: Re-staking — The price of agency3 savers
- Semantics of Staking 1: Liquefaction — The price of agency2 savers
- Crypto's Three Body Problem2 savers
- Superlinear Returns11 savers
- Based rollups—superpowers from L1 sequencing - Layer 2 - Ethereum Research3 savers
- Curius / Onboarding2621 savers
- Looking for Alice - by Henrik Karlsson - Escaping Flatland90 savers
- How to Do Great Work80 savers
- Thirty Observations at Thirty62 savers
- Quit Your Job58 savers
highlights — 1844
The information is “in the model” already, but it takes a very smart human to pull it out.
LLMs reward expertiseTao makes several leaps and suggestions himself. He almost never takes the model’s advice about where to go next
LLMs reward expertiseBoth funds give investors access to private companies, but they focus on companies at different stages. RVI primarily invests in late-stage private companies—well-established businesses that haven't gone public yet. RVII will invest in early and growth-stage companies, with some recently founded. Earlier entry generally involves greater risk, but the potential for meaningful growth if the company succeeds.
Robinhood Ventures Fund II FAQ | RobinhoodBNY is already the primary custodian for RLUSD’s reserves and provides custody and investment management for OpenEden’s tokenized Treasury fund.
24/7 financial rails: How BNY plans to eliminate the weekend lag in U.S. TreasuriesUnlike stablecoins, whose circulating supply is a direct proxy for demand — one token, one dollar — tokenized stocks move with their underlying equities, so market cap does not cleanly separate the effects of new tokens minted and existing tokens repricing.
3 charts on the tokenized stocks boom - a16z cryptoThe evaluation it was extracted from layers self-improvement on the same pieces: a playbook the agent rewrites and re-reads every turn, durable memory files, and an install_tool meta-tool that lets the agent write new ES-module tools composing the emulator primitives. Over two long runs the agent authored 14 tools (dialog mashing, battle recovery, pathing macros) and 2 skills entirely on its own.
Game & Third-Party Tool Integration | exo docsThe payoff: sessions you can stop, resume, fork, and rewind across runs, regardless of which coding agent is driving.
Executors & Harnesses | exo docsThe exoharness stores and orders events durably; executors interpret them to construct message history and higher-level behavior.
Data Model | exo docsAI products often focus so much on capabilities that they forget fundamentals like stability, speed, and reliability.
Hard Truths From the AI Trenches - Jason LiuNobody cares that your RAG system has 78.9% recall. They care that FancyCorp saved $2M using your product.
Hard Truths From the AI Trenches - Jason LiuYou are doing it wrong if you aren’t looking at lots of data.
- Your AI Product Needs Evalsusing raw agreement is generally not recommended and can be misleading when classes are imbalanced. Instead, you should typically measure precision and recall separately to get a more accurate picture of your judge’s alignment.
- Your AI Product Needs EvalsAfter bringing the model-based evaluator in line with the human, you must continue doing periodic exercises to monitor the model and human agreement.
- Your AI Product Needs EvalsUse the most powerful model you can afford. It often takes advanced reasoning capabilities to critique something well.
- Your AI Product Needs EvalsHaving humans periodically evaluate at least a sample of traces is a good idea. I often find that “correctness” is somewhat subjective, and you must align the model with a human.
- Your AI Product Needs EvalsFurthermore, there are tools like Lilac which uses AI to search and filter data semantically, which is incredibly handy for finding a set of similar data points while debugging an issue.
- Your AI Product Needs EvalsWe decided to make the final output editable by a human so that we could curate & fix data for fine-tuning. These tools can be built with lightweight front-end frameworks like Gradio, Streamlit, Panel, or Shiny in less than a day.
- Your AI Product Needs Evalsyou don’t necessarily need a 100% pass rate. Your pass rate is a product decision, depending on the failures you are willing to tolerate.
- Your AI Product Needs EvalsIf you streamline your evaluation process, all other activities become easy. This is very similar to how tests in software engineering pay massive dividends in the long term despite requiring up-front investment.
- Your AI Product Needs EvalsThat means the routing problem is worth learning.
Training an Agentic Router for Optimal Cost-Performance on SWE Tasks | Applied ComputeThis is how multi-model systems should improve: not by treating model selection as a one-time benchmark comparison, but by continuously shaping both the router and the routed models.
Training an Agentic Router for Optimal Cost-Performance on SWE Tasks | Applied ComputeIf Nemotron 3 Ultra is already the oracle choice for a large fraction of tasks, it can be post-trained on the regions where it underperforms and move the entire frontier.
Training an Agentic Router for Optimal Cost-Performance on SWE Tasks | Applied ComputeWith the right training infrastructure, a dataset that faithfully represents the production task, and a reward function that encodes how the product actually needs to behave, enterprises can post-train a specialized model that occupies the exact point on the cost-latency-performance Parento frontier required.
Unlocking Real-Time Bug Detection at Cognition | Applied ComputeIf all the value is accrued by only a few models, the political economy will simply not tolerate it. There is no societal permission for an AI future that hollows out entire industries.
A frontier without an ecosystem is not stableThe companies that build this early will have an advantage that is hard to replicate
A frontier without an ecosystem is not stableSeen this way, the swarm starts to resemble a compiler. A compiler translates source code down to machine code through a series of intermediate steps. The swarm does something similar with intent.
Agent swarms and the new model economics · CursorWhat was scarce in this experiment, and what we expect to be scarce in software engineering going forward, is the right description of intent.
Agent swarms and the new model economics · CursorIn the run that used GPT-5.5 for both planners and workers, the workers alone cost $9,373. In the run where Opus 4.8 did the planning and Composer 2.5 did the work, the entire worker fleet cost $411.
Agent swarms and the new model economics · CursorIn the old run, the biggest files kept growing for the entire run and its single hottest file collected 7,771 conflicts, touched by 1,173 different agents. In the new run, the most contested file in the whole codebase saw 47.
Agent swarms and the new model economics · CursorTraining models to write for their successors, where better capture leads to better rewards, is an interesting follow-up area of research.
Agent swarms and the new model economics · CursorStigmergy is the mechanism by which swarm organisms like ants and termites coordinate without direct communication. They shape the environment, and the environment shapes the next organism.
Agent swarms and the new model economics · CursorThe new system peaks at around 1,000 commits per second.
Agent swarms and the new model economics · CursorSet benchmarks for repeated work. Do not ask teams to "use a cheaper model" in the abstract. Test which cheaper setup still produces acceptable outcomes.
You're Spending Too Much on AI. You're Also Using Too Little. — Ramp Builders BlogThe companies that manage AI costs best build a culture where efficiency is treated as an engineering achievement, not a constraint.
You're Spending Too Much on AI. You're Also Using Too Little. — Ramp Builders BlogIf a human is waiting for your agent to get back to them, it makes sense to pay for speed.
You're Spending Too Much on AI. You're Also Using Too Little. — Ramp Builders Blogthese graphs are not constructed by benchmarking models against the kind of work your business likely does. These models are benchmarked against brutally hard problems, intentionally.
You're Spending Too Much on AI. You're Also Using Too Little. — Ramp Builders BlogArtifact Completeness: The model learned to properly use tools and always create an output artifact. This behavior flipped 185 relevant rubric criteria from failing to passing during the course of training. Specificity and Exactness: The base GLM-5.1 model suffers from poor calculations or referring to imprecise numbers, often rounding figures during math (like 1.9 to 2), which is punished by legal graders. This behavior flips 243 criteria from failing to passing. Grounding: Without training, the checkpoint sometimes hallucinates source document items or invents findings from outside the provi…
Training a State-of-the-Art Legal Agent with Harvey | Applied Computereal memories need to be updated — often in surprisingly subtle ways — as new information comes in and old information becomes stale. A compelling next step would be to understand reward shaping to train for this functionality end to end as well.
Neural Cheat Sheets: Learning to Summarize with Reinforcement Learning | Applied ComputeThe same policy improves substantially over baseline and essentially matches the accuracy it reaches when QuALITY is included in the training mix. The model has learned how to take useful notes and amortize reasoning in a token-efficient manner, not domain-specific heuristics or knowledge.
Neural Cheat Sheets: Learning to Summarize with Reinforcement Learning | Applied ComputeIt is starkly reminiscent of the kind of frenetic, stream-of-consciousness notes a student might put on a 1-page cheat-sheet before an important exam:
Neural Cheat Sheets: Learning to Summarize with Reinforcement Learning | Applied Computelatent methods store large tensors for each token whereas summaries store only the token itself, and are thus much smaller and cheaper
Neural Cheat Sheets: Learning to Summarize with Reinforcement Learning | Applied ComputeThey close most of the gap toward a full KV cache which is several orders of magnitude larger, and not model-agnostic.
Neural Cheat Sheets: Learning to Summarize with Reinforcement Learning | Applied ComputeA summary is good context if it suffices to answer unseen QA pairs, which ask about fine-grained details, higher-level inferences, and subtext, in the original document. The QA pairs have a verifiable ground truth answer, and are posed as multiple-choice questions.
Neural Cheat Sheets: Learning to Summarize with Reinforcement Learning | Applied ComputeLatent vector stores, eg. learned KV caches. These high-capacity document-specific vectors often yield the best downstream performance when used as context, but are uninterpretable and model-specific.
Neural Cheat Sheets: Learning to Summarize with Reinforcement Learning | Applied ComputeWithout a durable specification, one-shot agents risk drifting away from the broader goal. The adversarial loop gives the system a way to pull itself back toward its desired state.
Telos — Declarative software, managed for youWithout a durable specification, one-shot agents risk drifting away from the broader goal. The adversarial loop gives the system a way to pull itself back toward its desired state. Checkpoint 3 One-shot 16% Telos loop 100%
Telos — Declarative software, managed for youWe built a custom instruction proposer that encodes the constraint directly into the reflection prompt.
Optimizing GEPA for production: A test-driven approach to prompt engineering | DecagonSpending on a frontier reflection model represents only 5-10% of total optimization cost, and a weak reflector means you waste all those task model calls learning nothing.
Optimizing GEPA for production: A test-driven approach to prompt engineering | Decagonsmaller models completely fail at prompt optimization
Optimizing GEPA for production: A test-driven approach to prompt engineering | DecagonWe observed a clear inverted-U relationship between sample size and performance
Optimizing GEPA for production: A test-driven approach to prompt engineering | Decagon