Alex K
1 followers · 2 following · 236 views
on the atlas — 42
- FT1 savers
- How To Write With An LLM — A Final Ward2 savers
- Do Ten Times as Much - by Bryan Caplan - Bet On It11 savers
- Introducing Ship Beta1 savers
- Estimating the Productivity of an Autonomous AI Software Engineer | Cognition1 savers
- FT1 savers
- X1 savers
- FT1 savers
- Announcing Our Series B | Simile1 savers
- Linear Layouts: Robust Code Generation of Efficient Tensor Computation Using 𝔽_21 savers
- OpenAI Jalapeño: Better Than Nvidia Blackwell2 savers
- frontier-swe/tasks/postgres-sqlite-wire-adapter/instruction.md at main · Proximal-Labs/frontier-swe · GitHub1 savers
- The Verification Horizon: No Silver Bullet for Coding Agent Rewards1 savers
- Frontier-Bench2 savers
- Agent swarms and the new model economics · Cursor8 savers
- FT1 savers
- Precise Manipulation with Efficient Online RL5 savers
- Building a robotics research setup that lives next to my desk – dfdx labs1 savers
- Durability & Consistency - ZeroFS Documentation1 savers
- An Interview with MatX CEO Reiner Pope About LLM Chips2 savers
- Crawling a billion web pages in just over 24 hours1 savers
- Building a web search engine from scratch in two months with 3 billion neural embeddings15 savers
- Orbit - Ultra-efficient RL Pipeline1 savers
- Mirage Persistent Kernel: A Compiler and Runtime for Mega-Kernelizing Tensor Programs1 savers
- Kimi-K2.5 Inference Benchmark - Luminal1 savers
- Compiling Models to Megakernels - by Luminal and Joe Fioti1 savers
- Real-time LLM Inference on Standard Datacenter GPUs (3,000 tokens/s per request)1 savers
- State of Robot Learning, December 20257 savers
- Fara1.5 - A family of frontier computer use agent models - Microsoft Research1 savers
- Nemotron_Diffusion_Tech_Report_v1.pdf1 savers
- You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories1 savers
- SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost1 savers
- [2605.20873] PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models1 savers
- PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-Play1 savers
- Learnings from 100K Lines of Rust with AI | Cheng Huang’s corner1 savers
- What Makes Low-Bit Quantization-Aware Training Work for Reasoning LLMs? A Systematic Study1 savers
- Orthrus: Memory-Efficient Parallel Token Generation via Dual-View Diffusion1 savers
- The Matrix Calculus You Need For Deep Learning1 savers
- You don’t need to work on hard problems51 savers
- The case against conversational interfaces « julian.digital12 savers
- How to engineer luck - by George Mack - High Agency4 savers
- The Man Behind AlphaGo Thinks AI Is Taking the Wrong Path | WIRED2 savers
highlights — 12
First, given the rendered web page and the evaluation rubrics, an action planner generates a complete action list in a single pass, specifying the sequence of user interactions needed to exercise the target functionality. Second, a Playwright-based render server executes these actions in a live browser environment and records the resulting interaction trace, including screen recordings and state changes at each step. Third, a judge model evaluates sampled frames from the recordings together with the source code against the rubric criteria, producing the final score.
The Verification Horizon: No Silver Bullet for Coding Agent RewardsHowever, deploying a fully autonomous visual agent loop for judging is impractical under current constraints: multi-turn agent interactions incur high inference cost (He et al., 2026), and sequential decision-making introduces compounding errors that degrade evaluation stability. We therefore design a semi-automated agentic interactive judge that balances interaction coverage with efficiency and reliability.
The Verification Horizon: No Silver Bullet for Coding Agent RewardsBased on current trends Europeans could end up putting more fresh equity into American AI and satellite capabilities than European ones this year.
FTBecause the actor and critic operate on this compact representation, they can be represented with small networks that are trained directly on the robot, with hundreds of updates per second. This makes RL training responsive enough to improve the behavior after each attempt.
Precise Manipulation with Efficient Online RLThe training dynamics show why this separation matters. In the single-agent baseline, solve rates quickly saturate to near-perfect performance. This is not evidence of a strong curriculum. It means the proposer has found tasks its own solver can reliably solve. In PopuLoRA, solve rates oscillate instead of monotonically increasing.
PopuLoRA: Co-Evolving LLM Populations for Reasoning Self-PlayEnd-to-End User Story Execution: I still prefer to define the user stories myself. As an architect, I feel I have a better sense of what I’m building and how I’d like to build it. However, the delivery of a perfect execution is something I believe AI can handle increasingly well. Today, I still have to spend a fair amount of time steering the AI — telling it to continue when it pauses, suggesting refactoring, reviewing test coverage, and suggesting additional tests. I would prefer the AI take more autonomy to drive this end-to-end. Automated Contract Workflows: The flow of applying contracts s…
Learnings from 100K Lines of Rust with AI | Cheng Huang’s cornerIn my organization, there are no formal job descriptions or organizational charts. Responsibilities are defined in a general way, so that people are not circumscribed. All are permitted to do as they think best and to go to anyone and anywhere for help.
Doing a Job - The Management Philosophy of Adm. Hyman G. Rickoverno one involved in a job can divest himself of responsibility for its successful completion.
Doing a Job - The Management Philosophy of Adm. Hyman G. RickoverWeaknesses overlooked in oral discussion become painfully obvious on the written page.
Doing a Job - The Management Philosophy of Adm. Hyman G. RickoverI insist they report the problems they have found directly to me—and in plain English. This provides them unlimited flexibility in subject matter—something that often is not accommodated in highly structured management systems—and a way to communicate their problems and recommendations to me without having them filtered through others. The Defense Department, with its excessive layers of management, suffers because those at the top who make decisions are generally isolated from their subordinates, who have the first-hand knowledge.
Doing a Job - The Management Philosophy of Adm. Hyman G. RickoverThe man in charge must concern himself with details. If he does not consider them important, neither will his subordinates. Yet “the devil is in the details.” It is hard and monotonous to pay attention to seemingly minor matters. In my work, I probably spend about ninety-nine percent of my time on what others may call petty details.
Doing a Job - The Management Philosophy of Adm. Hyman G. RickoverWe still use our mouse to navigate and tell our computers what to do next, but routine actions are typically communicated in form of quick-fire keyboard presses: ⌘b to format text as bold, ⌘t to open a new tab, ⌘c/v to quickly copy things from one place to another, etc.
The case against conversational interfaces « julian.digital