✳flâneur — a map of the web's best reading
Al-Ekram Elahee Hridoy
0 followers · 3 following · 106 views
Open this reading profile →
on the atlas — 30
- MAI-Thinking-1: Building a Hill-Climbing Machine1 savers
- Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild | alphaXiv1 savers
- shareAI-lab/learn-claude-code: Bash is all you need - A nano claude code–like 「agent harness」, built from 0 to 1 ·1 savers
- pdf1 savers
- Your Evals Will Break and You Won't See It Coming - Lun Wang7 savers
- Lecture 01. Strong Models Don't Mean Reliable Execution | Learn Harness Engineering1 savers
- tencent/AutoCodeBenchmark · Datasets at Hugging Face1 savers
- Intent Formalization: A Grand Challenge for Reliable Coding in the Age of AI Agents | alphaXiv1 savers
- Intent Formalization: A Grand Challenge for Reliable Coding in the Age of AI Agents | alphaXiv1 savers
- MirrorCode: Evidence AI can already do some weeks-long coding tasks | Epoch AI1 savers
- [2605.02421] AOCI: Symbolic-Semantic Indexing for Practical Repository-Scale Code Understanding with LLMs1 savers
- Cognition | Multi-Agents: What's Actually Working1 savers
- BigO(Bench) -- Can LLMs Generate Code with Controlled Time and Space Complexity? | alphaXiv1 savers
- Multi-Teacher On-Policy Distillation: A New Post-Training Primitive | Notion1 savers
- Statistics for AI/ML, Part 4: pass@k and Unbiased Estimator1 savers
- Hidden Technical Debt of AI Systems: Agent Runtime1 savers
- How to Harness Coding Agents with the Right Infrastructure | Blog1 savers
- ZeroCoder: Can LLMs Improve Code Generation Without Ground-Truth Supervision? | alphaXiv1 savers
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks | alphaXiv1 savers
- Measuring what Matters: Construct Validity in Large Language Model Benchmarks | alphaXiv1 savers
- Frontier Coding Agents Can Now Implement an AlphaZero Self-Play Machine Learning Pipeline For Connect Four That Performs Comparably to an External Solver — LessWrong1 savers
- Code Generation and Repository-Level Software Engineering Benchmarks — A Field Guide to LLM Benchmarks | by Adnan Masood, PhD. | Medium1 savers
- AI excels at code competitions, struggles with real work1 savers
- [2501.01257] CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings1 savers
- CodeEvolve: an open source evolutionary coding agent for algorithmic discovery and optimization | alphaXiv1 savers
- Why AI-Generated Code Becomes Hard to Maintain and How to Fix It1 savers
- Surrender as a non-stupid life strategy11 savers
- Cursor and SpaceX: In search of a complete loop - kwokchain6 savers
- The ultimate guide to RL environments: building and scaling them in the LLM era - a Hugging Face Space by AdithyaSK5 savers
- Is Frontier Asynchronous RL Solved? — Luke J. Huang4 savers