✳flâneur — a map of the web's best reading
Our Problems · Proximal
proximal.ai · 1,361 words · saved by 1 readers
Proximal is a research lab for coding data. We build the data engine behind the next generation of autonomous coding agents.
We believe that training data is solved through creative engineering and research ideas. Here are some of the concrete technical problems we are working on right now: Index All Code on the Internet Useful coding data is spread across many public sources - GitHub alone has hundreds of millions of public repositories and more than 100M pull requests merged every year. Beyond this, GitLab, Bitbucket, public developer tool documentation, Stack Overflow, and other sources are extremely useful as seed data for data pipelines. Ideally, we would like to be able to run queries like this over this data:
Explore this link on the map →saved by
related reading
- Announcing Proximal · Proximalproximal.ai
- Cheap RL tasks will waste compute | Mechanize, Inc.mechanize.work
- PostTrainBenchposttrainbench.com
- Composer2.pdfcursor.com
- Good QC for RL Dataseancai.com
- RL Environments and RL for Science: Data Foundries and Multi-Agent Architecturesnewsletter.semianalysis.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- A Taxonomy of RL Environments for LLM Agentsleehanchung.github.io
- Adithya S K on X: "RL Coding Environments 101: Why Harbor Exists" / Xx.com
- Effective harnesses for long-running agents \ Anthropicanthropic.com
- IsoCompute Playbook: Optimally Scaling Sampling Compute for RL Training of LLMscompute-optimal-rl-llm-scaling.github.io
- Trust me bro, just one more RL scale up, this one will be the real scale up with the good environments, the actually legit one, trust me bro — AI Alignment Forumalignmentforum.org