flâneur — a map of the web's best reading

Tencent HY Research

hy.tencent.com · 3 words · saved by 1 readers

CL-bench project page:www.clbench.com CL-bench paper link: download CL-bench paper CL-bench code repository:Tencent-Hunyuan/CL-bench CL-bench data repository:tencent/CL-bench Over the past few years, language models have become astonishingly capable. Frontier systems can now solve International Mathematical Olympiad problems, navigate complex coding challenges, and pass professional exams that take humans years to prepare for. They excel at taking tests, spinning out long chains of reasoning to win at benchmarks. But as impressive as these feats are, they obscure a simple truth: being a "test-taker" is not what most people need from an AI. Look at our own daily work. A developer skims documentation for a tool they've never seen and immediately starts debugging. A player picks up a rulebook for a new game and learns by playing. A scientist sifts through complex experimental logs to derive a new theorem from fresh data. In all these cases, humans aren't relying solely on a fixed body of

Explore this link on the map →

saved by