How good are slop-vestigators? — LessWrong
lesswrong.com · saved by 2 readers
TLDR: • 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI a…
TLDR: • 1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI a…