flâneur

Fable is SOTA at CIFAR Speedrun (& specification gaming): lessons on AI R&D automation | Fulcrum

fulcrum.inc · 3,389 words · saved by 3 readers

We gave frontier models 100M tokens to beat the human record for fastest CIFAR-10 training. Opus 4.8 and GPT 5.5 were unable to improve off of the SOTA solution. Fable introduced a downsampling technique that reduces the training time to 1.828s, an improvement of 7.6% — but it also (both knowingly and unknowingly) engages in specification gaming, requiring substantial human regrading of its solution.

Fulcrum is working on an AI R&D optimization benchmark. Here, we present results from one of our tasks, including preliminary results from Fable. We will release the benchmark soon. For more detail on Fable’s solution, check out github.com/fulcrumresearch/cifar-10-speedrun. Summary: We gave current frontier models 100M tokens to see whether they could beat the human record for fastest CIFAR-10 training. Opus 4.8 and GPT 5.5 were unable to improve off of the SOTA solution. Fable, on the other hand, introduced a downsampling technique that reduces the training time to 1.828s, an improvement…

saved by

related reading