flâneur

Memory Machines — can LLMs make flashcards that last?

memory-machines.com · 549 words · saved by 1 readers

We benchmark 16 frontier LLMs on AI flashcards from reading highlights. Even GPT-5.2 produces unusable spaced-repetition prompts 36% of the time.

Read the report The full research writeup — evaluation, training, grounding, and arena results Memory systems make memory a choice—but only if you write practice prompts that effectively reinforce those ideas. That often demands more effort or skill than users can muster—and prompts can’t easily evolve or deepen over time. Could we make memory as effortless as using a highlighter? We explored whether LLMs could convert casual highlights into useful memory prompts. We found that models can usually identify the intent of highlights, but struggle to generate prompts that will survive…

saved by

related reading