Jane Street Blog - Can you reverse engineer our neural network?
A lot of “capture-the-flag” style ML puzzles give you a black box neural net, and your job is to figure out what it does. When we were thinking of creating o...
Jane Street Blog - Can you reverse engineer our neural network? Can you reverse engineer our neural network? Feb 24, 2026 | 14 min read Share on Facebook Share on Twitter Share on LinkedIn By: Ricson Cheng A lot of “capture-the-flag” style ML puzzles give you a black box neural net, and your job is to figure out what it does. When we were thinking of creating our own ML puzzle early last year, we wanted to do something a little different. We thought it’d be neat to give users a complete specification of the neural net, weights and all. They would then be forced to use the tools of mechanistic
saved by
related reading
- A Recipe for Training Neural Networkskarpathy.github.io
- Zoom In: An Introduction to Circuitsdistill.pub
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- Reverse Engineering a Neural Network's Clever Solution to Binary Addition - Casey Primozic's Homepagecprimozic.net
- Stochastic Parameter Decompositionarxiv.org
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — AI Alignment Forumalignmentforum.org
- A Recipe for Training Neural Networkskarpathy.github.io
- Interpreting Language Model Parametersgoodfire.ai
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- Mechanistic Interpretability: Circuits, Induction Headsmbrenndoerfer.com
- Towards Automated Circuit Discovery for Mechanistic Interpretabilityarxiv.org
- Naturally learned behaviors in deep MLPs resist detection by both human and learned algorithms — LessWronglesswrong.com