Jane Street Blog - Can you reverse engineer our neural network?
A lot of “capture-the-flag” style ML puzzles give you a black box neural net, and your job is to figure out what it does. When we were thinking of creating o...
Jane Street Blog - Can you reverse engineer our neural network? Can you reverse engineer our neural network? Feb 24, 2026 | 14 min read Share on Facebook Share on Twitter Share on LinkedIn By: Ricson Cheng A lot of “capture-the-flag” style ML puzzles give you a black box neural net, and your job is to figure out what it does. When we were thinking of creating our own ML puzzle early last year, we wanted to do something a little different. We thought it’d be neat to give users a complete specification of the neural net, weights and all. They would then be forced to use the tools of mechanistic
Explore this link on the map →saved by
related reading
- A Recipe for Training Neural Networkskarpathy.github.io
- Zoom In: An Introduction to Circuitsdistill.pub
- Neural networks and deep learningneuralnetworksanddeeplearning.com
- Circuit Tracing: Revealing Computational Graphs in Language Modelstransformer-circuits.pub
- Transformer Circuits Threadtransformer-circuits.pub
- Reverse Engineering a Neural Network's Clever Solution to Binary Addition - Casey Primozic's Homepagecprimozic.net
- Interpreting Language Model Parametersgoodfire.ai
- A Recipe for Training Neural Networkskarpathy.github.io
- Fact Finding: Attempting to Reverse-Engineer Factual Recall on the Neuron Level (Post 1) — AI Alignment Forumalignmentforum.org
- Naturally learned behaviors in deep MLPs resist detection by both human and learned algorithms — LessWronglesswrong.com
- Interpretability — LessWronglesswrong.com
- Mechanistic Interpretability, Variables, and the Importance of Interpretable Basestransformer-circuits.pub