TOPLOC: A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference
We introduce TOPLOC, a novel method for verifiable inference. TOPLOC employs a compact locality sensitive hashing mechanism for intermediate activations, which can detect unauthorized modifications to models, prompts, or compute precision with 100% accuracy in our empirical evaluations. The system maintains robustness across diverse hardware configurations, GPU types, tensor parallel dimensions, and attention kernel implementations, while achieving validation speeds up to 100× faster than the original inference. The polynomial encoding scheme used in TOPLOC reduces the memory overhead of generated commits by 1000×, requiring only 258 bytes of storage per 32 new tokens compared to the 262 KB required for storing token embeddings directly for Llama-3.1-8B-Instruct. This makes TOPLOC a practical solution for large-scale deployment. TOPLOC is easy to implement in modern inference engines and efficient enough to generate commits for all model inferences performed by an untrusted compute pro
TOPLOC - A Locality Sensitive Hashing Scheme for Trustless Verifiable Inference We introduce TOPLOC, a novel method for verifiable inference. TOPLOC employs a compact locality sensitive hashing mechanism for intermediate activations, which can detect unauthorized modifications to models, prompts, or compute precision with 100% accuracy in our empirical evaluations . The system maintains robustness across diverse hardware configurations, GPU types, tensor parallel dimensions, and attention kernel implementations, while achieving validation speeds up to 100× faster than the original inference. T
Explore this link on the map →related reading
- Defeating Nondeterminism in LLM Inference - Thinking Machines Labthinkingmachines.ai
- LLM Inference Performance Engineering: Best Practices | Databricks Blogdatabricks.com
- ⭐️ Fast LLM Inference From Scratchandrewkchan.dev
- Optimizing inference · Hugging Facehuggingface.co
- LLM-as-a-Verifier: A General-Purpose Verification Framework | alphaXivalphaxiv.org
- How is LLaMa.cpp possible?finbarr.ca
- The Compute Verification Postfirstscattering.com
- The Llama Hitchiking Guide to Local LLMs – hackerllamaosanseviero.github.io
- Mind the Trust Gap: Fast, Private Local-to-Cloud LLM Chat · Hazy Researchhazyresearch.stanford.edu
- Best practices to accelerate inference for large-scale production workloadstogether.ai
- Transformers Inference Optimization Toolset | AstraBlogastralord.github.io
- LLM Inference Economics from First Principlestensoreconomics.com