TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering | OpenReview
As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks. However, there is no standard approach to evaluate tamper resistance. Varied data sets, metrics, and inconsistent threat settings make it difficult to compare safety, utility, and robustness across different models and defenses. To this end, we introduce TamperBench, the first unified framework to systematically evaluate the tamper resistance of LLMs. TamperBench (i) curates a repository of weight-space fine-tuning attacks and latent-space representation attacks; (ii) allows for testing state-of-the-art tamper-resistance defenses; and (iii) provides both safety and utility evaluations. TAMPERBENCH requires minimal additional code to specify any fine-tuning configuration, alignment-stage defense method, and metric suite while ensuring end-to-end reproducibility. In this work,
Verifying your browser | OpenReview Verifying your browser Complete the check below to continue to OpenReview Please complete the verification above. Have an OpenReview account? Sign in to skip this check.
Explore this link on the map →related reading
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- LLM evaluation: a beginner's guideevidentlyai.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2312.12575] LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarksarxiv.org
- The bitter lesson of LLM evalsparsed.com
- [2402.04249] HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusalarxiv.org
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org