TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering | OpenReview
As increasingly capable open-weight large language models (LLMs) are deployed, improving their tamper resistance against unsafe modifications, whether accidental or intentional, becomes critical to minimize risks. However, there is no standard approach to evaluate tamper resistance. Varied data sets, metrics, and inconsistent threat settings make it difficult to compare safety, utility, and robustness across different models and defenses. To this end, we introduce TamperBench, the first unified framework to systematically evaluate the tamper resistance of LLMs. TamperBench (i) curates a repository of weight-space fine-tuning attacks and latent-space representation attacks; (ii) allows for testing state-of-the-art tamper-resistance defenses; and (iii) provides both safety and utility evaluations. TAMPERBENCH requires minimal additional code to specify any fine-tuning configuration, alignment-stage defense method, and metric suite while ensuring end-to-end reproducibility. In this work,
Verifying your browser | OpenReview Verifying your browser Complete the check below to continue to OpenReview Please complete the verification above. Have an OpenReview account? Sign in to skip this check.
related reading
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- A Safe Path to Open Weights - Thinking Machines Labthinkingmachines.ai
- [2506.17209] Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Fine-Tuning Lowers Safety and Disrupts Evaluation Consistencyarxiv.org
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Contentarxiv.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Announcing Safety Research Grantsthinkingmachines.ai
- LLM evaluation: a beginner's guideevidentlyai.com