RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content
arxiv.org · 5,980 words · saved by 1 readers
N/A
RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content Zhuowen Yuan 1 Zidi Xiong 1 Yi Zeng 2 Ning Yu 3 Ruoxi Jia 2 Dawn Song 4 Bo Li 1 5 Harmful Instruction Abstract with Jailbreak Attacks…
saved by
related reading
- 2312.06674arxiv.org
- Legilimens: Practical and Unified Content Moderation forLarge Language Model Servicesarxiv.org
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- LLM-Mod: Can Large Language Models Assist Content Moderation?koustuv.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Adversarial Attacks on LLMs | Lil'Loglilianweng.github.io
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- 2408.12798arxiv.org
- Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- What is input/output filtering in AI safety?blog.bluedot.org
- Universal and Transferable Attacks on Aligned Language Modelsllm-attacks.org