2312.06674
arxiv.org · 7,108 words · saved by 1 readers
N/A
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, Madian Khabsa GenAI at Meta We introduce Llama Guard, an LLM-based input-output safeguard model geared towards Human-AI…
saved by
related reading
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Contentarxiv.org
- Legilimens: Practical and Unified Content Moderation forLarge Language Model Servicesarxiv.org
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- gpt-4.pdfcdn.openai.com
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- What is input/output filtering in AI safety?blog.bluedot.org
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- GitHub - salesforce/AuditNLG: AuditNLG: Auditing Generative AI Language Modeling for Trustworthinessgithub.com
- 2405.01470arxiv.org
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com