Legilimens: Practical and Unified Content Moderation forLarge Language Model Services
arxiv.org · 5,975 words · saved by 1 readers
N/A
Legilimens: Practical and Unified Content Moderation for Large Language Model Services Jialin Wu∗ Jiangyi Deng∗ Shengyuan Pang Zhejiang University Zhejiang University Zhejiang University Hangzhou, Zhejiang,…
saved by
related reading
- RigorLLM: Resilient Guardrails for Large Language Models against Undesired Contentarxiv.org
- LLM-Mod: Can Large Language Models Assist Content Moderation?koustuv.com
- 2025.naacl-long.441.pdfaclanthology.org
- 2312.06674arxiv.org
- [2502.05209] Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilitiesarxiv.org
- What We’ve Learned From A Year of Building with LLMs – Applied LLMsapplied-llms.org
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Uncensored Modelserichartford.com
- Productizing Large Language Modelsblog.replit.com
- 2025: The year in LLMssimonwillison.net
- Refusal in LLMs is mediated by a single direction — LessWronglesswrong.com
- Things we learned about LLMs in 2024simonwillison.net