flâneur

RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content

arxiv.org · 5,980 words · saved by 1 readers

N/A

RigorLLM: Resilient Guardrails for Large Language Models against Undesired Content Zhuowen Yuan 1 Zidi Xiong 1 Yi Zeng 2 Ning Yu 3 Ruoxi Jia 2 Dawn Song 4 Bo Li 1 5 Harmful Instruction Abstract with Jailbreak Attacks…

saved by

related reading