MMHal Bench
arxiv.org · 7,872 words · saved by 1 readers
N/A
Preprint A LIGNING L ARGE M ULTIMODAL M ODELS WITH FACTUALLY AUGMENTED RLHF Zhiqing Sun∗♠ , Sheng Shen∗♣ , Shengcao Cao∗ ♢ Haotian Liu♡ , Chunyuan Li♮ , Yikang Shen△ , Chuang Gan†∇△ , Liang-Yan Gui†♢ Yu-Xiong Wang†♢ , Yiming Yang†♠ , Kurt Keutzer†♣ , Trevor Darrell†♣ ♣ UC Berkeley, ♠ CMU, ♢ UIUC, ♡…
saved by
related reading
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- 2403.09611.pdfarxiv.org
- HALVA: Hallucination Attenuated Language and Vision Assistantresearch.google
- Patterns for Building LLM-based Systems & Productseugeneyan.com
- Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts | RLHFlowrlhflow.github.io
- Unfamiliar Finetuning Examples Control How Language Models Hallucinatearxiv.org
- Language Models Learn to Mislead Humans via RLHFarxiv.org
- Features as Rewards: Using Interpretability to Reduce Hallucinationsgoodfire.ai
- [2203.02155] Training language models to follow instructions with human feedbackarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- rlhfbook.com/book.pdfrlhfbook.com
- 2307.12950.pdfarxiv.org