Constitutional AI: Harmlessness from AI feedback \ Anthropic
anthropic.com · saved by 1 readers
Constitutional AI trains a harmless but non-evasive assistant through self-critique and RL from AI feedback, with far fewer human labels.
Constitutional AI trains a harmless but non-evasive assistant through self-critique and RL from AI feedback, with far fewer human labels.