flâneur

Constitutional AI: Harmlessness from AI feedback \ Anthropic

anthropic.com · saved by 1 readers

Constitutional AI trains a harmless but non-evasive assistant through self-critique and RL from AI feedback, with far fewer human labels.

saved by