Risk from fitness-seeking AIs: mechanisms and mitigations — LessWrong
lesswrong.com · saved by 3 readers
Current AIs routinely take unintended actions to score well on tasks: hardcoding test cases, training on the test set, downplaying issues, etc. This…