Will reward-seekers respond to distant incentives? — LessWrong
Reward-seekers are usually modeled as responding only to local incentives administered by developers. Here I ask: Will AIs or humans be able to influ…
x Will reward-seekers respond to distant incentives? — LessWrong Fitness-seeking AIs Redwood Research AI Frontpage 57 Will reward-seekers respond to distant incentives? by Alex Mallen 16th Feb 2026 AI Alignment Forum 12 min read 4 57 Ω 33 Reward-seekers are usually modeled as responding only to local incentives administered by developers. Here I ask: Will AIs or humans be able to influence their incentives at a distance—e.g., by retroactively reinforcing actions substantially in the future or by committing to run many copies of them in simulated deployments with different incentives? If reward
Explore this link on the map →related reading
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward is not the optimization target — LessWronglesswrong.com
- Reward Is Not Enough — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- What failure looks like — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Recursive forecasting: Eliciting long-term forecasts from myopic fitness-seekers — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org