paper
Concrete Problems in AI Safety
A 2016 paper on practical research problems arising from machine-learning accident risk.
Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané authored this paper on harmful, unintended behavior in machine-learning systems.
One of its research topics is Reward Hacking. See the first submission for the historical date.
Sources
Pages that link here
Last updated 2026-10-07