paper

Concrete Problems in AI Safety

A 2016 paper on practical research problems arising from machine-learning accident risk.

Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan Mané authored this paper on harmful, unintended behavior in machine-learning systems.

One of its research topics is Reward Hacking. See the first submission for the historical date.

Sources

Last updated 2026-10-07