"The problem is not that we are building intelligent machines. The problem is that we are telling them exactly what we want. A machine that pursues a fixed objective does not coordinate with human values — it optimizes over them."
Author of the defining AI textbook used in universities worldwide, who then wrote its philosophical corrective: Human Compatible argues that the entire framework of specifying reward functions is wrong. Instead of telling AI what to optimize, we should build systems that remain uncertain about human preferences and coordinate with us to discover them over time.