DOSSIER · №46 OF 81 · ACADEMIC · ch06
← →
← Back to The 81
46 · A SR ACADEMIC

"The Human Compatible"

Stuart Russell

CHAPTER
ch06 · Consciousness as Pattern Recognition
TIER
Academic
STATUS
Living · Active
ACTIVE
1962 – present
AFFILIATION
UC Berkeley · Center for Human-Compatible AI
JURISDICTION
Artificial Intelligence / AI Safety / Provably Beneficial AI
COLD OPEN
"The problem is not that we are building intelligent machines. The problem is that we are telling them exactly what we want. A machine that pursues a fixed objective does not coordinate with human values — it optimizes over them."
OVERVIEW

Author of the defining AI textbook used in universities worldwide, who then wrote its philosophical corrective: Human Compatible argues that the entire framework of specifying reward functions is wrong. Instead of telling AI what to optimize, we should build systems that remain uncertain about human preferences and coordinate with us to discover them over time.

WHY THIS VOICE MATTERS TO KNOWWARE
Russell formalizes the three-body structure that safe AI requires: the machine, the human, and the coordination loop between them that keeps values aligned. His inverse reward design is coordination intelligence applied to alignment — the machine learns what to want by coordinating with human feedback, not by receiving a fixed target.
OUTSTANDING NOTES
  • ◦IJCAI Computers and Thought Award
  • ◦ACM Fellow
  • ◦Time 100 AI (2023)
CLASSIFICATION
SR
ACADEMIC · TIER A
ch06
METHODS & FRAMEWORKS
5 entries
Standard AI textbookInverse reward designCAIS co-founderProvably beneficial AICorrigibility
KEY WORKS
  1. 01 Artificial Intelligence: A Modern Approach (with Norvig)
  2. 02 Human Compatible
  3. 03 Research Defense Initiative
← All 81
← 045 047 →