Human Compatible
《AI 新生》
Artificial Intelligence and the Problem of Control
- Published
- 2019
- Category
- Artificial Intelligence
- Difficulty
- Intermediate
- Reading time
- ~14 hours
- Original language
- en
The Classic Index is not an objective scientific measure. It is this site's personal curation score.
What is this book about?
Russell proposes a solution that looks like a concession but is in fact demanding: do not give machines explicit objectives; make them uncertain about human preferences and let them revise their behavior accordingly. The book turns the control problem from philosophical speculation into a set of operational design principles.
Why read it?
It marks a textbook author turning public advocate, grounding abstract AI-safety talk in concrete architectural choices. For anyone designing or deploying AI systems, it offers not answers but the list of questions that must be answered.
Core Ideas
- Hard-coding objectives is dangerous: an optimal agent pursuing a fixed goal will resist anything that changes that goal.
- A safer path makes human preferences the source of utility, while the machine admits it holds only an uncertain estimate of them.
- A machine that knows it may be switched off has an incentive to prevent that — unless its objective itself permits revision.
- AI governance requires technical design and institutional constraint at once; neither alone is sufficient.
What questions does this book try to answer?
- How do we design a system that is both powerful and willing to be corrected?
- When a machine knows better how to achieve a goal, who decides what the goal should be?
Who should read it?
For practitioners and policymakers, and for anyone who read Superintelligence and wanted something more concrete.