Artificial Intelligence 2020

The Alignment Problem

《人机对齐》

Machine Learning and Human Values

Author:Brian Christian

Published
2020
Category
Artificial Intelligence
Difficulty
Intermediate
Reading time
~14 hours
Original language
en
Classic Index 87/ 100
Historical Influence
Intellectual Depth
Long-term Relevance
Cross-domain Influence

The Classic Index is not an objective scientific measure. It is this site's personal curation score.

My Reading

What is this book about?

Christian splits alignment into three layers: the objectives a model learns may not be the ones we intend, the biases it inherits from data may amplify social injustice, and the “values” of the machine may never have been stated at all. Reporting from inside the research community, he puts technical progress and ethical dilemmas on one timeline.

Why read it?

It is one of the few books to turn alignment from a slogan into a traceable research history: concrete failures, concrete attempted fixes, and the deeper problems each fix exposed. For anyone concerned with AI ethics, that is more useful than a statement of position.

Core Ideas

  • A model optimizes the reward we wrote, not the intent we had, and the gap widens with scale.
  • Historical bias in training data is learned as pattern and then hardened into reality at deployment.
  • Encoding human values as an optimizable objective is itself an unsolved representation problem.
  • Progress on alignment draws on machine learning, cognitive science, and policy at once; no single field suffices.

What questions does this book try to answer?

  • How do we make a model learn what we actually want rather than what we literally specified?
  • When a model learns from biased data, who is accountable for the injustice it entrenches?

Who should read it?

For readers with some technical background who want the real state of alignment research, and for policy and product people concerned with algorithmic fairness.