Don’t Let AI "Fight to the Death" for a Fixed Goal
In Human Compatible Russell noted that a fundamental risk of the traditional AI-system design paradigm is that once a system is set a concrete, fixed optimization goal, it may pursue that goal by any means, heedless even of deviation from humanity’s true intent. He advocated designing systems naturally uncertain about the goal they should pursue and thus always willing to correct by reference to humanity’s real-time feedback, an idea thereafter called "human-compatible" AI, one of the most influential theoretical frameworks in the AI-safety field.