"The Bitter Lesson": The More General, the More Compute-Reliant Method Often Wins Last
In the essay Sutton reviewed the development of several AI subfields such as computer chess, speech recognition, and computer vision, and found a recurring pattern: researchers at first always tend to encode human expertise and intuition into the system, indeed making progress in the short term, but in the long run the methods that give up reliance on human prior knowledge and instead do general search or learning with stronger compute and more data almost always overtake and do better. This observation is seen by many researchers as an important theoretical basis explaining why the "scaling path" of later large language models could overwhelm several earlier, more "ingenious" methods.