Marvin Minsky? Seymour Papert? Apocryphal?

Question for Quote Investigator: The remarkable breakthroughs in artificial intelligence in the twenty-first century are based on digital neural networks. This area of machine learning was pioneered by Warren McCulloch and Walter Pitts who created an artificial neuron model. Also, researcher Frank Rosenblatt moved the field forward with his work on perceptrons.
The field suffered a massive blow when influential researchers proved that a simple perceptron network was incapable of learning an XOR function. Frank Rosenblatt attempted to extend his research to multilayer perceptrons which were more capable. Sadly, Rosenblatt died in an accident in 1971 when he was 43 years old.
Ultimately, multilayer perceptron networks became the foundation of deep learning and modern advances in artificial intelligence. Would you please explore what early critics and respondents said about multilayer networks?
Reply from Quote Investigator: In 1969 by Marvin L. Minsky and Seymour A. Papert published “Perceptrons: An Introduction to Computational Geometry”. A second printing with corrections was published in 1972. These authors highlighted the limitations of simple perceptron networks. They also mentioned with skepticism a “many-layered” version. Now, these systems are called “multilayer neural networks”. Boldface added to excerpts by QI:1
The problem of extension is not merely technical. It is also strategic. The perceptron has shown itself worthy of study despite (and even because of!) its severe limitations. It has many features to attract attention: its linearity; its intriguing learning theorem; its clear paradigmatic simplicity as a kind of parallel computation.
There is no reason to suppose that any of these virtues carry over to the many-layered version. Nevertheless, we consider it to be an important research problem to elucidate (or reject) our intuitive judgment that the extension is sterile. Perhaps some powerful convergence theorem will be discovered, or some profound reason for the failure to produce an interesting “learning theorem” for the multilayered machine will be found.
Below are additional selected citations in chronological order.
In 1986 a landmark collection of papers was published under the title “Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Volume 1: Foundations”. The eighth chapter was “Learning Internal Representations by Error Propagation” by David Rumelhart, Geoffrey Hinton, and Ronald J. Williams. The authors reprinted the key passage about multilayered perceptrons from the 1969 book. They prefaced the passage with the following statement:2
In their pessimistic discussion of perceptrons, Minsky and Papert (1969) finally discuss multilayer machines near the end of their book.
Rumelhart, Hinton, and Williams presented the following riposte:
Although our learning results do not guarantee that we can find a solution for all solvable problems, our analyses and results have shown that as a practical matter, the error propagation scheme leads to solutions in virtually every case. In short, we believe that we have answered Minsky and Papert’s challenge and have found a learning result sufficiently powerful to demonstrate that their pessimism about learning in multilayer machines was misplaced.
In 1988 Seymour Papert published an article in the journal “Daedalus” which discussed the intent of the 1969 book:3
Did Minsky and I try to kill connectionism, and how do we feel now about its resurrection? Something more complex than a plea is needed. Yes, there was some hostility in the energy behind the research reported in Perceptrons, and there is some degree of annoyance at the way the new movement has developed; part of our drive came, as we quite plainly acknowledged in our book, from the fact that funding and research energy were being dissipated on what still appear to me (since the story of new, powerful network mechanisms is seriously exaggerated) to be misleading attempts to use connectionist methods in practical applications.
Also, in 1988 Minsky and Papert published a new edition of their 1969 book. The authors added a chapter titled “Prologue: A View from 1988” which included the following comment:4
In preparing this edition we were tempted to “bring those theories up to date.” But when we found that little of significance had changed since 1969, when the book was first published, we concluded that it would be more useful to keep the original text (with its corrections of 1972) and add an epilogue, so that the book could still be read in its original form.
Minsky and Papert commented on the impact of their 1969 book:5
One popular version is that the publication of our book so discouraged research on learning in network machines that a promising line of research was interrupted. Our version is that progress had already come to a virtual halt because of the lack of adequate basic theories, and the lessons in this book provided the field with new momentum—albeit, paradoxically, by redirecting its immediate concerns.
In 1988 Minsky and Papert contrasted classic symbol-based approaches to machine learning and network-based machine learning. Interestingly, they saw great potential for both approaches:6
This is why we see no reason to choose sides. We expect a great many new ideas to emerge from the study of symbol-based theories and experiments. And we expect the future of network-based learning machines to be rich beyond imagining.
In 1994 David Freedman published “Brainmakers: How Scientists Are Moving Beyond Computers To Create a Rival To the Human Brain”. Freedman presented his opinion about the impact of the 1969 book:7
When Minsky and Papert described the perceptron’s inability to perform certain simple tasks, they were focusing on perceptrons without hidden units. As for neural networks with hidden units, the book essentially dismissed them — without careful analysis — as being of only marginally more promise and too difficult to train. “We consider it to be an important research problem to elucidate (or reject) our intuitive judgment that the extension is sterile,” sniffed Minsky and Papert in the book.
This brief statement was the most consequential one of the book. It said, in essence, that the only research project in neural networks worth carrying out was to prove the intuitively obvious (to Minsky and Papert’s thinking) point that multilayer neural networks were as big a waste of time as single-layer perceptrons.
In 2019 Melanie Mitchell published “Artificial Intelligence: A Guide for Thinking Humans”. Mitchell reprinted the key passage about multilayered perceptrons from the 1969 book. She prefaced the passage with the following statement:8
The limitations Minsky and Papert proved for simple perceptrons were already known to people working in this area. Frank Rosenblatt himself had done extensive work on multilayer perceptrons and recognized the difficulty of training them. It wasn’t Minsky and Papert’s mathematics that put the final nail in the perceptron’s coffin; rather, it was their speculation on multilayer neural networks.
In conclusion, Marvin Minsky and Seymour Papert deserve credit for the following provocative judgement about multilayered perceptrons in 1969: “the extension is sterile”. In 1988 the pair was still skeptical. They wrote: “little of significance had changed since 1969”.
Image Notes: Illustration depicting an abstract network of connections from geralt at Pixabay. Image has been cropped and resized.
Acknowledgement: Great thanks to the anonymous person whose inquiry led QI to formulate this question and perform this exploration.
- 1969 Copyright (1972 Corrections), Perceptrons: An Introduction to Computational Geometry by Marvin L. Minsky and Seymour A. Papert, Expanded Edition, Chapter 13: Perceptions and Pattern Recognition, Quote Page 231 and 232, The MIT Press, Cambridge, Massachusetts. (Verified with scans) ↩︎
- 1986 Copyright, Parallel Distributed Processing: Explorations in the Microstructure of Cognition, Volume 1: Foundations by David E. Rumelhart, John J. McClelland, and the PDP Research Group, Chapter 8: Learning Internal Representations by Error Propagation by D. E. Rumelhart, G. E. Hinton, and R. J. Williams, Start Page 318, Quote Page 361, A Bradford Book: The MIT Press, Cambridge, Massachusetts. (Verified with scans) ↩︎
- 1988 Winter, Daedalus: Journal of the American Academy of Arts and Sciences, Volume 117, Number 1, One AI or Many? by Seymour Papert, Start Page 1, Quote Page 4, American Academy of Arts and Sciences, Cambridge, Massachusetts. (Verified with scans) ↩︎
- 1988, Perceptrons: An Introduction To Computational Geometry by Marvin L. Minsky and Seymour A. Papert, Expanded Edition, Prologue: A View from 1988, Quote Page vii, The MIT Press, Cambridge, Massachusetts. (Verified with scans) ↩︎
- 1988, Perceptrons: An Introduction To Computational Geometry by Marvin L. Minsky and Seymour A. Papert, Expanded Edition, Prologue: A View from 1988, Quote Page xii, The MIT Press, Cambridge, Massachusetts. (Verified with scans) ↩︎
- 1988, Perceptrons: An Introduction To Computational Geometry by Marvin L. Minsky and Seymour A. Papert, Expanded Edition, Prologue: A View from 1988, Quote Page xiv, The MIT Press, Cambridge, Massachusetts. (Verified with scans) ↩︎
- 1994 Copyright, Brainmakers: How Scientists Are Moving Beyond Computers To Create a Rival To the Human Brain by David Freedman, Chapter 3: The Art of Thought, Quote Page 74, Simon & Schuster, New York. (Verified with scans) ↩︎
- 2019 Copyright (2020 paperback), Artificial Intelligence: A Guide for Thinking Humans by Melanie Mitchell, Chapter 1: The Roots of Artificial Intelligence, Quote Page 32, Picador, New York. (Verified with scans) ↩︎