Abstract: Minsky and Papert’s Perceptrons in 1969 mentions the limitations of the perceptron, namely the XOR problem, and this has often been cited as a contributing factor for the first AI winter. These problems have largely been overcome with MLPs / deeper networks, but recent focus on the Transformer architecture, while successful in many ways, introduces new limitations in what can be expressed by the Transformer architecture. These limitations echo the XOR problem in some interesting ways. I will briefly cover some of the present media coverage, history of AI research, then I will try to motivate why these limitations are important, and then give a quick overview of what the limitations entail, with a focus on an XOR-like problem. Finally, I will give some examples of what pitfalls to avoid if you're trying to use these models in your project.
Biography: Shawn is a research scientist at the MIT-IBM Watson AI Lab. He's been staring at neural networks for the past 10 years.