LSTM: The Comeback Story? [Prof. Sepp Hochreiter]

From the creator

SPONSOR MESSAGES: *** CentML offers competitive pricing for GenAI model deployment, with flexible options to suit a wide range of models, from small to large-scale deployments. Check out their super fast DeepSeek R1 hosting! https://centml.ai/pricing/ Sepp Hochreiter, the inventor of LSTM networks, makes a forceful case that large language models are fundamentally limited -- they store and retrieve human knowledge rather than reason about it. He walks through the technical history from the original 1991 vanishing gradient problem through to his latest work on xLSTM, explaining how exponential gating and matrix memory replace the old sigmoid bottleneck to create architectures that can compete with Transformers on sequence modeling while maintaining fixed memory footprints. The conversation covers the practical advantages of xLSTM for industrial applications like robotics simulation and discrete element modeling, where constant memory consumption matters more than raw benchmark performance. Hochreiter also discusses the neuro-symbolic Pi AI project at JKU Linz, his view that genuine reasoning requires something beyond pattern matching on training data, and why he believes the current scaling paradigm will hit a wall. The technical deep-dive into FlashAttention comparisons, recurrent vs attention-based architectures, and the evolution from sigmoid to exponential gating mechanisms gives a clear picture of where LSTM-style architectures fit in the current landscape. --- REFERENCES: Paper: [00:00:13] Long Short-Term Memory (Original LSTM Paper) https://direct.mit.edu/neco/article-abstract/9/8/1735/6109/Long-Short-Term-Memory [00:04:18] Kolmogorov Complexity https://link.springer.com/article/10.1007/BF02478259 [00:19:38] xLSTM: Extended Long Short-Term Memory https://arxiv.org/abs/2405.04517 [00:22:53] FlashAttention: Fast and Memory-Efficient Attention https://arxiv.org/abs/2205.14135 [00:36:08] Mamba: Linear-Time Sequence Modeling https://arxiv.org/abs/2312.00752 [00:55:03] Core Knowledge Theory of Human Cognition https://www.harvardlds.org/wp-content/uploads/2017/01/SpelkeKinzler07-1.pdf Book: [00:53:18] Thinking, Fast and Slow https://www.amazon.com/Thinking-Fast-Slow-Daniel-Kahneman/dp/0374533555 Company: [01:00:18] NXAI: Industrial AI with xLSTM https://investinaustria.at/en/blog/nxai-top-researchers-in-austria-develop-solutions-for-industrial-companies/ --- LINKS: Full Transcript: https://app.rescript.info/share/50e3f79d116b1a92aa40819b6cb40eb2 Download PDF transcript: https://app.rescript.info/api/public/sessions/bdb8e291a0e4514d/pdf Prof. Sepp Hochreiter https://www.nx-ai.com/ https://x.com/hochreitersepp https://scholar.google.at/citations?user=tvUH3WMAAAAJ&hl=en

Choose to Build with AI
Matched to Neural Networks

AI Maker Residence 3

The third AI workshop taught by our legendary teacher, Nick Sarafa. This is a full-day hands-on training workshop for purposeful co-creation with AI using Claude Code. Imagine having access to hundreds of billions of dollars of computing power and knowing exactly how to make it work for you through the power of super intelligence.

◆ Fri 09 Oct 2026 ◆ KOKO Cafe, London ◆ With Nick Sarafa
AI Maker Residence 3
Live event
AI Maker Residence 3
Fri 09 Oct 2026

More like this

Running one yourself?

List your AI event,
wherever it is.

A meetup, a workshop, a hackathon, a conference. Any city, or online. Tell us about it and it lands in front of people already learning this stuff.