
Juergen Schmidhuber, Scientific Director of the Swiss AI Lab, speaking at the AI for Good Global Summit, at the International Telecommunication Union, Geneva, Switzerland, June 7-9, 2017.
العربية
1/2
Jürgen Schmidhuber (born January 17, 1963 in Munich) is a German and Swiss computer scientist, researcher, and artist.[83,84,92,99] He is prominent in the fields of artificial intelligence, deep learning, artificial neural networks, digital physics, and low-complexity art, with contributions including generalizations of Kolmogorov complexity and the Speed prior.[99,75]
Schmidhuber is best known for his foundational and highly cited work on Long Short-Term Memory (LSTM), a neural network architecture that became the dominant technique for various natural language processing tasks in research and commercial applications during the 2010s.[1,97,11,96,85] He also introduced principles of dynamic neural networks, meta-learning, generative adversarial networks, and linear transformers, all of which are widely used in modern AI.[33,32] Media outlets have described him as a leading AI pioneer and referred to him by various honorifics, including the father of modern AI, father of deep learning, godfather of AI, and papa of famous AI products, though he personally considers Oleksiy Hryhorovych Ivakhnenko to be the father of artificial intelligence.[7,8,2,5,3,4,6,11,18,19,20,32,92,116,117,69] According to Google Scholar data, his work received over 100,000 citations between 2016 and 2021.[1]
Schmidhuber completed his university education at the Technical University of Munich in Germany, where he served as a professor of cognitive robotics and taught from 2004 to 2009, while also directing its Cognitive Robotics Laboratory.[87,34,86,70,118] Regarding his tenure at the Università della Svizzera italiana in Lugano, Switzerland, most sources state that he became a professor of artificial intelligence there in 2009, whereas Mandarin Chinese Wikipedia states that he held this professorship from 2004 to 2009. Since 1995, he has served as scientific director, co-director, or deputy director at the Dalle Molle Institute for Artificial Intelligence Research (IDSIA) in Manno, Lugano District, Canton of Ticino, Switzerland.[1,17,87,75] On October 1, 2021, he joined King Abdullah University of Science and Technology (KAUST) in Saudi Arabia as director of the Artificial Intelligence Initiative and professor of computer science in the Computer, Electrical, and Mathematical Sciences and Engineering (CEMSE) division.[1,92]
Education and Career
Jürgen Schmidhuber was born in Munich to Hans Schmidhuber and Karolin (née Sickinger), and is married with two children. Inspired by science fiction, particularly Arthur C. Clarke and Isaac Asimov, he began to study computer theory, logic, and the algorithmic structuring of knowledge.[21] After graduating from high school, he studied computer science and mathematics at the Technical University of Munich from 1983, earning his diploma in 1987.[31] He earned his doctorate in computer science at the Technical University of Munich in 1991 under Wilfried Brauer and Klaus Schulten, focusing on dynamic neural networks and the fundamental spatiotemporal learning problem.[22,31,34,35] Dynamic neural networks and fast weight programmers proposed by him in 1991 contained core ideas of the modern Transformer architecture.[23] He worked as a postdoctoral researcher at the University of Colorado Boulder in 1991–1992 before completing his habilitation at the Technical University of Munich in 1993 on net architectures, objective functions, and the chain rule.
Schmidhuber worked as a senior assistant and Privatdozent at the Technical University of Munich, where he also headed the Cognitive Robotics Laboratory as an associate professor from 2004 to 2009. Since 1995, he has served as the scientific director of the Dalle Molle Institute for Artificial Intelligence Research (IDSIA) in Lugano.[85] From 2003 to 2021, he was a professor at the Scuola universitaria professionale della Svizzera italiana in Manno.[17,24] Sources differ on his tenure as professor of artificial intelligence at the Università della Svizzera italiana in Lugano, with German sources stating he was a full professor from 2009 to 2024 and remains an adjunct professor, while English, Galician, and Albanian sources state he held the position from 2009 to 2021.[17,24,5,86] Since October 2021, he has served as the director of the AI Initiative at King Abdullah University of Science and Technology (KAUST).
Schmidhuber, along with students such as Sepp Hochreiter, Felix Gers, Fred Cummins, and Alex Graves, developed advanced versions of Long Short-Term Memory (LSTM) recurrent neural networks. Initial LSTM results were presented in Hochreiter's 1991 diploma thesis analyzing the vanishing gradient problem, the term was formally introduced in 1995, and the standard architecture was introduced in 2000. Bloomberg Businessweek noted that LSTM is 'the most commercial AI achievement, used in everything from predicting diseases to composing all types of music.' In 2011, his team at IDSIA with Dan Cireșan achieved a fourfold GPU acceleration of convolutional neural networks compared to central processing units, winning at least four competitions.
In 2014, Schmidhuber co-founded NNAISENSE to develop commercial AI applications in finance, heavy industry, and self-driving cars, serving as its chief scientist and as president from 2014 to 2017.[17] Sepp Hochreiter, Jaan Tallinn, and Marcus Hutter served as advisors to NNAISENSE, which raised capital funding in January 2017 despite having sales under $11 million in 2016 due to a focus on research over revenue.[9,10,17] Schmidhuber's goal is to create artificial general intelligence by sequentially training a single AI on diverse tasks, though as of 2026 he stated that NNAISENSE's focus had shifted to asset management.[11,34,36]
Research and Contributions
Schmidhuber has published extensively across machine learning, neural networks, Kolmogorov complexity, digital physics, robotics, low-complexity art, and the theory of beauty. In the 1980s, standard backpropagation performed poorly in deep learning tasks involving long credit assignment paths. To address this issue, Schmidhuber proposed a hierarchy of recurrent neural networks (RNNs) in 1991 that were pre-trained level-by-level via self-supervised learning, using predictive coding to learn internal representations across multiple time scales.[36] This hierarchy could be collapsed into a single RNN by distilling a higher-level chunker network into a lower-level automatizer network, allowing a chunker network to solve a deep learning task exceeding a depth of 1000 by 1993.[34,38] Recurrent neural networks developed in his research group efficiently solved previously intractable tasks, including context-sensitive language recognition, robot control in partially observable environments, music composition, handwriting recognition, and speech processing. Schmidhuber refers to his RNNs incorporating Long Short-Term Memory (LSTM) as deep learning networks.[26] In 1991, Schmidhuber supervised the diploma thesis of his student Sepp Hochreiter at TU Munich—which Schmidhuber considered one of the most important documents in machine learning history—analyzing and overcoming the vanishing gradient problem using the neural history compressor.[25,41,100] This work led to the creation of Long Short-Term Memory (LSTM) networks, with the name first appearing in a 1995 technical report prior to the widely cited 1997 paper co-authored by Hochreiter and Schmidhuber.[25,42,101] The standard LSTM architecture was introduced in 2000 by Felix Gers, Schmidhuber, and Fred Cummins, followed by 'vanilla LSTM' using backpropagation through time in 2005 with Alex Graves.[43,46,40,102,103,104] In 2006, the Connectionist Temporal Classification (CTC) training algorithm was introduced and applied to end-to-end speech recognition using LSTM.[46,40]
LSTM networks became foundational for Google's smartphone speech recognition system starting in 2015, enabling continuous training on sounds and words using brain-like network structures with additional connections that preserve sound over time. Beyond speech recognition, Google integrated LSTM into Google Translate, the smart assistant Allo, and AI research such as DeepMind's AlphaGo, with one of DeepMind's founders having studied under Schmidhuber in Lugano.[106,105,34] LSTM was subsequently adopted by Apple for Siri and iPhone QuickType, Amazon for Alexa, and Facebook for executing roughly 4.5 billion daily automated translations in 2017.[107,108,109,110,111,43,42,36,119,105,120,106] Bloomberg Businessweek characterized LSTM as perhaps the most commercial AI achievement, applied to diverse domains ranging from medical disease prediction to music composition.[11] To overcome training accuracy degradation when stacking 20 to 30 layers in deep networks, Rupesh Kumar Srivastava, Klaus Greff, and Schmidhuber introduced the Highway Network in May 2015 using LSTM principles, which enabled hundreds of layers and led to the Residual Neural Network (ResNet) variant in December 2015.[51,47,48,49,35,54,55,50,53,56,52] Schmidhuber also introduced fast weights programmers in 1992—where a slow feedforward network controls fast weights of another network via outer products, later shown equivalent to the unnormalized linear transformer—and proposed adversarial neural networks with artificial curiosity in 1991, anticipating Generative Adversarial Networks.[34,39,35,37,57,58]
In 2011, Schmidhuber's team at IDSIA, led with postdoc Dan Ciresan, achieved major accelerations of deep convolutional neural networks (CNNs) on graphics processing units (GPUs).[60] While earlier GPU implementations achieved 4x speedups over CPUs, Ciresan's team achieved a 60x speedup, winning the first superhuman performance in a computer vision contest in August 2011 and winning four more image competitions through September 2012.[62,64,59,10,63,112,114,113,115,121,122,111,11,123] While these fast and deep GPU-accelerated CNNs became central to computer vision, English sources credit their underlying design to earlier work by Kunihiko Fukushima, whereas Chinese sources attribute it to Yann LeCun et al.[60,61] In theoretical computer science, Schmidhuber proposed the Gödel Machine in 2003, a theoretical construct that uses an asymptotically optimal theorem prover to rewrite its software whenever it proves doing so will improve future performance.[27] He explored digital physics and computable universes, implementing Konrad Zuse's 1967 hypothesis via the 'Great Programmer' and showing that the simplest program computes all universes. His work on non-halting converging programs led to generalizations of Kolmogorov complexity, limit-computable probability measures, Super-Omegas, and optimal inductive inference, though sources differ on whether the associated metric is named the 'Speed Prior' or 'speed matters'.
In 2014, Schmidhuber co-founded Nnaisense to commercialize AI technologies across finance, heavy industry, and autonomous vehicles, with advisors including Sepp Hochreiter, Jaan Tallinn, and Marcus Hutter.[63,64] Nnaisense raised its first external funding round in January 2017, maintaining a focus on research over revenue after reporting under $11 million in sales in 2016. Schmidhuber's long-term goal is to build universal artificial intelligence by sequentially training a single network on multiple narrow tasks, though skeptics observe that long-standing industry efforts by companies like IBM and Arago have yet to yield strong AI. His achievements earned him the INNS Helmholtz Award in 2013, the IEEE CIS Neural Networks Pioneer Award in 2016, and the NVIDIA Pioneers of AI Research Award for his lab in 2016. In 2023, Elon Musk publicly remarked 'Schmidhuber invented everything,' a phrase frequently cited in biographies and lectures to illustrate his legacy in artificial intelligence.[28]
Credit Disputes
Schmidhuber has controversially argued that he and other researchers have been denied adequate recognition for their contributions to the field of deep learning, in favor of Geoffrey Hinton, Yoshua Bengio, and Yann LeCun, who shared the 2018 Turing Award for their work in deep learning. In a scathing 2015 article, he argued that Hinton, Bengio, and LeCun "heavily cite each other" but "fail to credit the pioneers of the field."[98,67]
In a statement to The New York Times, Yann LeCun wrote that "J2rgen is manically obsessed with recognition and keeps claiming credit he doesn't deserve for many, many things... It causes him to systematically stand up at the end of every talk and claim credit for what was just presented, generally not in a justified manner." Schmidhuber replied that LeCun made this statement "without any justification, without providing a single example", and published details of numerous priority disputes with Hinton, Bengio, and LeCun.[67,68,65]
The term "schmidhubered" has been jokingly used in the artificial intelligence community to describe Schmidhuber's habit of publicly challenging the originality of other researchers' work, a practice seen by some as a "rite of passage" for young researchers. Some suggest that Schmidhuber's significant accomplishments have been underappreciated due to his confrontational personality.[34,66]
Recognition and Awards
In 2013, Jürgen Schmidhuber received the Helmholtz Award from the International Neural Network Society, and in 2016 he received the Neural Networks Pioneer Award from the IEEE Computational Intelligence Society for "pioneering contributions to deep learning and neural networks."[15,16,13,14,34,126] He is a member of the European Academy of Sciences and Arts.[15,16,13,14,75,76,34,77]
Schmidhuber has been referred to as the "father of modern AI", the "father of generative AI", and the "father of deep learning".[6,71,72,70] However, Schmidhuber himself has called Alexey Grigorevich Ivakhnenko the "father of deep learning", giving credit to many even earlier AI pioneers.[41,74] The New York Times ran a profile under the headline "When A.I. Matures, It May Call Jürgen Schmidhuber 'Dad'", highlighting his early work on deep learning and his long-term vision for self-improving AI.[34,69]
Views and Opinions
Since the 1970s, Schmidhuber has aimed to create intelligent machines capable of learning and improving on their own to become smarter than him within his lifetime. He differentiates between two types of AI: tool AI, such as applications for improving healthcare, and autonomous AIs that set their own goals, perform their own research, and explore the universe. Having worked on both types for decades, he expects the next stage of evolution to be self-improving AIs that will succeed human civilization as the next stage in the universal increase towards ever-increasing complexity, and he expects AI to colonize the visible universe.[35] Additionally, Schmidhuber is a proponent of open-source AI and believes that open-source models will become competitive against commercial closed-source AI.
In 2025, Schmidhuber dated the technological singularity—the convergence point of human inventions—to the year 2042.[30,29] Furthermore, as a consequence of what he views as inevitably advancing automation and the accompanying loss of gainful employment jobs, Schmidhuber sees the necessity of an unconditional basic income.
According to reports by The Guardian, in a scathing 2015 article, Schmidhuber complained that fellow deep learning researchers Geoffrey Hinton, Yann LeCun, and Yoshua Bengio heavily cite each other while failing to credit the pioneers of the field.[12,124] He claimed that they understated the contributions of Schmidhuber himself and other early machine learning pioneers, including Alexey Grigorevich Ivakhnenko, who published the first deep learning networks in 1965.[124] Yann LeCun denied this accusation, stating instead that Schmidhuber continues to claim credit he does not deserve. Schmidhuber objected to LeCun's statement, claiming that LeCun did not support his assertion with examples, and listed many of his own achievements to refute it.[125]
Schmidhuber's Hypothesis
The German computer scientist's most important contribution to digital philosophy is the so-called "Schmidhuber hypothesis", set forth in the 1997 paper A Computer Scientist's View of Life, the Universe, and Everything.[89] The starting point of the discussion begins with the idea that a long time ago, the Great Programmer wrote a program that launched all possible—meaning computable—universes in his Great Computer.[90] Our Universe is likewise launched by the Great Programmer, and its state can be described by a low number of bits, reflecting a theory of beauty based on the concept of simplicity.
Despite this regularity, whereby slices of buttered bread falling are always turned toward the floor rather than toward the ceiling, Schmidhuber notes that deviations also exist.[91] These deviations lead many to believe that the world is partially random and therefore incomputable. However, according to Schmidhuber, our inability to decode the state of our Universe does not affect the possibilities of the Great Programmer, who can instead examine it at any moment.
.jpg)