YouTube3m· Nov 2024· cataloged

The Grand Theory of Intelligence – Gwern


What this covers

Full Episode: https://youtu.be/a42key59cZQ Transcript: https://www.dwarkeshpatel.com/p/gwern-branwen Apple Podcasts: https://podcasts.apple.com/us/podcast/gwern-branwen-how-an-anonymous-researcher-predicted/id1516093381?i=1000676859036 Spotify: https://open.spotify.com/episode/46H5dTtYaj1L55UAy9XXaY?si=e2f5f418f3994c62

Me on Twitter: https://twitter.com/dwarkesh_sp

Source description (no synthesized summary yet).

Sharpest takeaway

Intelligence is fundamentally search over Turing machines of varying length and complexity, not a unified 'fluid intelligence' or master algorithm; variation in human intelligence reflects differences in compute capacity for searching longer machine specifications.

  • All learning and scaling involve searching over progressively longer and more numerous Turing machine specifications
  • No localized 'IQ gland' exists because intelligence is distributed across learned specialized problem solutions recombined for fluid reasoning
  • Human general intelligence evolved slowly and rarely because direct genetic encoding of solutions is more efficient than expensive, glitchy search processes

The claims · ranked17 claims · weighted by value

This asset isn't compiled yet

You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.

0.73

The grand parsimonious theory of intelligence is that all intelligence is search over Turing machines of various lengths, and learning and scaling involve searching over more and longer Turing machines and applying them to specific cases.

definitionhigh valuecontestednovelty 3/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

the 10,000 foot view of intelligence that I think the successive scaling points to is that all intelligence is is search over Turing machines and I think anything that happens can be described by Turing Machines of various lengths and all that we're doing when we're doing learning or when we're doing scaling is that we're searching over more and longer turning machines and we're applying them in specific case

0.73

Human-level intelligence is rare and took a long time to evolve because small Turing machines could always be encoded more directly and efficiently by genes with sufficient evolution, making specialized hardcoded solutions superior to the costly, slow, unreliable search process that intelligence requires.

causalhigh valuecontestednovelty 3/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

I would actually just say that it helps explain why human level intelligence isn't such a great idea and so rare to evolve because any small turning machine could always be encoded more directly by your genes right with sufficient Evolution you have these organisms where like their entire neural network is just hardcoded by the genes so if you could do that obviously that's way better than some sort of colossally expensive Ive unreliable glitchy search process like what humans Implement right which takes whole days in some cases to learn whereas you know you it could be hardwired in right from birth

0.69

For many organisms and ecological niches, general-purpose intelligence is not adaptive because it does not pay to be intelligent; there are more direct, efficient solutions to environmental problems than general-purpose learning mechanisms.

causalhigh valueestablishednovelty 1/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

for many creatures like it just doesn't pay to be intelligent because that's not actually adaptive um there are better ways to solve the problem than a general purpose intelligence

0.69

All intelligence is fundamentally search over Turing machines, and anything that happens can be described by Turing machines of various lengths; learning and scaling involve searching over more and longer Turing machines and applying them to specific cases.

definitionhigh valuecontestednovelty 3/4durability 2/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

all intelligence is is search over Turing machines and I think anything that happens can be described by Turing Machines of various lengths and all that we're doing when we're doing learning or when we're doing scaling is that we're searching over more and longer turning machines and we're applying them in specific case

0.68

For many creatures, it does not pay to be intelligent because general-purpose intelligence is not actually adaptive; in static niches, in environments where intelligence is expensive, or for short-lived organisms, evolving a specialized tailor-made solution is superior to evolving a general-purpose learning mechanism.

causalhigh valuecontestednovelty 2/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

I think for many creatures like it just doesn't pay to be intelligent because that's not actually adaptive um there are better ways to solve the problem than a general purpose intelligence so in any kind of Niche where it's like static or where intelligence will be super expensive or where you don't have much time because you're a short-lived organism is going to be really hard to evolve a general purpose learning mechanism when you could instead evolve one that's just tailor made to the specific problem that you encounter

0.68

Intelligence took a long time to evolve in humans and remains rare because direct genetic encoding of solutions is always more efficient and adaptive than general-purpose search processes, which are expensive, unreliable, and slow (taking whole days to learn).

causalhigh valuecontestednovelty 2/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

any small turning machine could always be encoded more directly by your genes right with sufficient Evolution you have these organisms where like their entire neural network is just hardcoded by the genes so if you could do that obviously that's way better than some sort of colossally expensive Ive unreliable glitchy search process like what humans Implement right which takes whole days in some cases to learn whereas you know you it could be hardwired in right from birth

0.68

In ecological niches where the environment is static, where intelligence would be prohibitively expensive, or where organisms are short-lived, it is evolutionarily very difficult to evolve general-purpose learning mechanisms; specialized solutions tailored to specific problems are strongly favored instead.

causalhigh valuecontestednovelty 2/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

in any kind of Niche where it's like static or where intelligence will be super expensive or where you don't have much time because you're a short-lived organism is going to be really hard to evolve a general purpose learning mechanism when you could instead evolve one that's just tailor made to the specific problem that you encounter

0.63

Variation in human intelligence reflects differences in computational capacity for search over Turing machines; more intelligent people simply have more compute to search over more Turing machines for longer.

causalhigh valuecontestednovelty 2/4durability 2/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

it doesn't really feel like they've got this long tale of touring machines that they've learned uh how does this picture account for variation in human intelligence when we talk about more or less intelligence it's just that they have more compute in order to do search over more Turing machines for longer

0.63

The brain is a gigantic ensemble of small models tailored to solve the ever-escalating number of tiny problems encountered; large brain functions can be decomposed into these specialized models, just as large neural networks contain extractable small models for specific tasks.

causalhigh valuecontestednovelty 2/4durability 2/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

with a large neural network model you can always pull out kind of a small model which does a specific task equally well because that's all the large model is right it's just a gigantic Ensemble of small models tailored to the ever escalating number of tiny problems that You've been feeding them

0.61

There is no general Master algorithm and no special intelligence fluid; instead there are a tremendous number of special cases that we learn and then code into our brains.

factualhigh valuecontestednovelty 2/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

I think otherwise there's kind of you know there's no General Master algorithm and there's no special intelligence fluid it's just a tremendous number of special cases that we learn and then code into our brains

0.61

No IQ gland exists in the brain because there is nowhere that eliminates fluid intelligence when damaged; fluid intelligence emerges from the brain's learning of individual specialized problems and their recombination rather than from a dedicated neural substrate.

causalhigh valuecontestednovelty 2/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

we going to find any IQ gland there's nowhere in the brain where if you hit it you eliminate fluid intelligence I just think that you know it'll turn out that you know this doesn't exist because what your brain is doing is a lot of learning individual specialized problems and then once those individual problems are learned then they get recombined for fluid

0.61

No localized 'IQ gland' or anatomical structure exists in the brain that, if removed or damaged, would eliminate fluid intelligence, because fluid intelligence is not a unified property but rather the recombination of many learned specialized solutions.

causalhigh valuecontestednovelty 2/4durability 3/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

we going to find any IQ gland there's nowhere in the brain where if you hit it you eliminate fluid intelligence I just think that you know it'll turn out that you know this doesn't exist because what your brain is doing is a lot of learning individual specialized problems and then once those individual problems are learned then they get recombined for fluid

0.60

Large neural network models can always be decomposed into small specialized models, each solving specific tasks equally well, because the large model is itself just a gigantic ensemble of small task-tailored models addressing an escalating number of tiny problems.

factualhigh valueestablishednovelty 1/4durability 2/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

with a large neural network model you can always pull out kind of a small model which does a specific task equally well because that's all the large model is right it's just a gigantic Ensemble of small models tailored to the ever escalating number of tiny problems that You've been feeding them

0.56

There is no General Master algorithm and no special intelligence fluid; intelligence is instead composed of a tremendous number of special cases that organisms learn and then code into their brains.

causalhigh valuecontestednovelty 2/4durability 2/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

there's no General Master algorithm and there's no special intelligence fluid it's just a tremendous number of special cases that we learn and then code into our brains

0.56

Variation in human intelligence reflects differences in compute capacity; individuals with more intelligence have more compute in order to do search over more Turing machines for longer.

causalhigh valuecontestednovelty 2/4durability 2/4· Unidentified Speaker — The Grand Theory of Intelligence – Gwern [IMkqAUsVAfc]

when we talk about more or less intelligence it's just that they have more compute in order to do search over more Turing machines for longer

0.45

The interlocutor observes that smart friends seem to possess general horsepower or enhanced cognitive capacity rather than a collection of specialized Turing machines, which appears more compatible with the master algorithm view than the Turing machine search view.

factualhigh valuespeaker onlynovelty 1/4durability 2/4· Interlocutor

when I think about the way in which my smart friends are smart it kind of just feels like a more um like a general horsepower kind of thing right they've just got more juice and that seems more compatible with this masteral algorithm perspective whereas if with this touring machine perspective I don't know it doesn't really feel like they've got this long tale of touring machines that they've learned

0.39

The intuition that smart people have 'general horsepower' or greater 'juice' (fluid intelligence) is less compatible with the Turing machine search model and more compatible with a unified master algorithm perspective.

factualhigh valuespeaker onlynovelty 1/4durability 2/4· Unknown Speaker (Conversation Partner)

when I think about I don't know when I think about the way in which my smart friends are smart it kind of just feels like a more um like a general horsepower kind of thing right they've just got more juice and that seems more compatible with this masteral algorithm perspective whereas if with this touring machine perspective I don't know it doesn't really feel like they've got this long tale of touring machines that they've learned