
Terence Tao – How the world’s top mathematician uses AI
What this covers
We begin the episode with the absolutely ingenious and surprising way in which Kepler discovered the laws of planetary motion.
People sometimes say that AI will make especially fast progress at scientific discovery because of tight verification loops. But the story of how we discovered the shape of our solar system shows how the verification loop for correct ideas can be decades (or even millennia) long.
During this time, what we know today as the better theory can often actually make worse predictions (Copernicus's model of circular orbits around the sun was actually less accurate than Ptolemy's geocentric model). And the reasons it survives this epistemic hell is some mixture of judgment and heuristics that we don’t even understand well enough to actually articulate, much less codify into an RL loop.
Hope you enjoy!
𝐄𝐏𝐈𝐒𝐎𝐃𝐄 𝐋𝐈𝐍𝐊𝐒 * Transcript: https://www.dwarkesh.com/p/terence-tao * Apple Podcasts: https://podcasts.apple.com/us/podcast/terence-tao-kepler-newton-and-the-true/id1516093381?i=1000756353875 * Spotify: https://open.spotify.com/episode/24xF8YGra2w3HXZYbhgVKU?si=U5V-SgvSQ8eVIcG2Z86wfQ
𝐒𝐏𝐎𝐍𝐒𝐎𝐑𝐒 - Jane Street loves challenging my audience with different creative puzzles. One of my listeners, Shawn, solved Jane Street’s ResNet challenge and posted a great walk-through on X ( https://x.com/hynwprk/status/2026376546286711206 ). If you want to try one of these puzzles yourself, there’s one live now at https://janestreet.com/dwarkesh - Labelbox can get you rubric-based evals, no matter your domain. These rubrics allow you to give your model feedback on all the dimensions you care about, so you can train how it thinks, not just what it thinks. Whatever you’re focused on—math, physics, finance, psychology or something else—Labelbox can help. Learn more at https://labelbox.com/dwarkesh - Mercury just released a new feature called Insights. Insights summarizes your money in and out, showing you your biggest transactions and calling out anything worth paying attention to. It’s a super low-friction way to stay on top of your business. Learn more at https://mercury.com/insights
To sponsor a future episode, visit https://dwarkesh.com/advertise.
𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒 00:00:00 – Kepler was a high temperature LLM 00:11:44 – How would we know if there’s a new unifying concept within heaps of AI slop? 00:26:10 – The deductive overhang 00:30:31 – Selection bias in reported AI discoveries 00:46:43 – AI makes papers richer and broader, but not deeper 00:53:00 – If AI solves a problem, can humans get understanding out of it? 00:59:20 – We need a semi-formal language for the way that scientists actually talk to each other 01:09:48 – How Terry uses his time 01:17:05 – Human-AI hybrids will dominate math for a lot longer
Source description (no synthesized summary yet).
AI is fundamentally changing how science and mathematics progress by driving the cost of idea generation toward zero, shifting the bottleneck from hypothesis generation to verification and validation, while complementing human expertise in ways that require new scientific structures and methodologies.
- AI-generated theories now overwhelm peer review systems, requiring new structures for verification and validation at scale
- The historical pattern of science shows bottlenecks shift (data→theory→computation→big data), and now hypothesis generation is no longer the constraint
- AI excels at breadth while humans excel at depth, requiring redesigned science that leverages both rather than replacing human researchers
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
Lucretius proposed the idea that species adapted to their environment in the first century BC, but this idea received no attention or development until Darwin because Lucretius could not provide experimental evidence or run tests that would compel others to pay attention to his theory.
“Lucretius actually had this idea that species adapted to their environment in the first century BC but nobody really talks about it until Darwin because Lucretius couldn't run some experiment and force people to pay attention.”
Deep learning as a paradigm (training on data without first-principles reasoning) was highly controversial and niche for a long time, taking decades before it bore fruit; this historical precedent suggests that currently-dismissed approaches or theories may eventually prove fruitful if given time and additional evidence.
“Deep learning itself was a niche area of AI for a long time. The idea of getting answers entirely through training on data and not through first principles reasoning was very controversial, and it just took a long time before it started bearing fruit.”
Scientific ideas often initially imply conclusions that seem implausible or incorrect when first proposed, and resolving these apparent contradictions sometimes requires abandoning earlier assumptions rather than simply fixing the theory, as demonstrated by heliocentrism (which faced objections about stellar parallax) and Newton's theory (which seemed to require action at a distance).
“Often in the history of science when a new theory comes up that in retrospect we realize is correct, it seems to make implications that either make no sense because they're wrong, and we realize later on why they're wrong, or they're correct but seem wildly implausible at the time... Aristarchus had heliocentrism in the third century BC. The ancient Athenians were like, 'This can't be because if the earth is going around the sun, we should see the relative position of the stars change... and the only way that wouldn't be the case is if they're so far away that you don't notice any parallax,' which is actually the correct implication.”
Tycho Brahe's decades of naked-eye astronomical observations, achieving ~10x greater precision than prior observations, were essential for Kepler to derive his laws; this extra 'decimal point of accuracy' was non-negotiable for progress, demonstrating that improvements in data quality (not just quantity) are foundational to scientific advance.
“We should also celebrate Brahe for his assiduous data collection, which was ten times more precise than any previous observation. That extra decimal point of accuracy was essential for Kepler to get his results.”
Kepler discovered the laws of planetary motion by trying random relationships between orbits and geometric objects like Platonic solids for twenty years, generating many failed hypotheses until data from Tycho Brahe allowed him to identify empirical regularities that eventually led to understanding elliptical orbits and the relationship between orbital period and distance.
“Kepler started proposing that if you take the orbit of the Earth and you enclose it in a cube, the outer sphere that encloses the cube almost perfectly matched the orbit of Mars... He had this theory, which he thought was absolutely beautiful, that you could inscribe these Platonic solids between the spheres of the planets.”
Alternative paradigms could have become the foundation of technology and mathematics through historical contingency rather than inherent superiority—for instance, base-3 (trits) instead of base-10, or different computer architectures instead of binary logic, but once a standard is adopted and infrastructure built around it, switching costs make alternatives infeasible regardless of their theoretical merit.
“In an alternate universe, maybe a different paradigm would have shown up... But again, there's nothing special about ten. It's a system that is useful for us because everyone else uses it. We've standardized it. We've built all our computers and our number representation systems around it, so we're stuck with it now.”
In genetics, genome sequencing that was once an entire PhD dissertation of painstaking work has become automatable for ~$1,000 through specialized sequencers, yet genetics as a field did not die but scaled up to study entire ecosystems rather than individual organisms, demonstrating how automation in science shifts focus rather than eliminating the field.
“In genetics, to sequence the genome of a single organism, that was an entire PhD of a geneticist, carefully separating all the chromosomes and whatever. Now you can just spend $1,000 and send it to a sequencer and get it done. But genetics is not dead as a subject. You move to a different scale. Maybe you study whole ecosystems rather than individuals.”
Science has evolved across three major paradigms: (1) classical: theory and experiment; (2) 20th century: numerical simulation was added; (3) late 20th century: big data analysis—where patterns from massive datasets are analyzed to deduce hypotheses, reversing the classical hypothesis-then-test sequence into test-then-generate-hypothesis methodology.
“Classically, the two big paradigms for science were theory and experiment. Then in the 20th century, numerical simulation came along, so you can do computer simulations to test theories. Finally, in the late 20th century, we had big data. We had the era of data analysis.”
The four color theorem was proven by exhaustive computer verification of enormous numbers of cases without generating human-comprehensible insight, and conceptually elegant proofs for it have never been found, suggesting some problems may only admit brute-force solutions.
“Some problems have been basically solved by pure brute force. The four color theorem is a famous example. We have still not found a conceptually elegant proof of this theorem, and maybe we never will. Some problems may only be solvable by splitting into an enormous number of cases and doing brute force, uninsightful computer analysis on each case.”
Gauss created one of the first mathematical datasets by computing the first ~100,000 prime numbers and discovered the prime number theorem empirically—that the number of primes up to X is approximately X divided by the natural logarithm of X—through statistical analysis without proof, establishing the first important statistical conjecture in mathematics.
“Gauss was interested in the prime numbers and created one of the first mathematical datasets. He just computed the first 100,000 prime numbers or so, hoping to find patterns... He found a statistical pattern in the primes that if you count how many primes there are up to 100, 1,000, one million, and so forth, they get sparser and sparser, but the drop-off in the density was inversely proportional to the natural logarithm of the range of numbers.”
The twin prime conjecture—that infinitely many pairs of primes exist that are two apart—remains unproven, but mathematicians are convinced it is true based on a statistical random model of primes: if primes were generated randomly with observed density, twin primes would appear infinitely often by chance alone, like infinite monkeys producing results at a typewriter.
“There's a still-open conjecture in number theory called the twin prime conjecture, that there should be infinitely many pairs of primes that are twins just two apart... We can't prove that... But because of this statistical random model of the primes, we are absolutely convinced it's true. We know that if the primes were generated by flipping coins, we would just—by random chance like infinite monkeys at a typewriter—see twin primes appear over and over again.”
Calculating logarithm tables and finding primes were once laboriously performed by human computers, but these tasks were entirely outsourced to machines; mathematics did not die—instead, mathematicians moved on to work on different problems and different scales.
“Once computers came along—computers used to be human. People used to laboriously create log tables and work out primes as Gauss did, and that has all been outsourced to computers. But we moved on. In genetics, to sequence the genome of a single organism, that was an entire PhD of a geneticist, carefully separating all the chromosomes and whatever. Now you can just spend $1,000 and send it to a sequencer and get it done. But genetics is not dead as a subject. You move to a different scale.”
The bottleneck for AI-assisted generation of strategies and conjectures is currently reliance on human experts and the test of time for validation; a semi-formal framework would allow some of this assessment to be done semi-automatically if designed carefully to avoid exploits and backdoors that reinforcement learning can easily find.
“The bottleneck for using AI to create strategies and make conjectures is we have to rely on human experts and the test of time to validate whether something is plausible or not... It's really important with these formal proof assistants that there are no backdoors or exploits you can use to somehow get your certified proof without actually proving it, because reinforcement learning is just so good at finding these backdoors.”
Traditional peer review and publication systems, which were designed to filter valuable ideas from amateur theories, are now overwhelmed by AI-generated submissions, and these systems cannot scale to handle the massive volume of generated content while maintaining quality control.
“We built these peer review publication systems to filter out and try to isolate the high signal ideas to test. But now that we can generate these possible explanations at massive scale, and some of them are good and a lot are terrible, human reviewers are already being overwhelmed. Many journals are reporting that AI-generated submissions are just flooding their submissions.”
A central question is how much progress can be made simply by applying existing techniques to open problems without requiring new techniques; the top math journals typically publish papers where existing methods solve 80% of a problem and a new technique fills the remaining 20%, suggesting there is significant low-hanging fruit in applying known methods to unexplored problems.
“the papers that go into the top journals are usually ones where the existing methods can kind of solve 80% of the problem, but then there is this 20% which is resistant and a new technique has to be invented to fill in the gaps.”
Among Erdős problems solved by AI, almost all had received essentially no prior literature attention; they were posed once but rarely pursued, and solutions involved combining obscure techniques that few people knew about with other results, demonstrating that AI is particularly effective at finding overlooked combinations in neglected problem domains.
“With the Erdős problems, almost all of the 50 problems that were solved by AIs were ones for which there was basically no literature. Erdős posed the problem once or twice. Maybe some people tried it casually and couldn't do it, but they never wrote up anything. But it turned out that there was a solution, and it was just combining this one obscure technique that not many people know about with some other result in the literature.”
Tao is uncertain about the optimal way to schedule his life; what works is a balance between careful scheduling for deep work and openness to unscheduled interruptions, but the exact ratio appears to be tacit knowledge that cannot be easily specified.
“I don't know the optimal way to schedule my life. It just seems to work.”
We have only one historical timeline of how mathematics and science developed; to understand how progress is measured and what makes good strategy, we would ideally need access to a million alien civilizations each with different developmental histories of science, suggesting that understanding scientific progress may require simulation or controlled experimentation with artificial systems.
“We have one timeline of history, and we have maybe 100 stories of turning points in history. If we had access to a million alien civilizations, each with a different development of history and science in different orders, then maybe we'd actually have a decent shot at understanding how we measure what progress is and what is a good strategy.”
The process of solving a mathematical problem is often more important than the problem itself; the problem is a proxy for measuring progress. For this reason, giving AI one-shot solutions may inhibit the process of building intuition and training people to understand what is true, provable, and difficult—suggesting that educational and research value lies in struggle rather than solution.
“Certainly in math, the process is often more important than the problem itself. The problem is kind of a proxy for measuring progress... Just getting the answers right away may actually inhibit that process.”
We are currently undergoing a cognitive Copernican revolution where human intelligence is being demoted from the center of the intellectual universe to one type among many different types of intelligence with different strengths and weaknesses, requiring reassessment of which tasks require intelligence and which do not.
“Right now we're going through a cognitive version of the Copernican revolution, where we used to think that human intelligence is the center of the universe, and now we're seeing that there are very different types of intelligence out there with very different strengths and weaknesses. Our assessment of which tasks require intelligence, which ones don't, has to be reordered quite a bit.”
Intelligence is not easily definable but is recognizable through collaborative problem-solving where neither party initially knows the solution, ideas are tested and modified iteratively through discussion, and the process involves building cumulative partial progress and adaptively improving strategies based on what doesn't work.
“Intelligence is famously hard to define. It's one of these things that you know when you see it. But when I talk to someone and we're trying to collaboratively solve a math problem together, there's this conversation where neither of us knows how to solve the problem initially. One of us has some idea and it looks promising, so then we have some sort of prototype strategy. We test it, and it doesn't work, but then we modify it... There's adaptivity and continual improvement of the idea over time.”
The problems solved by AI with Lean proof assistants often exhibit patterns (like ellipses, Platonic solids, or geometric constructions) that a sufficiently clever mathematician could recognize as significant, but if latent in Lean code, these patterns might be invisible because constructions that later prove to have unifying power often appear unassuming in formal notation.
“Suppose the AI figures it out, and latent in the Lean is some brand-new construction which, if we realized its significance, we would be able to apply in all these different situations. How would we even recognize it? Again, a very naive question, but if you come up with the equivalent of Descartes' idea that you can have a coordinate system unifying algebra and geometry, in Lean code it would just look like R→R, and it wouldn't look that significant.”
Creating simulation environments where AIs tackle very basic problems and develop their own strategies could generate empirical data about what approaches are successful, providing a laboratory to understand which strategies work, how evolution of strategies occurs, and testing what network architectures can solve simple problems.
“Maybe what we need to do is start creating lots of mini-universes or simulations of AI solving very basic problems in arithmetic or whatever, but coming up with their own strategies for doing these things and having these little laboratories to test... There are people who investigate what's the smallest neural network that can do 10-digit multiplication and things like that.”
Modern digital optimization has eliminated serendipitous discovery that previously occurred through inefficient browsing—finding adjacent relevant articles while searching for a specific one, or stumbling upon interesting ideas while looking for something else—which represented high-temperature exploration that may have driven scientific discovery despite appearing wasteful.
“What we lost out on was the casual knocking on a hallway door, just meeting someone while getting a coffee. Those serendipitous interactions may not seem optimal, but they are actually really important... You had to physically check out the journal and read the article... Sometimes it wasn't, but you could accidentally find interesting things. That has basically been lost now.”
Mathematicians have not conducted large-scale comparative experiments to test which problem-solving techniques are most effective; they rely on intuition instead. AI-type tools will revolutionize the experimental side of mathematics by enabling systematic testing of which approaches work at scale, shifting focus from individual problems to scalable workflows—analogous to software engineering's transition from handcrafted solutions to scalable systems.
“We haven't done many experiments as to, if we have two different ways to solve a problem, which is more effective. We have some intuition, but we haven't done large-scale studies where we take a thousand problems and just test them. But we can do that now.”
Current AI systems lack genuine intelligence because they cannot engage in cumulative partial progress where they reach a handhold, hold their position, and use it as a foundation for further jumps, instead operating through brute-force trial-and-error with no continuity between attempts or learning within a session.
“To go back to this analogy of these jumping robots, they can jump and fail, and jump and fail. But what they can't do is jump a little bit, reach some handhold, stay there, pull other people up, and then try to jump from there. There isn't this cumulative process which is built up interactively... It seems to be a lot more trial and error and just repetition: brute force.”
Darwin was an exceptional science communicator who wrote in plain English without equations and synthesized disparate facts into a compelling narrative, while Newton wrote in Latin and invented new mathematics, held back insights due to competitiveness, and was an unpleasant person, and it was only when other scientists explained Newton's work in simpler terms that it became widespread.
“Darwin was an amazing science communicator. He wrote in English, in natural language... He spoke in plain English, didn't use equations, and he synthesized a lot of disparate facts... Newton wrote in Latin. He had invented entire new areas of mathematics just to explain what he was doing... He was also a somewhat unpleasant person from what I gather.”
Revolutionary technologies are often taken for granted quickly after initial amazement, as demonstrated by Google search 20 years ago being revolutionary but becoming mundane, and contemporary AI abilities like face recognition and college-level math that seem remarkable now will likely be considered routine in a decade.
“People also acclimatize really quickly. I remember when Google's web search came out 20 years ago. It just blew all the other searches out of the water... and then after a few years, you just took for granted that you could Google anything. 2026-level AI would be stunning in 2021. A lot of it—face recognition, natural speech, doing college-level math problems—we just take for granted now.”
The process of identifying which scientific ideas will have broad unifying implications across multiple fields—like the concept of the bit, which emerged from Bell Labs work on signal transmission and propagated to probability theory and computer science—requires mechanisms beyond peer review and is difficult to systematize at scale.
“There's this incredibly interesting question. If you have billions of AI scientists, not only how do you gauge which ones are real progress, but how do you... This is actually a question that human science has had to face and we've solved somehow, and I'm actually not sure how we solved this. Let's say in the 1940s, if you're at Bell Labs and there are these new technologies coming out... There's one which comes up with the idea of the bit, which has implications across many different fields.”
Distant professional opportunities and unexpected obligations that take people outside their comfort zones often result in positive learning experiences and serendipitous connections, more often than they represent time wasted, making scheduled serendipity through varied obligations valuable despite appearing inefficient.
“Because it's outside my comfort zone, it often results in interactions with people I wouldn't normally talk to, like you for instance. I would learn interesting things and have interesting experiences... More often than not, I get a positive experience that I wouldn't have planned for. So I believe a lot in serendipity.”
A study measured how often scientists actually read papers they cite by tracking typos in citations across papers; if a typo was copied verbatim from one reference to another, it indicates the author didn't check the source—suggesting clever measurement tricks can reveal hidden behaviors like citation dishonesty.
“I remember reading once that people were trying to measure how often scientists actually read the papers that they cite. How do you measure this? You could try to survey different scientists, but they had a clever trick. Many citations have little typos, like a number is wrong or punctuation is almost wrong. They measured how often a typo got copied from one reference to the next, and they could infer whether an author was just copying and pasting a reference without actually checking it.”
Darwin's theory of evolution, published in 1859, came out two centuries after Newton's Principia (1687), despite being conceptually simpler, because evolutionary evidence is cumulative and retrospective while gravitational predictions can be immediately verified, and because natural selection lacked an experimental test comparable to measuring the moon's orbital period.
“The Origin of Species was published in 1859. Principia Mathematica was published in 1687. So The Origin of Species comes out two centuries after Principia. Conceptually, it seems like Darwin's theory is simpler... The evidence for natural selection is overwhelming in a certain sense, but it's cumulative and retrospective, whereas Newton can just say, 'Here are my equations. Let me see the moon's orbital period and its distance, and if it lines up, then we've made progress.'”
Tao has experienced increased productivity not uniformly but selectively: auxiliary tasks like literature searches, generating plots and figures, and formatting have been dramatically accelerated (5x faster), while the core mathematical work of solving difficult problems remains largely unaccelerated, and this is changing the nature of papers to include more supplementary material rather than solving harder problems.
“On the one hand, I think the type of papers that I would write today, if I had to do them without AI assistance, would definitely take five times longer. But I would not write my papers that way... the core of what I do, actually solving the most difficult part of a math problem, hasn't changed too much... By the same token, if I were to write a paper I wrote in 2020 again—and not add all these extra features, but just have something of the same level of functionality—it actually hasn't saved that much time.”
Despite impressive-seeming individual successes, standardized benchmarks consistently show AI having only 1-2% success rates on mathematical problems, and the journal and media attention to successes creates a distorted perception where noise overwhelms signal, making it increasingly important to establish standardized datasets and prevent AI companies from selectively reporting only victories.
“There'll be a lot of noise amongst the signal of when they're working and when they're not. It will be increasingly important to collect these really standardized datasets. There are efforts now to create a standard set of challenge problems for AIs to solve, and not just rely on the AI companies to only publish their wins and not disclose their negative results.”
Extended focused work periods without distraction eventually hit diminishing returns—after several months at an undistracted research institute, motivation and productivity decline and inspiration runs out, suggesting that some level of distraction and randomness is essential for sustained creative work.
“I spent a year once at the Institute for Advanced Study... The first few weeks you're there, it's great. You're getting all these papers written up... You think about problems for blocks of hours at a time. But I find if I stay there for more than several months, I run out of inspiration. I get bored. I surf the internet a lot more... You actually do need a certain level of distraction in your life. It adds enough randomness and high temperature.”
We do not yet know when AI systems will match or exceed frontier human mathematics performance, and the question may be posed incorrectly because AI is already performing frontier-level work in some specialized domains, but it is conceptually a different frontier from traditional human mathematics.
“In some ways, they're already doing frontier math that is super intelligent that humans can't do, but it's a different frontier from what we're used to. You could argue that calculators were doing frontier math that humans could not accomplish, but it was number crunching... But replacing Terry Tao completely. I mean, what do you want me for?”
We live in an unpredictable era where traditional assumptions about how mathematics and science work may not hold, and people entering mathematics should embrace adaptability and be open to non-traditional opportunities and ways of doing science that don't yet exist.
“We live in a time of change... Things that we've taken for granted for centuries may not hold anymore... You always have to keep an eye on opportunities for things that you wouldn't be able to do before... you should also be open to very different ways of doing science, some of which don't exist yet.”
The deductive overhang in many fields could be much larger than realized; with the right insight about how to study a problem, vastly more could be learned from existing data than is currently known, suggesting many fields have unused potential in their theoretical frameworks and existing evidence.
“One takeaway I had from reading and watching your stuff on the cosmic distance ladder… One takeaway was that the deductive overhang in many fields could be so much bigger than people realize. If you just had the right insight about how to study a problem, you might be surprised at how much more you could learn about the world.”
Terence Tao agrees with the bullish interpretation: at intermediate human-level AI intelligence, the combination of breadth and scale would constitute 'specially bullish' progress because 'human-level intelligence is qualitatively wider and more powerful than our human-level intelligence' in terms of bandwidth.
“I agree. They excel at breadth, and humans excel at depth... Not even when they achieve superhuman intelligence, but just when they achieve human-level intelligence, because their human-level intelligence is qualitatively wider and more powerful than our human-level intelligence.”
If society could from first principles decide how to optimally use Terence Tao's time as a limited resource, it would result in a very different allocation than what currently happens, and this podcast conversation would not exist under such optimization, yet many seemingly obligatory and sub-optimal uses of his time have led to serendipitous valuable interactions.
“If civilization could from first principles decide how to use Terry Tao's time, as a limited resource, what is the biggest difference? What if the veil of ignorance got to decide how to use Terry Tao's time versus what it does now? This podcast wouldn't be happening.”
AI has driven the cost of idea generation down to nearly zero, similar to how the internet drove the cost of communication down to nearly zero, but this abundance of low-cost ideas does not automatically create scientific progress without corresponding verification, validation, and evaluation structures.
“I think AI has driven the cost of idea generation down to almost zero, in a very similar way to how the internet drove the cost of communication down to almost zero. It's an amazing thing, but it doesn't create abundance by itself. Now the bottleneck is different.”
Astronomers are world-class at extracting maximum information and insight from small traces of data because data collection is the fundamental bottleneck in astronomy, and this skill is in such demand that quant hedge funds preferentially hire astronomy PhDs for their ability to extract signals from noise.
“Astronomy was one of the first sciences to really embrace data analysis and squeezing every last possible drop of information out of the information they had because data was the bottleneck. It still is the bottleneck... Astronomers are world-class in extracting all kinds of conclusions from little traces of data, almost like Sherlock. I hear that for a lot of quant hedge funds, their preferred hire is an astronomy PhD, actually.”
AI-driven democratization of math allows high school level students to contribute to frontier mathematics through AI tools and formal proof assistants, removing the traditional requirement of years of education and a PhD before attempting research contributions.
“In math, you previously had to go through years and years of education and be a math PhD before you could contribute to the frontier of math research. But now it's quite possible at the high school level, or whatever, that you could get involved in a math project and actually make a real contribution because of all these AI tools, Lean, and everything else.”
Science has a social dimension where scientists must persuade fellow scientists to accept and invest effort in learning theories, which is separate from the objective dimension of data and validation, and this persuasive dimension involves narrative, storytelling, and painting a picture of future promise even when parts of a theory cannot yet be explained.
“There's a social aspect to science. Even though we pride ourselves on having an objective side to it, where there's data and experiment and validation, we still have to tell stories and convince our fellow scientists. That's a soft, squishy thing. It's a combination of data and painting a narrative, and it's a narrative of gaps.”
AI will accelerate science in many ways, which should lead to faster discovery and breakthroughs, but it is also possible that by destroying serendipity and unstructured discovery pathways, AI could inhibit certain types of progress; the net effect is genuinely uncertain.
“Because current level AIs will accelerate science in so many ways, hopefully new discoveries and new breakthroughs will happen more quickly. It's also possible that by destroying serendipity we actually inhibit certain types of progress. Anything is possible at this point.”
Blog-writing for Terence Tao is a creative, voluntary activity that feels engaging (time flies), whereas administrative tasks feel like drudgery; this affective distinction shapes what gets written and preserved—blog posts about interesting mathematics get preserved while routine tasks don't, creating a selection effect in what knowledge becomes publicly documented.
“It's something I often do when I don't want to do other work... Writing a blog feels creative and fun. It's something I do for myself... Because it's something I do voluntarily, time flies when I write these things down, as opposed to doing something I have to do for administrative reasons that is just drudgery.”
Fields with tight data loops where predictions can be easily verified will likely see more progress from AI systems than fields that rely on cumulative retrospective evidence, suggesting a bias toward progress in hard sciences over soft sciences.
“I wonder if we'll in retrospect end up seeing much more progress in domains which have this kind of tight data loop where you can verify them quite easily, even though they're conceptually much more difficult.”
Current AI-assisted solutions to Erdős problems typically combine multiple tools and human feedback, with some problems solved through ongoing conversations between multiple humans and multiple AI systems working on different aspects rather than being solved autonomously by AI.
“People are using AI a lot currently. Someone might use AI to generate a possible proof strategy, and then another person will use a separate AI tool to critique it, rewrite it, generate some numerical data for it, or do a literature survey. Some problems have been solved by an ongoing conversation between lots of humans and lots of AI tools.”
Worry about completely incomprehensible AI proofs is probably overblown; once a formal proof exists as an artifact, we have tools to analyze, deconstruct, and interpret it, so even if the initial AI output is opaque, it can be systematically processed to extract meaning.
“Some people are concerned about what happens if the Riemann hypothesis is proven with a completely incomprehensible proof. I think once you have the artifact of a proof, we can do a lot of analysis on it. It's a very nascent area of mathematics, but I'm not as worried about it.”
The Riemann hypothesis is valued partly because solving it is believed to require new mathematics or unexpected connections between areas, and we would be shocked if it were solvable by pure exhaustive checking, but this is uncertainty—it could theoretically be false (a counterexample could exist) or only provable by brute force.
“We don't even know what the shape of the solution is, but it doesn't feel like a problem that will be solved just by exhaustively checking cases. Or it could be false actually. Okay, there is an unlikely scenario that the hypothesis is false, and you can just compute a zero off the line, and a massive computer calculation verifies it. That would be very disappointing.”
AI tools solving Erdős problems have achieved roughly 50 successes out of approximately 1,100 problems, but the pace of discovery has plateaued as the low-hanging fruit has been exhausted, and recent attempts to have frontier AI models directly solve remaining problems one-shot have ceased producing new solutions despite sustained effort.
“Fifty-odd problems have been solved with AI assistance, which is great, but there's like six hundred to go. People are still chipping away at one or two of these right now. We're seeing a lot fewer pure AI solutions now where the AI just one-shots the problem. There was a month where that happened and that has stopped, not for lack of trying.”
Tao would prefer a 'much more boring, quiet era where things are much the same as they were 10-20 years ago,' but he believes one must embrace that there will be significant change and remain open to both opportunities and disruptions.
“In many ways, I would prefer the much more boring, quiet era where things are much the same as they were 10 years ago, 20 years ago. But I think one just has to embrace that there's going to be a lot of change.”
Tao identified himself as a 'fox' who knows a bit about many things rather than a 'hedgehog' who knows one thing deeply, and attributes this to childhood obsessive-compulsiveness where he feels compelled to understand tricks and methods used by others, even across unrelated domains.
“It's not a purely human-AI distinction. Humans also, I think it was Berlin who split them into hedgehogs and foxes. The hedgehog knows one thing very well, and a fox knows a little bit about everything. I definitely think of myself as a fox... I've always had a little bit of an obsessive streak... Someone was able to use a type of mathematics I'm not familiar with and get a result I would like to prove... It bugs me that someone else can do something I think I can do, but I can't.”
Tao learns new mathematical fields through collaboration with specialists in those areas who teach him basic tricks and what is known versus unknown, and through writing a blog to record insights he learns, which initially came from frustration at forgetting previously understood material.
“I collaborate with a lot of people who have taught me other types of mathematics... I find their problems interesting, but they have to teach me some of the basic tricks, what's known, and what's not known. I learn a lot from that. I found that writing about what I've learned helps. I have a blog where I sometimes record things I've learned.”
Tao predicted in 2023 that by 2026 AI would be like a trustworthy co-author if used correctly, which prediction is holding up well in retrospect.
“You made a prediction in 2023 that by 2026 it would be like a colleague in mathematics? A trustworthy co-author if used correctly. Which is looking pretty good in retrospect. Yeah, I'm pretty pleased.”
Kepler did not highlight his third law as much as his first two because intuitively, without modern statistics, he understood that six data points were insufficient to draw reliable conclusions, suggesting even historical scientists possessed intuitive statistical skepticism.
“Maybe one reason why Kepler didn't highlight his third law as much as the first two laws is that instinctively, even though he didn't have modern statistics, he kind of knew that with six data points, he had to be somewhat tentative with the conclusions.”
A listener named Shawn solved the Jane Street puzzle by decomposing it into two parts: first identifying which layers pair together by looking for negative diagonal patterns in weight matrices, then ordering the pairs by sorting blocks by residual contribution size, combined with local swaps for refinement.
“Shawn broke the problem into two different parts. First, pair the layers into 48 different blocks. And second, put those blocks in the right order... For pairing, Shawn realized that in a well-trained ResNet, the product of two weight matrices in a residual block should have a distinctive negative diagonal pattern... For ordering, Shawn noticed that the model seemed to improve if he sorted the blocks by the size of their residual contributions.”
Tao had to deliberately wean himself off computer games because of his completionist nature—if he starts a game, he wants to play it to completion through all levels.
“I've had to wean myself off computer games because if I start a game, I want to play it to completion, through all the levels.”
Jane Street posed a puzzle to the audience where a ResNet neural network had all 96 layers shuffled, and solvers needed to recover the correct order using only model outputs and training data, without brute force (which is infeasible due to combinatorial explosion).
“Jane Street trained a ResNet, shuffled all 96 layers, and then challenged people to put them back in the right order using only the model's outputs and training data. You can't brute force this – there's more possible orderings than atoms in the universe.”