
Is AI's "intelligence" an illusion? | GZERO World with Ian Bremmer
What this covers
Is ChatGPT all it’s cracked up to be? Will truth survive the evolution of artificial intelligence?
Sign up for GZERO Daily (free newsletter on global politics): https://rebrand.ly/gzeronewsletter Subscribe to GZERO on YouTube: http://bit.ly/2TxCVnY
Is ChatGPT all it’s cracked up to be? Will truth survive the evolution of artificial intelligence?
On GZERO World with Ian Bremmer, cognitive scientist and AI researcher Gary Marcus breaks down the recent advances––and inherent risks––of generative AI.
AI-powered, large language model tools like the text-to-text generator ChatGPT or the text-to-image generator Midjourney can do magical things like write a college term paper in Klingon or instantly create nine images of a slice of bread ascending to heaven.
But there’s still a lot they can’t do: namely, they have a pretty hard time with the concept of truth, often presenting inaccurate or plainly false information as facts. As generative AI becomes more widespread, it will undoubtedly change the way we live, in both good ways and bad.
“Large language models are actually special in their unreliability,” Marcus says on GZERO World, “They're arguably the most versatile AI technique that's ever been developed, but they're also the least reliable AI technique that's ever gone mainstream.”
Marcus sits down with Ian Bremmer to talk about the underlying technology behind generative AI, how it differs from the “good old-fashioned AI” of previous generations, and what effective, global AI regulation might look like.
Watch new episodes of GZERO World with Ian Bremmer every week on YouTube https://www.youtube.com/playlist?list=PLwpOQsKKqZGNfKZ7zx127SGsnqMY4mPWi or at gzeromedia.com/gzeroworld and on US public television. Check local listings.
GZERO Media is a multimedia publisher providing news, insights and commentary on the events shaping our world. Our properties include GZERO World with Ian Bremmer, our newsletter GZERO Daily, Puppet Regime, the GZERO World Podcast, In 60 Seconds and GZEROMedia.com
#GZEROWorld #ArtificialIntelligence #ChatGPT
Source description (no synthesized summary yet).
Large language models like ChatGPT are fundamentally unreliable autocomplete systems that create an illusion of capability while lacking genuine understanding, making them unsuitable for safety-critical applications without human oversight and necessitating new regulatory frameworks including national AI agencies and FDA-style safety reviews.
- LLMs analyze statistical relationships between words rather than concepts or real-world entities, leading to hallucinations and false claims they present with unwarranted confidence
- Combining LLMs with fact-checking or logic systems is theoretically possible but practically difficult because it requires solving the 75-year-old AI problem of translating natural language into verifiable logical form
- Current regulatory infrastructure is inadequate; governments need dedicated AI agencies and deployment oversight similar to FDA drug approval before systems affecting millions of users are released
This asset isn't compiled yet
You're seeing its claims, ranked. Compile it to build the argument threads, weight them, and check each claim against your library — the full view.
Combining large language models with fact-checking systems (like Google or Wikipedia searches) is theoretically appealing but practically difficult because it requires translating the LLM's natural language output into formal logic, a problem AI researchers have been struggling with for 75 years without complete success.
“People are trying, but the reality is it's kind of like apples and oranges. You know, they both look like fruit, but they're really different things. And in order to do this, quote, 'quick search,' what you really need to do is to analyze the output of the large language model. Basically take sentences and translate them into logic.”
Large language models analyze the statistical relationship between words in their training data rather than analyzing concepts, ideas, or entities in the world, functioning essentially as 'autocomplete on steroids' that predict the most likely next word based on learned patterns.
“They are analyzing something, but what they're analyzing is the relationship between words, not the relationship between concepts or ideas or entities in the world. And so they're basically like autocomplete on steroids.”
Siri is engineered very differently from large language models: Siri is carefully designed to do a few specific things reliably through APIs that hook into external systems, whereas LLMs are generalist systems that pretend to do everything without actually connecting to the world, creating an illusion of capability.
“Siri was very carefully engineered to do only a few things and do them really well, and it has APIs to hook out into the world... Siri only does a few things. It'll control your lights if you have the right kind of light switches, it will lock your door if you have the right kind of door. But it doesn't just make stuff up.”
Creating AI that is genuinely truthful and factual to warrant human confidence similar to autonomous vehicle safety standards requires fundamental changes in approach, involving a retreat from current scaling strategies and investment in techniques that modern AI has abandoned for 30 years.
“I think we're fairly far, but the main thing is, I would think of it in terms of climbing mountains in the Himalayas. You're at one peak, and you see that there's this other peak that's higher than you and the only way you're going to get there is if you climb back down. And that's emotionally painful... you're going to have to give up the money that you're making now in order to make a long-term commitment to doing something that feels foreign and different. Nobody really wants to do that.”
Current LLM systems lack transparency regarding their training data sources and mechanisms of political influence, which combined with their potential to affect everyone's political opinions creates a public interest justifying mandatory oversight even for users who did not consent to the system.
“These systems are going to affect people's political opinions. So everybody, even if they signed up or not, is going to be affected by what these systems do. We have no transparency, we don't know what data they're trained on, we don't know how they're going to influence the political process, and so that affects everybody and so we should have some oversight of that.”
Despite significant investment in driverless car technology, after 7 years since 2016 the industry remains plagued by the 'outlier problem'—scenarios not represented in training data that the systems cannot handle—exemplifying how capital investment alone does not guarantee technological breakthroughs.
“In 2016, I said, 'You know, even though these things look good right now, I'm not sure they're going to be commercialized anytime soon because there's an outlier problem.' And the outlier problem is there's always some scenarios you haven't seen before, and the driverless cars continue to be plagued by this stuff seven years later.”
Large language models are the most versatile AI technique ever developed but also the least reliable AI technique that has ever gone mainstream, making them unique in requiring human oversight for nearly all practical applications.
“Large language models are actually special in their unreliability. They're arguably the most versatile AI technique that's ever been developed, but they're also the least reliable AI technique that's ever gone mainstream.”
Every nation should establish its own dedicated AI agency or cabinet-level position to coordinate oversight across existing agencies and respond to the rapidly moving risks posed by AI development.
“First thing is, I think every nation has to have its own AI agency or cabinet-level position, something like that, in recognition of how fast things are moving and in recognition of the fact that you can't just do this, like, with your left hand. You can't just say to all the existing agencies, 'Yeah, you can just do a little bit more and handle AI.' There's so much going on.”
The digital order, emerging or emerging imminently, is controlled by technology companies rather than governments, creating a new locus of global power outside traditional state structures.
“Now, the third global order may not be quite here yet, but it is right around the corner. And I'm talking about the digital order, which is not run by governments but rather by technology companies.”
Old-fashioned symbolic AI (good old-fashioned AI) based on logic, symbols, and mathematics has been dismissed prematurely and should be reconsidered because it is much better than neural networks at truth-telling when given a limited set of facts, as it reasons within those facts without hallucinating.
“I think once people start, first of all, taking seriously old-fashioned AI, sometimes people call it 'good old-fashioned AI,' it's totally out of favor right now, but it was, I think, dismissed prematurely.”
Ian Bremmer argues that there is no longer a single global order to lead because the United States has declined as a unilateral global power, and instead there are three distinct global orders: security (dominated by the US), economic (shared power), and digital (controlled by technology companies).
“But if there isn't just one global order anymore, how many are there? Actually, I'd say three. First off, there's the global security order, and in that realm, at least, the United States and its allies are the most powerful players.”
Gary Marcus previously made seven predictions about GPT-4 in an essay titled 'What to Expect When You're Expecting... GPT-4,' and all seven predictions proved correct.
“I wrote an essay called 'What to Expect When You're Expecting... GPT-4.' I made seven predictions and they were all right, and one of them was that GPT-4 would continue to hallucinate.”
The United States remains the sole superpower in military/security terms, being the only country capable of projecting military force to every corner of the world, with China and Russia as distant second-tier powers whose military capabilities do not match US dominance.
“The US is the only country in the world that can send its soldiers and its sailors to every corner of that world. No one else close. Sure, China's making a go of it with its warships coveting Taiwan's coastline like a jealous lover. And Russia's military had some swagger before it fell apart in Ukraine. But as long as nuclear war remains synonymous with suicide, the US, yes the US, stays alone at the top of the military flagpole.”
Technology companies are releasing explosive, productive, and disruptive technologies without pause buttons, creating risks that governance struggles to manage in real time.
“There is no pause button on these explosive productive and disruptive technologies.”
Technology companies' current advertising models turn citizens into products and drive hate and misinformation into societies, creating negative externalities that justify governance intervention.
“Will they proceed with advertising models that turn citizens and products and drive hate and misinformation into our societies?”
Even if GPT-5 is incrementally better than GPT-4, the crucial question is whether it becomes reliable enough for deployment in safety-critical domains like medicine, and current progress provides no guarantee that future versions will cross this threshold.
“So programmers are good, but take medicine. It's not clear that even GPT-5 is going to be a reliable source of medicine. 5, or 4, is actually better than 3. but is it reliable enough is an interesting question. You know, each year's driverless car is better than the last, but is it reliable enough?”
A Tesla vehicle ran into a jet aircraft because the jet was not in the vehicle's training set, demonstrating that the autonomous vehicle system failed to learn the abstract principle that driving requires avoiding large objects, instead relying on pattern matching to specific objects it had seen before.
“You know, I gave an example then that Google had just maybe solved, at that point, a problem about recognizing piles of leaves on the road. And I said there's going to be a huge number of these problems, we're never gonna solve them, and we still haven't. So there was a Tesla that ran into a jet not that long ago because a jet wasn't in its training set, it didn't know what to do with it. It hasn't learned the abstract idea that driving involves not running into large objects, and so if there's a particular large object that isn't in its training set, it doesn't know what to do.”
Siri's design philosophy emphasizes not overselling capability—when it cannot do something, it says so rather than attempting to answer and potentially giving false information.
“Every once in a while, they roll out an update and now we can ask it about sports scores or movie scores. At first you couldn't. But it's very narrowly engineered.”
The false claim by a language model that Tesla CEO Elon Musk died in a fatal car accident on March 18th, 2018 demonstrates that LLMs cannot actually analyze the world, because a system with genuine understanding would recognize the contradiction with overwhelming evidence that Musk was alive, active on Twitter daily, and continuously in the news.
“A good example of this was one of these systems saying on March 18th, 2018, Tesla CEO Elon Musk died in a fatal car accident. Well, we know that a system that could actually analyze the world wouldn't say this. We have enormous data that Elon Musk is still alive. He didn't die in 2018. He tweets every day. He's in the news every day.”
Companies should not merely publish research papers about AI systems affecting 100 million users; instead, independent experts should examine safety claims to ensure systems meet minimum safety standards, similar to FDA drug approval processes.
“You don't just publish a paper, you have people examine it. We need to do the same thing. You can't just put something out. And if you affect 100 million users, you really affect everybody.”
In the global economic order, power is shared among the US, EU, China, India, and Japan, with the US and China so economically interdependent that neither Washington nor Beijing can achieve economic dominance, preventing the emergence of a unilateral economic superpower.
“There's a global economic order, and here, power is shared. The United States, of course, is still a very robust global economy, but is unable to use its dominant position militarily to tell other countries what to do economically. And so when it comes to China, the two countries are so economically interdependent that neither Washington nor Beijing can get the upper hand. The EU, by the way, the European Union, has the largest common market. India and Japan, very much in the mix as well.”
Large language models like ChatGPT-5 and ChatGPT-6 will continue to hallucinate significantly unless they involve radical new machinery beyond simply scaling up neural networks trained on more data.
“I wrote an essay called 'What to Expect When You're Expecting... GPT-4.' I made seven predictions and they were all right, and one of them was that GPT-4 would continue to hallucinate. And I will go on record now as saying GPT-5 will, unless it involves some radical new machinery. If it's just a bigger version trained on more data, it will continue to hallucinate. And same with GPT-6.”
Large language models are actually illiterate in the meaningful sense—they cannot read with high comprehension or genuinely understand the content being discussed, despite producing fluent text that mimics understanding.
“Another way to put it all together is these things are actually illiterate. We don't know how to build an AI system that can actually read with high comprehension, really understand the things that are being discussed.”
Digital identity is now determined by nature, nurture, and algorithm, reflecting the increasing role of algorithmic systems in shaping human identity and experience beyond genetic and environmental factors.
“When I was growing up, it was nature and nurture that determined our identities, but now it's nature, nurture, and algorithm, and there is no pause button on these explosive productive and disruptive technologies.”
Economic incentives may finally exist to drive improvement in LLM truthfulness through the desire for 'chat search' to work well—combining ChatGPT-style interfaces with search functionality—because users can clearly see the value and express frustration with current failures.
“I think there's a chance maybe that finally we have the right economic incentive, which is people want what I call 'chat search' to work, which is you type in a search to ChatGPT. It doesn't work that well right now, but everybody can see how useful and valuable that would be. And maybe that will put enough money and enough kind of frustration because it's not working to get people to kind of turn the boat.”
API is jargon for 'Application Program Interface,' a technical term for connecting machines to external systems and functions.
“API, explain API for everybody. - Application program interface. It's jargon for hooking your machine up to the world.”
Smart people are concerned that generative AI tools will reshape society in both positive and negative ways, justifying serious study and governance of these technologies.
“Many of the smartest people I know think these tools will reshape the way we live in both good ways and bad, and that is the subject of my interview today with psychologist, cognitive scientist, and NYU Professor Emeritus Gary Marcus.”
Ian Bremmer opens the interview by noting that ChatGPT and similar systems 'can churn out a two-hour movie script or Picasso-style painting in, like, just an instant,' establishing the apparent capability that the interview then critiques.
“Those chatbots like ChatGPT that you've surely heard about by now. You know, the ones that can churn out a two-hour movie script or Picasso-style painting in, like, just an instant.”
OpenAI's release of ChatGPT-4 demonstrates that the AI industry is still only at 'the tip of the AI iceberg,' suggesting significant further development is expected.
“With the recent rollout of OpenAI's ChatGPT-4, it's clear we're still only at the tip of the AI iceberg.”
Ian Bremmer humorously jabs that 'no, a chatbot did not write this script,' suggesting common concerns that AI systems might be replacing human work while insisting on human authorship of the episode introduction.
“The truth is, I couldn't do this show without my beloved and irreplaceable team of producers and editors whose families rely on... No, hey, this is not what I approved for the teleprompter. Anyway... No, a chatbot did not write this script.”