High-Schoolers Outsmart World’s Top AI Models

The Rise of AI in the International Mathematical Olympiad

The most advanced artificial intelligence models have entered the world’s most prestigious competition for young mathematicians and achieved a breakthrough that once seemed impossible. Despite their impressive performance, they still fell short of the world’s brightest teenagers.

Every year, a few hundred elite high school students from around the globe gather at the International Mathematical Olympiad (IMO). This year, they were joined by companies like Google DeepMind and others in the artificial intelligence industry. These tech giants came to participate in one of the ultimate tests of reasoning, logic, and creativity.

The IMO is known for its challenging exam, which spans two days with three increasingly difficult problems each day, giving participants over four hours to solve them. The questions cover algebra, geometry, number theory, and combinatorics. To tackle them, you need to be a math genius—just understanding the problems can be a brain workout.

Because these problems are both complex and unconventional, the annual math test has become a useful benchmark for measuring AI progress. In recent years, leading research labs have dreamed of a day when their systems would meet the standard for an IMO gold medal, which was seen as the AI equivalent of a four-minute mile.

However, no one knew when or if this milestone would be reached—until now. Earlier this month, an AI model from Google DeepMind earned a gold-medal score by perfectly solving five out of six problems. Another dramatic twist came when OpenAI also claimed gold, even though it wasn’t officially part of the event. Both companies described their achievements as significant steps toward the future, even if they weren’t quite there yet.

What makes this event particularly remarkable is that 26 students scored higher than the AI systems. Among them were four U.S. team members, including Qiao (Tiger) Zhang, a two-time gold medalist from California, and Alexander Wang, who brought home his third straight gold from New Jersey. Wang is one of the most decorated young mathematicians of all time and is a high-school senior who could aim for another gold next year.

But in a year, he might be dealing with a different equation altogether. “I think it’s really likely that AI is going to be able to get a perfect score next year,” Wang said. “That would be insane progress,” Zhang added. “I’m 50-50 on it.”

So, will this be remembered as the last IMO when humans outperformed AI? “It might well be,” said Thang Luong, the leader of Google DeepMind’s team.

DeepMind vs. OpenAI

Until recently, what happened in Australia would have sounded about as likely as koalas doing calculus. But the impossible began to feel almost inevitable last year when DeepMind’s models built for math solved four problems and scored 28 points for a silver medal, just one point short of gold. This year, the IMO officially invited a select group of tech companies to their own competition, giving them the same problems as the students and having coordinators grade their solutions with the same rubric.

These companies were eager for the challenge. AI models are trained on vast amounts of information, so if anything has been done before, the chances are they can figure it out again. However, they can struggle with problems they’ve never seen before.

As it turns out, the IMO process is specifically designed to come up with those original and unconventional problems. In addition to being novel, the problems must also be interesting and beautiful, according to IMO president Gregor Dolinar. If a problem is similar to any other published problem, it gets tossed. By the time students take the exam, the list of a few hundred suggested problems has been whittled down to six.

Meanwhile, the DeepMind team kept improving the AI system it would bring to the IMO, an unreleased version of Google’s advanced reasoning model Gemini Deep Think, and it was still making tweaks in the days leading up to the competition.

The effort was led by Thang Luong, a senior staff research scientist who narrowly missed getting to the IMO in high school with Vietnam’s team. He finally made it to the IMO last year—with Google. Before returning this year, DeepMind executives asked about the possibility of gold. Luong told them to expect bronze or silver again.

He adjusted his expectations when DeepMind’s model nailed all three problems on the first day. The simplicity, elegance, and readability of those solutions astonished mathematicians. The next day, as soon as Luong and his colleagues realized their AI creation had crushed two more proofs, they also realized that would be enough for gold.

They celebrated their monumental accomplishment by doing one thing the other medalists couldn’t: They cracked open a bottle of whiskey.

To keep the focus on students, the companies at the IMO agreed not to release their results until later this month. But as soon as the Olympiad’s closing ceremony ended, one company declared that its AI model had struck gold—and it wasn’t DeepMind. It was OpenAI.

The company wasn’t a part of the IMO event, but OpenAI gave its latest experimental reasoning model all six problems and enlisted former medalists to grade the proofs. Like DeepMind’s, OpenAI’s system flawlessly solved five and scored 35 out of 42 points to meet the gold standard.

After the OpenAI victory lap on social media, the embargo was lifted and DeepMind told the world about its own triumph—and that its performance was certified by the IMO.

Not long ago, it was hard to imagine AI rivals dueling for glory like this.

In 2021, a Ph.D. student named Alexander Wei was part of a study that asked him to predict the state of AI math by July 2025—that is, right now. When he looked at the other forecasts, he thought they were much too optimistic. As it turned out, they weren’t nearly optimistic enough. Now he’s living proof of just how wrong he was: Wei is the research scientist who led the IMO project for OpenAI.

The Problem of Problem 6

The only thing more impressive than what the AI systems did was how they did it. Google called its result a major advance, though not because DeepMind won gold instead of silver. Last year, the model needed the problems to be translated into a computer programming language for math proofs. This year, it operated entirely in “natural language” without any human intervention. DeepMind also crushed the exam within the IMO time limit of 4 ½ hours after taking several days of computation just a year ago.

You might find all of this completely terrifying—and think of AI as competition. The humans behind the models see them as complementary. “This could perhaps be a new calculator,” Luong said, “that powers the next generation of mathematicians.”

Speaking of that next generation, the IMO gold medalists have already been overshadowed by AI. So let’s put them back in the spotlight.

Qiao Zhang is a 17-year-old student in Los Angeles on his way to MIT to study math and computer science. As a young boy, his family moved to the U.S. from China, and his parents gave him a choice of two American names. He picked Tiger over Elephant.

His career in competitive math began in second grade, when he entered a contest called the Math Kangaroo. It ended this month at the math Olympics next to a hotel in Australia with actual kangaroos.

When he sat down at his desk with a pen and lots of scratch paper, Zhang spent the longest amount of time during the exam on Problem 6. It was a problem in the notoriously tricky field of combinatorics, the branch of mathematics that deals with counting, arranging, and combining discrete objects, and it was easily the hardest on this year’s test. The solution required the ingenuity, creativity, and intuition that humans can muster but machines cannot—at least not yet.

“I would actually be a bit scared if the AI models could do stuff on Problem 6,” he said.

Problem 6 did stump DeepMind and OpenAI’s models, but it wasn’t just problematic for AI. Of the 630 student contestants, 569 also received zero points. Only six received the full credit of seven points. Zhang was proud of his partial solution that earned four points—which was four more than almost everyone else.

At this year’s IMO, 72 contestants went home with gold. But for some, a medal wasn’t their only prize. Zhang was among those who left with another keepsake: victory over the AI models.

(As if it weren’t enough that he can bend numbers to his will, he also has a way with words and wrote this about his IMO experience.)

In the end, the six members of the U.S. team piled up five golds and one silver, finishing second overall behind the Chinese after knocking them off the top spot last year.

There was once a time when such precocious math students grew up to become professors. (Or presidents—the recently elected president of Romania was a two-time IMO gold medalist with perfect scores.) While many still choose academia, others get recruited by algorithmic trading firms and hedge funds, where their quantitative brains have never been so highly valued. This year, the U.S. team was supported by Jane Street while XTX Markets sponsored the whole event. After all, they will soon be competing with each other—and with the richest tech companies—for their intellectual talents.

By then, AI might be destroying mere humans at math. But not if you ask Junehyuk Jung.

A former IMO gold medalist himself, Jung is now an associate professor at Brown University and visiting researcher at DeepMind who worked on its gold-medal model. He doesn’t believe this was humanity’s last stand, though. He thinks problems like Problem 6 will flummox AI for at least another decade.

And he walked away from perhaps the most significant math contest in history feeling bullish on all kinds of intelligence.

“There are things AI will do very well,” he said. “There are still going to be things that humans can do better.”

Leave a Comment