header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

Predicting World Cup Knockout Stage: Why Such Discrepancy in AI Performance?

Read this article in 14 Minutes
DeepSeek and Gemini shone the brightest in the Netherlands vs. Morocco match, while Grok and QWERTY excelled at predicting the exact scores of popular games.
Original Title: "Predicting World Cup Knockout Stage: Why Such Discrepancy in AI Performance?"
Original Author: Asher, Odaily Planet Daily


Before each World Cup match, I always have AI make predictions, and almost every model sounds knowledgeable and detailed.


Some talk about team value, some dissect group stage data, some analyze injuries and tactics, and some even directly provide score predictions, extra time scenarios, and penalty shootouts. At first glance, ChatGPT, Grok, Qianwen, DeepSeek, Gemini, and Claude all seem to understand football.


However, as a user of prediction markets, what I truly care about is not which model provides a more comprehensive analysis, but which one is more worth referencing.


As the World Cup enters the knockout stage, Odaily Planet Daily has started from the first match, asking as similar questions as possible to different AI models before each game, and then comparing the actual results afterward to see which models merely made decent analyses and which ones actually captured the game's trajectory in advance.


So far, in the concluded World Cup knockout matches, Canada secured a 1-0 victory over South Africa, Brazil narrowly defeated Japan 2-1, Germany was dragged into a penalty shootout by Paraguay and eliminated, and the Netherlands also fell to Morocco on penalties. In the Belgium vs. Senegal match, the game was even more intense, ending 2-2 and then seeing a turnaround in extra time, maximizing the uncertainty of the knockout stage.


DeepSeek and Gemini Shine by Predicting the Morocco Showdown


One of the most memorable moments so far is DeepSeek and Gemini's prediction for the Netherlands vs. Morocco match. This match was actually quite easy to get wrong—on paper, the Netherlands were stronger and had a more complete lineup, and many models knew that Morocco would be a tough opponent. However, they still ultimately believed that the Netherlands would advance.


What set DeepSeek and Gemini apart was that they did not stop at "this match will be close," but they also outlined the subsequent plot. Gemini directly predicted a 1-1 draw in regular time, with Morocco winning in a penalty shootout. The result was indeed a 1-1 draw during regulation time, and in the end, Morocco won the penalty shootout 3-2 to eliminate the Netherlands. They didn't just guess the direction but rather accurately predicted how the match would go into a penalty shootout, who would have the last laugh, and so on.


Gemini's prediction for the Netherlands vs. Morocco match


DeepSeek is also very close. It predicts that this regular time match is likely to be 1:1 or 0:0, the game may go all the way to extra time or even penalties, and it leans towards Morocco pulling off an upset through defense and counterattacks.


DeepSeek predicts the Netherlands vs. Morocco match


After this match, DeepSeek and Gemini's presence is fully felt. Especially Gemini, this time it's not like making pre-match predictions, but more like having seen the script of the match in advance.


Grok and Qianwen consistently hit specific scores, showing stronger stability than imagined


In addition to DeepSeek and Gemini shining in this Morocco match, Grok and Qianwen also have their presence. Their most prominent aspect is in some matches where the outcome is relatively clear, not only predicting the advancing team correctly but also forecasting the specific scores quite close to the final result.


South Africa vs. Canada is an example. Before the match, most AI models favored Canada, but the disagreement was whether Canada would win easily. Grok gave a prediction of Canada winning 1:0 before the match, and Qianwen also predicted a one-goal victory. In the end, Canada did only pass through with one goal and did not achieve the expected big win.


Qianwen predicts the South Africa vs. Canada match


Brazil vs. Japan is similar. Most AI models believed Brazil was stronger, but whether Japan could hold the match was the key. Grok and Qianwen both predicted the score to be 2:1, and in the end, the match indeed ended with Brazil's narrow 2:1 victory. What they predicted correctly was not just "Brazil will win" so simply, but that Japan could cause enough trouble for Brazil.


Ivory Coast vs. Norway, this match also saw accurate predictions from both sides. With Norway having Haaland, it's understandable to see the direction of advancement, but Ivory Coast's physical confrontation and wing attacks will also not make the game completely one-sided. Grok and Qianwen both predicted a 2:1 win for Norway, and the final score fell exactly into this "script."


Grok predicts the Ivory Coast vs. Norway match


The strengths of Grok and Thousand Whys lie in delving deeper into popular matches. They didn't script major twists like Morocco eliminating the Netherlands in advance, but in matches such as Canada, Brazil, Norway, and France, they provided more precise insights into the direction of victory and the score. In other words, they may not always excel at predicting upsets, but they are very good at judging whether the favored team will dominate or barely win.


ChatGPT doesn't offer many miracle scores, but its match analysis is quite accurate


ChatGPT didn't foresee a Morocco penalty shootout victory over the Netherlands like Gemini did, nor did it consistently hit specific scores like Grok and Thousand Whys. However, its strength lies in many matches where the pre-match view favors a strong team, ChatGPT will more clearly indicate that this match might not be that straightforward.


The Brazil-Japan match is an example. ChatGPT predicted Brazil to advance but didn't portray the match as an easy domination by Brazil. Instead, it mentioned that Japan's pressing, running, and discipline would make Brazil uncomfortable, with the chance to score first or equalize. The Ivory Coast-Norway match is similar. ChatGPT predicted Norway to advance but pointed out early on that it wouldn't be an easy game, as the Ivory Coast's physicality, flank attacks, and counterattacking abilities would cause trouble.


ChatGPT predicts the match between England and the Democratic Republic of the Congo


The strength of ChatGPT lies not in predicting the score accurately every time but in often being able to foresee where the resistance in the match will be. It is suitable for understanding the game but not for those looking for a precise final score prediction. It can describe the process accurately, but when it comes to predicting a major upset, it lacks a bit of decisiveness.


Germany's elimination becomes a collective AI model disaster


If the previous matches showcased each model's highlights, the Germany-Paraguay match was a collective disaster.


Before the match, all AI models sided with Germany. ChatGPT, Grok, Thousand Whys, Gemini, Claude—everyone supported Germany, with most score predictions centered around 2-0, 3-0, or 3-1. The reasons were also very consistent: they all believed Germany had stronger overall strength, better squad depth, and more attacking prowess.


However, the result was a problem. The AI models underestimated Paraguay's ability to drag the match into a quagmire. Germany couldn't resolve the battle in regular time, couldn't break the deadlock in extra time, and was ultimately dragged into a penalty shootout by Paraguay, leading to their elimination.


Who is the Most Accurate Predictor Right Now?


From the knockout stage matches that have already concluded, the characteristics of different models have begun to emerge.


DeepSeek and Gemini stand out the most. They were not only able to predict the advancement of popular teams like Brazil and France, but also provided valuable insights in more difficult matches. In the Netherlands vs. Morocco match, their key advantage was the courage to predict Morocco's upset victory and the subsequent penalty shootout. Especially Gemini, which directly predicted Morocco's penalty shootout advancement, was truly impressive.


Grok and Inquisitor are more like "score-oriented players". They correctly predicted many specific scores, especially performing well in matches involving Canada, Brazil, Norway, and France. However, when facing traditional powerhouses like Germany and the Netherlands, they still tended towards the favorites in the end.


ChatGPT and Claude, on the other hand, are more like "analysis-oriented players". Their reasoning is comprehensive, and their overall direction is mostly on track, capable of pointing out some overtime risks. However, the issue is that they can often recognize when a match will be difficult but are hesitant to conclude in favor of the underdog. This was evident in the Netherlands vs. Morocco match where, despite foreseeing the risk of extra time and penalties, they ultimately leaned towards the Netherlands.


Therefore, instead of rushing to ask which model understands football the best, it is better to consider which scenarios each of them is more suitable for.


Original Article Link


Welcome to join the official BlockBeats community:

Telegram Subscription Group: https://t.me/theblockbeats

Telegram Discussion Group: https://t.me/BlockBeats_App

Official Twitter Account: https://twitter.com/BlockBeatsAsia

举报 Correction/Report
Choose Library
Add Library
Cancel
Finish
Add Library
Visible to myself only
Public
Save
Correction/Report
Submit