header-langage
简体中文
繁體中文
English
Tiếng Việt
한국어
日本語
ภาษาไทย
Türkçe
Scan to Download the APP

25 models collectively lose to humans: AI still can't understand many common sense aspects of life

Beating AI News Flash: Scale Labs, the research arm of Scale AI, has partnered with Elorian to release a new benchmark called Humanity's Sixth Sense (HSS), specifically designed to test whether AI can discern implicit information from images and videos. The 522 open-ended questions cover 288 images and 234 videos. For example, whether two cars can pass between each other, why a woman suddenly slows down while chasing a bus, and who holds more sway in a given situation.


The research team tested 25 multimodal models and had 20 human participants answer the questions. Human accuracy reached 93.1%, while GPT-6 Astra, which ranked first, scored only 53.6% even with maximum reasoning effort. GPT-6.1 Sol and Claude Opus 5.5 scored 46.6% and 44.6% respectively, and the median score across all models was just 30.9%.


The researchers analyzed 8,573 model failures and found that 94% were related to missing key clues, misidentifying subjects, or being unable to infer implicit relationships in the visuals. Only about 5% were classified as logical reasoning errors. Social understanding proved especially difficult, with 21 of the 25 models performing worst on this type of question. Video questions were also generally harder than image questions.


Increasing the amount of reasoning did not necessarily help. Models consumed an average of about 4,000 reasoning tokens per question, and for some questions, the longer they thought, the worse their scores became. The research team also tried having Agents zoom in, crop, and repeatedly inspect the visuals, which improved scores somewhat, but still left a clear gap compared with humans.

举报 Correction/Report
Correction/Report
Submit
Add Library
Visible to myself only
Public
Save
Choose Library
Add Library
Cancel
Finish