Visual reasoning puzzles designed by Russian scientist Mikhail Bongard 60 years ago were placed back on the AI evaluation table this week by researcher Matthew Hodges. We notice a less-than-optimistic signal: today's strongest multimodal large models — including GPT-5, Claude, and Gemini — collectively lost points on this old test. This is a reminder to the entire industry: the gap between AI "seeing" and "understanding" may be wider than vendors are willing to admit.
What This Is
Bongard Problems are a set of visual puzzles: each problem presents two groups of figures, distinguished only by some abstract rule — for example, "whether the figure is closed" or "whether it contains three lines" — and you must find that hidden feature yourself. Born in 1967, they have long been a classic prop in cognitive science and philosophy for discussing what "understanding" really means.
Why are they being dusted off again? Because mainstream visual benchmarks have turned into a data arms race, with models scoring high by memorizing training distributions. Bongard Problems, by contrast, stress cross-domain abstract reasoning — exactly the traditional weak spot of large models. Hodges pointed out in his blog that swapping out a set of figures or changing the way questions are posed can shift scores by as much as 30%.
Industry View
Supporters believe these puzzles hit the bullseye. University of Toronto professor and cognitive scientist Melanie Mitchell has tracked model performance on Bongard Problems for years, finding that models at the level of GPT-4V and Claude 3 still score below 60% accuracy — far below the near-100% level of humans. She argues this shows that mainstream benchmarks have simply failed to adequately cover "visual understanding."
Opposing views are equally pointed. One camp of researchers argues that Bongard Problems themselves are too anthropocentric — designed around human intuition, they may not measure what AI truly needs; forcing them as a benchmark introduces the bias of "measuring machines with a human ruler." A more pragmatic camp argues that model performance is unstable because the way the problems are presented (line thickness, contrast, prompts) interferes too much, jumbling "abstract capability" together with "perceptual robustness," so the scores cannot tell which is the real problem.
Impact on Regular People
For enterprise IT: If you are evaluating "visual AI" vendors, don't just look at the standard demo set — ask one more question: what happens when the scenario requires seeing through surface rules? This is the blind spot where today's models collectively fail.
For individual professionals: When using AI to make charts or illustrations, it can handle "looking good," but the abstract judgment of "why this chart makes the point better than that one" still needs human backup. Don't be fooled by the tool.
For the consumer market: Over the next year, a wave of "AI image viewing" and "AI photo Q&A" consumer products will launch. The hidden cost exposed by Bongard Problems will become a watershed — the experience gap between products that pass and those that fail will be very tangible.