What This Is
Nonobench v1.2 is a logic-reasoning benchmark maintained by overseas developer mauricekleine that tests AI on Nonogram puzzles (logic grid puzzles where you reconstruct a black-and-white image using row and column number clues). The latest version ran 43 models, and three findings stand out. First, closed-source GPT-6 Astra scored a perfect 30/30 on 15×15 puzzles—the first perfect score on this benchmark. Second, the top open-source performer, DeepSeek V4 Pro, posted an overall score of 83% and tied for fourth place, but on the newly added 20×20 Hard mode (10 puzzles, each with a unique solution), every open-source model—including it—scored 0. Third, on the closed-source side, Opus 5.5 got 8/10 right on Hard mode, which is the current ceiling.
Industry View
Bulls read the results as puncturing a myth: "can chat" and "can reason" are two different abilities. Nonograms reward no memorization, only step-by-step deduction, and open-source models can't deliver—which means they remain unreliable for serious tasks like contract review and workflow automation. We must also air the counter-view: critics point out that Hard mode has only 10 puzzles—sample size is tiny, so the "total wipeout" may be overblown. Nonograms themselves are a narrow, game-like task, and performance on this niche puzzle set doesn't necessarily predict reliability in real office work. There's also an easily overlooked detail: DeepSeek V4.1 Flash ran the entire benchmark for just $0.84—cheap, yes, but cheap enough that submitting a blank paper has no practical value. The open-source camp still needs to prove it can run, and run right.
What It Means for the Rest of Us
For enterprise IT: when selection involves workflow automation or multi-step rule reasoning, open-source models still can't "complete tasks independently" today—we recommend closed-source as the primary choice, with open-source reserved for edge cases or backup. For individual professionals: it's fine to let AI draft your emails or summarize documents, but letting it independently review contracts, make judgments, or run workflows still requires a human in the loop—don't be fooled by demo videos. For the consumer market: this test fires another warning shot. There's a real gap between vendors' "thinking AI" marketing and actual capability. When picking a paid AI product, don't just watch the vendor demo—benchmark it on your own real tasks.