NonobenchDeepSeek
43 LLMs Take the Same Logic Test: Open-Source Flunks, Top Closed-Source Hits 80%
Nonobench tested 43 LLMs on Nonogram logic puzzles. Every open-source model scored 0 on Hard mode; the top closed-source model hit only 80%.
Sep 27·2 min read