20,168 Python debugging tasks went live quietly this week, designed to evaluate local large models' (small models running on your own computer, offline) ability to fix code bugs. What it really targets is a chronic problem in AI evaluation: models score perfectly on visible test sets, then fall apart when moved to hidden test cases.
What this is
Failure Map, an open test set aimed at local code models. The 20,168 Python debugging tasks span 254 categories, and each task ships with a standard-library implementation, explicit contract (pre-agreed input/output rules), known-failing fix samples, and executable boundary checks.
Three representative tasks — deduplication logic mistakenly removes valid events; cache expiry time is off by one unit, so legitimate records get treated as stale; pagination code flips > to >=, returning the same record twice. Problems are free to download, with unified prompt templates and scoring scripts provided.
Industry view
Supporters argue that running code tasks on local models is cheap, code never leaves the enterprise network (a must-have for privacy-sensitive scenarios), and the community urgently needs a public, reproducible evaluation benchmark to stop vendors from patting themselves on the back.
Criticisms and risks are not in short supply. First, the authors themselves admit they "haven't tested any local model yet" — this is essentially an empty exam paper, inviting the community to fill in the scores; the credibility of distribution will depend on data accumulating over time. Second, the authors repeatedly stress that "passing visible tests" does not equal "passing hidden cases," and the implied accusation is that many industry benchmarks effectively let the model see the answers before testing — scores are systematically inflated. Third, running hidden tests to produce meaningful results requires compute and engineering investment that small teams and indie developers can't afford, which could give rise to a new round of evaluation hegemony. Fourth, once 20K problems are published, will they be specifically optimized against during fine-tuning? An old issue in the NLP world, called "data contamination."
Impact on regular people
For enterprise IT: When buying "AI coding" tools, require vendors to produce scores in a closed-book setting (without exposing the test set). Otherwise, no matter how pretty the numbers look, it's still just vendor talk.
For working professionals: More and more people are using local models to write scripts and check bugs; they're more likely to first verify on their own code with a public benchmark like this before pushing to production.
For the consumer market: Over the next year, local models claiming "coding ability exceeding GPT-4" will keep popping up. Ordinary users can't reproduce these benchmarks, so don't take visible scores at face value.