Back to home
AI Evaluation
2 articles tagged with this topic
QwenTongyi Qianwen
Qwen 27B Fixed a Real Git Repo — But Its Eval Method Is the Real Story
A dev tested Qwen 27B on a real Git repo. It fixed code, ran tests, recovered. The real story: the eval method—and 6 criteria China AI lacks.
4d ago2 min read
Claude CodeAI Evaluation
AI botched my quote by 30% — the free 5-step check OpenAI teams use
Shreya and Hamel's free method (4500+ users from OpenAI/Google teams) turns AI output from 'gut feel' to 'checkable'. No code needed, 2-3 hours to set
6d ago2 min read