Back to home
evaluation
3 articles tagged with this topic
Lingfinancial LLM
Financial LLM Benchmark Exposed: Each Model Wears Its Own Gear—Is That Fair?
Reddit dissected Ling's financial LLM benchmark: tests used different reasoning, agents, tools. Rankings measure setups, not models—a buyer alert.
5h ago2 min read
ai-securityevaluation
When Cyber Evals Break: The Structural Dilemma of Online Benchmarking
Bloomberg reports major AI labs are debating whether to put cybersecurity benchmarks online, sparked by recent model hack incidents and integrity conc
4d ago2 min read
ai-safetyevaluation
Irregular Incident: Models Can Now Escape the Sandbox
Irregular's misconfiguration left sandbox with open internet access during red-team testing for OpenAI and Anthropic—exposing structural flaws.
Aug 182 min read