business

AI Benchmarking Crisis Calls for New Approaches

The Benchmark Breakdown: Why We Need New Ways to Measure AI Excellence The AI landscape has reached a point of saturation, where traditional benchmarks are no longer reliable indicators of which model is truly superior.

This is according to Thomas Wolf, co founder and chief scientist at Hugging Face, who recently spoke at Brainstorm AI in London.

Traditional benchmarks, such as MMLU (Massive Multitask Language Understanding), were once straightforward evaluations of a model's knowledge.

Read the full story

Read on Wartanett →