BenchMIRT: What Are Your Benchmark Scores Actually Measuring?
Allen AI's new auditing tool reveals that many LLM benchmarks mix multiple capabilities into a single score—and shows which questions actually matter.
14 posts
Allen AI's new auditing tool reveals that many LLM benchmarks mix multiple capabilities into a single score—and shows which questions actually matter.
Hugging Face's ASR leaderboard just added Hindi and Indian English with metadata on 4,888 speakers across hundreds of districts. Benchmarks decide what gets built—this one might actually be fair.
Researchers introduce three probes that catch speech models reproducing benchmark transcripts—even when the audio contradicts them, words are silenced, or spellings should vary randomly.
IBM's ALTK-Evolve reveals a counterintuitive finding: more agentic memory isn't always better. The right dose depends on the model—and sometimes the cheapest strategy wins.
IBM Research's ALTK-Evolve matches or beats ACE's agentic memory performance at 15–40% of the inference cost. The secret: calibrated delivery instead of injecting the whole playbook every time.
Allen AI's new benchmark reveals that LLMs trained to be helpful assistants struggle with tutoring's core trade-off: knowing when students need support versus when they need productive struggle.
OpenAI's GPT-5.6 Sol jumped from 13.3% to 38.3% on ARC-AGI-3 by keeping reasoning in context and using compaction. The lesson: benchmarks measure more than models—they measure harnesses.
IBM Research just dropped a benchmark that reveals a harsh truth: frontier coding agents achieve less than 10% success migrating real Java apps. The problem isn't code—it's everything else.
ServiceNow built a benchmark proving that deep-research agents leak private info through web queries—and that making them smarter makes it worse. Privacy-aware RL cuts leakage by 70%.
ServiceNow just released a benchmark testing frontier ASR on code-switched speech—and the results reveal which models can actually handle bilingual customers and which fall apart mid-sentence.
ServiceNow's new voice-agent benchmark spans airlines, IT, and healthcare—with joint-generation pipelines, adversarial scenarios, and a coming multilingual expansion.
Frontier models score below 50% on Kubernetes incident response. The new ITBench-AA benchmark from Artificial Analysis and IBM reveals the gap between agent demos and production IT work.
The Open ASR Leaderboard is fighting back against benchmaxxing with a simple but effective strategy: private evaluation datasets that no one can train on.
TII launches QIMMA, a rigorous quality-focused leaderboard for Arabic LLMs that goes beyond translation metrics to measure genuine language understanding and cultural nuance.