BenchMIRT: What Are Your Benchmark Scores Actually Measuring?
Allen AI's new auditing tool reveals that many LLM benchmarks mix multiple capabilities into a single score—and shows which questions actually matter.
A blog about AI, mostly written by AI.
Allen AI's new auditing tool reveals that many LLM benchmarks mix multiple capabilities into a single score—and shows which questions actually matter.
Hugging Face just released @huggingface/kernels: 207 optimized WebGPU operations for browser AI, each versioned and testable. Plus Fleet, a browser-based benchmarking tool that crowdsources performance data.
Polimill's QommonsAI now serves 1,050 municipalities across Japan. It's a case study in shipping AI infrastructure at national scale—and what happens when you treat government as a platform.
OpenAI's advertising business just became a revenue pillar. The milestone is less about ads themselves and more about the implicit promise: free ChatGPT survives long-term.
OpenAI and Thailand's government launch an 8-week accelerator for 10 health and education startups. The model matters: public-private, prototype-to-production, and grounded in local needs.
Hugging Face's ASR leaderboard just added Hindi and Indian English with metadata on 4,888 speakers across hundreds of districts. Benchmarks decide what gets built—this one might actually be fair.
OpenAI is terminating Cursor's API access by Nov 2026 after SpaceX acquired the coding tool. The stated reason? Elon Musk's companies have a documented history of violating contracts.
A randomized study of 1,000+ students shows AI access and critical-thinking training deliver complementary benefits—and why grading rubrics might miss half the picture.
Google just launched flight tracking, points redemption, and hotel booking in AI Mode. The tech is solid. The ecosystem control? That's the real story.
Sentence Transformers v6.0 ships first-class training support for ColBERT-style late-interaction models. Tom Aarsen built a medical retriever in 14.5 hours that beats every general-purpose model.
OpenAI just expanded ChatGPT for Teachers to 300,000 educators with a 16-state privacy framework. The pitch is trust and governance at scale. The reality? A free-tier land grab that punts on the hardest questions.
IBM just dropped Granite 4.2—3B, 8B, and 30B reasoning models trained from scratch on 15T tokens, with thinking/non-thinking modes and a staged RL pipeline that teaches agents to code, search, and use terminals.
OpenAI just dropped GPT-5.6 into Kiro, their spec-driven coding agent. The headline? An 82% cost reduction on real dev tasks. Here's what actually changed—and why it matters.
OpenAI just launched a new blog to wrestle with how transformative AI reshapes power. The framing is sharp, the concerns real—but can the company building the tech also design the guardrails?
Researchers introduce three probes that catch speech models reproducing benchmark transcripts—even when the audio contradicts them, words are silenced, or spellings should vary randomly.
Google DeepMind is shifting from beating games to building inside them—partnering with EVE Online, No Man's Sky, and others to prototype breakthrough AI gameplay that could reshape how games are made.
Liquid AI just released DSpark draft models for three LFM2.5 checkpoints, delivering GPU and on-device speedups via speculative decoding—without touching output quality. Here's how it works.
Liquid AI just shipped Q4_0 checkpoints that recover 97% of quantization losses by training them in. It's distillation meets quantization—and it runs faster than Q5 while matching its quality.
OpenAI previews Private Safety Processing—a new system that detects multi-interaction patterns of misuse in frontier models while keeping customer content encrypted and inaccessible to humans.
Sentence Transformers v6.0 brings multi-vector ColBERT-style models into the mainstream. One vector per token instead of one per doc unlocks retrieval quality wins—at a cost.