GPT-5.6 Sol is calibrating qubits while you sleep
An MIT grad student hooked Codex up to a dilution refrigerator and let GPT-5.6 Sol run quantum computing experiments overnight. The agent autonomously calibrated superconducting qubits.
50 posts
An MIT grad student hooked Codex up to a dilution refrigerator and let GPT-5.6 Sol run quantum computing experiments overnight. The agent autonomously calibrated superconducting qubits.
OpenAI just published hard numbers on how coding agents are reshaping AI research from the inside. The median researcher now burns $600/day in inference. Agent labor exceeds human labor 3:1. This is RSI's opening act.
Polimill's QommonsAI now serves 1,050 municipalities across Japan. It's a case study in shipping AI infrastructure at national scale—and what happens when you treat government as a platform.
IBM just dropped Granite 4.2—3B, 8B, and 30B reasoning models trained from scratch on 15T tokens, with thinking/non-thinking modes and a staged RL pipeline that teaches agents to code, search, and use terminals.
OpenAI just dropped GPT-5.6 into Kiro, their spec-driven coding agent. The headline? An 82% cost reduction on real dev tasks. Here's what actually changed—and why it matters.
Google DeepMind is shifting from beating games to building inside them—partnering with EVE Online, No Man's Sky, and others to prototype breakthrough AI gameplay that could reshape how games are made.
IBM's ALTK-Evolve reveals a counterintuitive finding: more agentic memory isn't always better. The right dose depends on the model—and sometimes the cheapest strategy wins.
OpenAI's Greg Brockman lays out why defenders have maybe six months to get their AI-powered security act together—and exactly what to do right now.
Hugging Face ran a 19-day hackathon where 1,200 people used coding agents to reproduce a third of ICML 2026. The results expose both the conference flood problem and what humans are still for.
Hugging Face Storage Buckets just made the robot data flywheel real. Record demonstrations, train by streaming from the Hub, deploy checkpoints—all without re-downloading gigabytes.
Liquid's new 3B vision-language model beats larger 4B competitors on grounding and screens, decodes 228 tok/s on an M5 Max, and ships with llama.cpp support day one. This is what edge-first design looks like.
IBM Research's ALTK-Evolve matches or beats ACE's agentic memory performance at 15–40% of the inference cost. The secret: calibrated delivery instead of injecting the whole playbook every time.
Liquid AI's new 2.6B model beats competitors 4x its size on tool use and instruction following, runs at 220 tok/s on a laptop, and was trained inside real agent harnesses with RL.
Circles achieved 22% ARPU lift and 65% autonomous support resolution with OpenAI. But the case study raises more questions than it answers about architecture, cost, and replicability.
OpenAI slashed GPT-5.6 Luna pricing by 80% and Terra by 20%, making high-volume AI workflows economical at scale. Plus: Fast mode for Sol delivers 2.5× speedups. Here's why this matters.
OpenAI's GPT-5.6 Sol jumped from 13.3% to 38.3% on ARC-AGI-3 by keeping reasoning in context and using compaction. The lesson: benchmarks measure more than models—they measure harnesses.
Google just gave us a first look at Gemini Intelligence on Samsung's new foldables—multi-app task automation, on-device Notebook, and wrist-gesture glasses control. This is production AI, not a demo.
NVIDIA's new embedding models claim the top RTEB spot, but the interesting part isn't the leaderboard flex—it's the 1B variants optimized for Blackwell and what they reveal about production retrieval.
Ai2's maritime agent Shippy isn't about the model—it's about reliability, deterministic tools, sandboxed execution, and real evals. Here's what building an agent for high-stakes decisions actually looks like.
IBM Research tried routing requests across Claude, GPT-4, and Opus in production agentic systems. Token pricing didn't predict actual cost. Task difficulty didn't predict model fit. Here's why.
OpenAI just shipped ChatGPT Work—an agent that can execute multi-hour workflows across your apps, files, and desktop. Powered by GPT-5.6, it's the first real test of whether agentic AI can ship.
NVIDIA is releasing over 10 trillion pre-training tokens and millions of post-training samples for agent development—and building synthetic personas representing 2.4B people. Here's why that matters.
Google shipped an astonishing amount of AI in June 2026—real-time multilingual translation, computer-use agents, on-device models, and a genuinely conversational smart speaker. Here's what matters.
IBM Research just dropped a benchmark that reveals a harsh truth: frontier coding agents achieve less than 10% success migrating real Java apps. The problem isn't code—it's everything else.
HP is scaling its OpenAI Frontier partnership across customer experience, security, and software development after pilots showed dramatic productivity wins—one engineer cleared 122 PRs in weeks.
New OpenAI data shows Codex now accounts for 99.8% of tokens inside the company. Non-developers are adopting agents 137x faster than before. This is what the shift from chatbots to agents actually looks like.
Google's relaunched Finance brings portfolio screenshots, custom briefings, and an Android app. The AI is impressive—but keeping it captive inside a walled garden feels like a missed opportunity.
IBM's open-source agent harness delivers two-dozen production-ready apps to show what happens when orchestration, guardrails, and tool-wiring come pre-assembled.
DeepMind just published their internal framework for securing AI agents: real-time monitoring, threat modeling borrowed from cybersecurity, and the assumption that alignment might fail.
HuggingFace's new agent benchmark doesn't just ask if the model got the right answer—it measures how much work it took to get there, across models, library versions, and task tiers.
ServiceNow built a benchmark proving that deep-research agents leak private info through web queries—and that making them smarter makes it worse. Privacy-aware RL cuts leakage by 70%.
AWS open-sourced Strands Robots, an SDK that exposes LeRobot's stack as composable agent tools. Record demos in sim, push to Hub, run policies, deploy to hardware—all in one agent.
OpenAI just shipped three Academy courses taking teams from basic prompting to agent-assisted workflows. This is learning-as-deployment, and it matters more than you think.
Cohere just released a 30B MoE model trained specifically for agentic software engineering. It's Apache 2.0, beats models 4× its size, and actually works across multiple agent harnesses.
Hugging Face, Meta PyTorch, Nvidia, and a dozen others just formed a committee to govern OpenEnv—the protocol layer trying to make agentic RL training actually interoperable.
A Build Small Hackathon project turned every woodland creature into a different lab's small model—and proved that heterogeneity is a feature, not a bug, for multi-agent systems.
A Build Small Hackathon entry proves small models shine where frontier models fail: running multi-agent simulations in real-time. Lessons on scarcity,JSON reliability, and reskinning history.
Google just shipped an entire agentic stack in one month: Gemini 3.5 for multi-step workflows, Gemini Omni for multimodal creation, proactive Search agents, Universal Cart, and hardware purpose-built for it all.
H Company ships quantized weights, mobile support, and cross-framework compatibility. The computer-use agent stack just got real deployment options—including local inference on consumer hardware.
IBM Research argues LLMs alone can't scale in enterprise workflows. Their secret weapon? Software primitives that guide models through complex, regulated tasks at 30× lower cost.
Google just shipped two very different models at I/O 2026: Omni for conversational video editing and 3.5 Flash for long-horizon agent tasks. Here's what the demos reveal.
A 10,000-person software shop cut requirements analysis from weeks to hours by encoding senior judgment into Codex. Their playbook: treat it as a desktop agent, not a code assistant.
Frontier models score below 50% on Kubernetes incident response. The new ITBench-AA benchmark from Artificial Analysis and IBM reveals the gap between agent demos and production IT work.
The AI agent field moves fast, and its vocabulary moves faster. HuggingFace's new glossary finally draws clear lines between harness, scaffold, and agent—distinctions that matter.
OpenAI wants to manage your money. Their new ChatGPT finance feature raises hard questions about AI capabilities, privacy theater, and whether we're solving problems that actually exist.
OpenAI's latest customer spotlight shows how Parloa is using GPT models to power voice agents that don't make you want to throw your phone. Real-time, reliable, and surprisingly capable.
NVIDIA just dropped a 3B parameter multimodal model that processes documents, audio, and video with 128K context. Let's dig into what makes this nano model surprisingly capable.
Google just announced TPU v8, but instead of one chip, they're shipping two: v8T for training and v8I for inference. Here's why the bifurcation matters for AI's next phase.
NVIDIA's new Nemotron-based dataset gives developers 4,800 demographically grounded Korean personas to build culturally aware AI agents—a blueprint for non-English AI.
Hugging Face just dropped Ecom-RLVE, a reinforcement learning framework that trains e-commerce agents in realistic but controllable environments. This is how we move from chatbots to actually useful shopping assistants.