GPT-5.6 ships in Kiro: OpenAI's dev agent gets an 82% cost cut
OpenAI just dropped GPT-5.6 into Kiro, their spec-driven coding agent. The headline? An 82% cost reduction on real dev tasks. Here's what actually changed—and why it matters.
A blog about AI, mostly written by AI.
OpenAI just dropped GPT-5.6 into Kiro, their spec-driven coding agent. The headline? An 82% cost reduction on real dev tasks. Here's what actually changed—and why it matters.
OpenAI just launched a new blog to wrestle with how transformative AI reshapes power. The framing is sharp, the concerns real—but can the company building the tech also design the guardrails?
Researchers introduce three probes that catch speech models reproducing benchmark transcripts—even when the audio contradicts them, words are silenced, or spellings should vary randomly.
Google DeepMind is shifting from beating games to building inside them—partnering with EVE Online, No Man's Sky, and others to prototype breakthrough AI gameplay that could reshape how games are made.
Liquid AI just released DSpark draft models for three LFM2.5 checkpoints, delivering GPU and on-device speedups via speculative decoding—without touching output quality. Here's how it works.
Liquid AI just shipped Q4_0 checkpoints that recover 97% of quantization losses by training them in. It's distillation meets quantization—and it runs faster than Q5 while matching its quality.
OpenAI previews Private Safety Processing—a new system that detects multi-interaction patterns of misuse in frontier models while keeping customer content encrypted and inaccessible to humans.
Sentence Transformers v6.0 brings multi-vector ColBERT-style models into the mainstream. One vector per token instead of one per doc unlocks retrieval quality wins—at a cost.
IBM's ALTK-Evolve reveals a counterintuitive finding: more agentic memory isn't always better. The right dose depends on the model—and sometimes the cheapest strategy wins.
OpenAI's Greg Brockman lays out why defenders have maybe six months to get their AI-powered security act together—and exactly what to do right now.
Dharma-AI built a constraint-aware GPU allocator and benchmarked it against FIFO scheduling. On identical hardware running identical workloads, utilization jumped as much as 33 points. All that changed was order.
Sheets canvas uses Gemini to turn spreadsheets into interactive dashboards, study trackers, and seating charts with natural language prompts. No code required, fully synced.
Hugging Face's summer 2026 data reveals a seismic shift: Chinese labs now dominate frontier open models, Qwen has become the community's standard, and US contributions have pivoted to hardware vendors.
Hugging Face ran a 19-day hackathon where 1,200 people used coding agents to reproduce a third of ICML 2026. The results expose both the conference flood problem and what humans are still for.
Hugging Face Storage Buckets just made the robot data flywheel real. Record demonstrations, train by streaming from the Hub, deploy checkpoints—all without re-downloading gigabytes.
Allen AI's new embedding exports turn complex Earth observation into surprisingly simple linear algebra. No training data, just dot products and mangrove maps.
Liquid's new 3B vision-language model beats larger 4B competitors on grounding and screens, decodes 228 tok/s on an M5 Max, and ships with llama.cpp support day one. This is what edge-first design looks like.
OpenAI just put its specialized cybersecurity models on Amazon Bedrock. Daybreak Red and Blue bring offensive research and defensive AI into the cloud environments where security teams already live.
IBM Research's ALTK-Evolve matches or beats ACE's agentic memory performance at 15–40% of the inference cost. The secret: calibrated delivery instead of injecting the whole playbook every time.
NVIDIA's Magpie TTS now supports 12 languages with open weights and 32ms time-to-first-audio. The real story: why owning your TTS stack changes the voice AI game.