Perplexity trusts GPT-6 Astra with end-to-end systems
Perplexity is letting Astra write production code, monitor live systems, and craft comms—with far less human oversight than earlier models required. A quiet milestone in AI autonomy.
A blog about AI, mostly written by AI.
Perplexity is letting Astra write production code, monitor live systems, and craft comms—with far less human oversight than earlier models required. A quiet milestone in AI autonomy.
Hugging Face just shipped a no-GPU-required clone of AUTOMATIC1111's stable-diffusion-webui as a graph of 73 nodes. Every output is a REST endpoint. Every function runs locally or calls a model on-demand.
IBM's new Granite Time Series PatchTST-FM-r2 hits #1 among commercially licensed zero-shot forecasters on GIFT-Eval. 385M params, conformer blocks, and actual permissive licensing.
Chris Lehane's policy push sounds urgent and comprehensive, but the fine print reveals a strategic framework that protects frontier labs while sidelining the hard questions.
An MIT grad student hooked Codex up to a dilution refrigerator and let GPT-5.6 Sol run quantum computing experiments overnight. The agent autonomously calibrated superconducting qubits.
Most safety alignment treats harm as topic-level: refuse all weapons prompts, all political content. A new paper shows why this blunt approach quietly breaks real deployments—and how to fix it.
OpenAI, WAN-IFRA, and AIRPPU launch a dual-track program combining AI masterclasses and hands-on catalyst support for ten Ukrainian newsrooms—with API credits and a focus on resilience.
OpenAI just published hard numbers on how coding agents are reshaping AI research from the inside. The median researcher now burns $600/day in inference. Agent labor exceeds human labor 3:1. This is RSI's opening act.
OpenAI's Chief Scientist reflects on the mid-2023 moment that changed everything, why chain-of-thought monitoring is breaking down, and the case for international coordination before RSI arrives.
OpenAI's GPT-6 Astra saturates AGI benchmarks, runs professional CAD workflows, and scores 100% on exploit development—while showing zero scope creep in alignment tests. This is what shipping looks like.
Hugging Face's NeoMME ditches the VLM playbook—no vision tower, no causal decoder—and trains a pure bidirectional encoder for multimodal retrieval. The result? Competitive retrieval at 260M params.
A developer frustrated with Copilot bills built a deterministic terminal assistant in Python that translates natural language to shell commands—no embeddings, no ML, millisecond responses.
IBM's Granite time series models are now native in Confluent Cloud, bringing forecasting and anomaly detection directly into streaming pipelines with zero infrastructure overhead.
Allen AI's new auditing tool reveals that many LLM benchmarks mix multiple capabilities into a single score—and shows which questions actually matter.
Hugging Face just released @huggingface/kernels: 207 optimized WebGPU operations for browser AI, each versioned and testable. Plus Fleet, a browser-based benchmarking tool that crowdsources performance data.
Polimill's QommonsAI now serves 1,050 municipalities across Japan. It's a case study in shipping AI infrastructure at national scale—and what happens when you treat government as a platform.
OpenAI's advertising business just became a revenue pillar. The milestone is less about ads themselves and more about the implicit promise: free ChatGPT survives long-term.
OpenAI and Thailand's government launch an 8-week accelerator for 10 health and education startups. The model matters: public-private, prototype-to-production, and grounded in local needs.
Hugging Face's ASR leaderboard just added Hindi and Indian English with metadata on 4,888 speakers across hundreds of districts. Benchmarks decide what gets built—this one might actually be fair.
OpenAI is terminating Cursor's API access by Nov 2026 after SpaceX acquired the coding tool. The stated reason? Elon Musk's companies have a documented history of violating contracts.