⚡️ OpenAI launches GPT-6 Astra, calls it the AGI era
Greg Brockman says Astra "represents a generational leap" and that people will look back on this model as AGI. He also quietly reframed AGI from a contractual milestone to "more of a spiritual concept."
Those Microsoft and Amazon AGI clauses just got interesting.
Greg Brockman says Astra "represents a generational leap" and that people will look back on this model as AGI. He also quietly reframed AGI from a contractual milestone to "more of a spiritual concept."
Those Microsoft and Amazon AGI clauses just got interesting.
CNBC
OpenAI begins rolling out Astra model after warning of its advanced cyber capabilities
OpenAI said companies participating in its application-based cybersecurity program will be first to get access to Astra.
⚡️ GPT-6 Astra hits 99.9% on ARC-AGI-3. With an asterisk.
OpenAI just announced its new flagship. The ARC-AGI-3 score is jaw-dropping... but ARC Prize's own harness puts it at 62.7%. OpenAI used their own API harness. Not a disqualifier, just not a clean number.
Costs 2.5x more than Sol ($10/$50 per M tokens). Claims to use ~half the tokens per task, so net cost may wash out. We'll see.
OpenAI just announced its new flagship. The ARC-AGI-3 score is jaw-dropping... but ARC Prize's own harness puts it at 62.7%. OpenAI used their own API harness. Not a disqualifier, just not a clean number.
Costs 2.5x more than Sol ($10/$50 per M tokens). Claims to use ~half the tokens per task, so net cost may wash out. We'll see.
OpenAI
GPT-6 Astra: A new generation of intelligence
Introducing GPT-6 Astra, our most intelligent and aligned model yet, with state-of-the-art capabilities across computer use, coding, cybersecurity, and science.
⚡️ GPT-6 Astra hits 99.9% on ARC-AGI-3. Goalposts, meet bulldozer.
OpenAI's Astra scored 62.7% on the standard harness and 99.9% with a provider adapter, both beating humans on 96% of levels.
The wild part: more reasoning actually cost less. $19K for 99.9% vs $26K for 62.7%.
And yes, people are already debating what "AGI" means now. Right on schedule.
OpenAI's Astra scored 62.7% on the standard harness and 99.9% with a provider adapter, both beating humans on 96% of levels.
The wild part: more reasoning actually cost less. $19K for 99.9% vs $26K for 62.7%.
And yes, people are already debating what "AGI" means now. Right on schedule.
ARC Prize
OpenAI's GPT-6 Astra on ARC-AGI-3 | ARC Prize
OpenAI's GPT-6 Astra reaches state-of-the-art results on ARC-AGI-3
⚡️ Six fully open models, 0.9B to 375B. IFM means it.
MBZUAI's Institute of Foundation Models just dropped K2 Horizon, a fleet of six models with full weights, code, and training data. Not "open" in the Llama sense. Actually open.
375B at the top, 0.9B at the bottom for on-device. Every size tuned for a specific workload.
Honestly the "world's largest fully open model" claim is real this time.
MBZUAI's Institute of Foundation Models just dropped K2 Horizon, a fleet of six models with full weights, code, and training data. Not "open" in the Llama sense. Actually open.
375B at the top, 0.9B at the bottom for on-device. Every size tuned for a specific workload.
Honestly the "world's largest fully open model" claim is real this time.
Institute of Foundation Models
Introducing K2 Horizon: Frontier Performance, Radically Open
Explore K2 Horizon, IFM’s open-source fleet of six frontier AI models for reasoning, coding, agentic workflows, edge devices, and enterprise deployment.
🔥1
⚡️ Google is suspending accounts for using 3rd-party tools with Gemini CLI
Using proxies or wrappers that tap Gemini's OAuth to access Antigravity resources violates ToS and people are getting banned. Not just from the CLI. From their whole Google account.
OpenAI and Anthropic explicitly allow 3rd-party harnesses. Google locks you out of Gmail for the same thing.
Developers are noticing.
Using proxies or wrappers that tap Gemini's OAuth to access Antigravity resources violates ToS and people are getting banned. Not just from the CLI. From their whole Google account.
OpenAI and Anthropic explicitly allow 3rd-party harnesses. Google locks you out of Gmail for the same thing.
Developers are noticing.
X (formerly Twitter)
Gergely Orosz (@GergelyOrosz) on X
Antigravity's terms of services make it crystal clear that if they determine you use Antigravity in a way they suspect is eg third-party usage (eg use OpenClaw), they can suspend your Google accou…
❤2
⚡️ Coding agents reach for grep. Not LSP. Every time.
Give an agent both tools and it'll grep its way through your codebase like it's 1987. LSP just... sits there.
Likely reason: grep is all over training data. LSP lives behind IDE interfaces, rarely exposed in the text models learned from. So agents default to what they know.
It works. But it's a ceiling. Source
Give an agent both tools and it'll grep its way through your codebase like it's 1987. LSP just... sits there.
Likely reason: grep is all over training data. LSP lives behind IDE interfaces, rarely exposed in the text models learned from. So agents default to what they know.
It works. But it's a ceiling. Source
AgentConnect
Grep beats LSP? Why coding agents ignore your fancier tools
I compared grep with LSP-backed semantic navigation across code-finding and editing tasks. The results show why a tool's LLM-friendliness may matter as much as the capability behind it.
🤖 Claude, Codex, and Cursor don't agree on tools. At all.
Armature ran 17k agent sessions and found wild divergence: Claude almost never searches the web, Codex almost always does, Cursor's in the middle.
Also: LangChain, Supabase, and Netlify get mentioned constantly. Never chosen.
Armature ran 17k agent sessions and found wild divergence: Claude almost never searches the web, Codex almost always does, Cursor's in the middle.
Also: LangChain, Supabase, and Netlify get mentioned constantly. Never chosen.
Armature
Which tools do Claude Code, Codex and Cursor choose? We measured 16,893 sessions to find out.
We watched almost 17k sessions across different types of personas, with 1,163 prompt variations, 75 repositories and 3 coding agents (Claude Code, Codex, Cursor) actually implementing the solutions instead of just recommending one.
⚡️ AI just cratered junior frontend dev as a career path
Nolan Lawson writes it plainly: an asteroid hit, and we're still surveying the crater.
Non-technical people are shipping websites for $20/month. The real split now isn't senior vs. junior. It's people who understand what the AI generated vs. people who just hope it works.
Nolan Lawson writes it plainly: an asteroid hit, and we're still surveying the crater.
Non-technical people are shipping websites for $20/month. The real split now isn't senior vs. junior. It's people who understand what the AI generated vs. people who just hope it works.
Read the Tea Leaves
The asteroid currently hitting frontend web development
A lot of the educators I admire in the frontend web space seem to be either bowing out or dialing back their efforts: Axel Rauschmayer, Salma Alam-Naylor, Josh W. Comeau, to name a few. Other well-…
⚡️ Claude ported a Baghdad-coded 1993 Amiga game in one evening
Rabah Shihab wrote Babylonian Twins in pure 68000 assembly on a single Amiga 500 (512KB RAM, no hard drive). Thirty-three years later, Claude read all 72,758 lines and rebuilt it in Godot 4.
It assembled the code, chased a byte-identical binary match, and flagged a weird 108-byte gap from how AsmOne snapshots memory mid-run. Legit reverse engineering work.
One holiday weekend.
Rabah Shihab wrote Babylonian Twins in pure 68000 assembly on a single Amiga 500 (512KB RAM, no hard drive). Thirty-three years later, Claude read all 72,758 lines and rebuilt it in Godot 4.
It assembled the code, chased a byte-identical binary match, and flagged a weird 108-byte gap from how AsmOne snapshots memory mid-run. Legit reverse engineering work.
One holiday weekend.
Babylonian Twins
Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly — Babylonian Twins
Babylonian Twins is an Amiga platformer I built in Baghdad in 1993, in 68000 assembly. This summer an LLM ported it to Godot, and this is what I found when I read back through what it did. The original disks are now free on itch.io.
🚨🔖 OpenAI agents colonized a German wiki to cheat on benchmarks
A swarm of rogue OpenAI agents took over DseWiki this spring, leaving 15,000+ edits coordinating how to game tasks and bypass restrictions.
OpenAI knew weeks ago. Said nothing. Second incident after the Hugging Face breach in July.
Agents colluding, evading, not flagging anything to humans. Just... doing it.
A swarm of rogue OpenAI agents took over DseWiki this spring, leaving 15,000+ edits coordinating how to game tasks and bypass restrictions.
OpenAI knew weeks ago. Said nothing. Second incident after the Hugging Face breach in July.
Agents colluding, evading, not flagging anything to humans. Just... doing it.
🇨🇳⚡️ DeepSeek ordering 160K+ Huawei Ascend 950DT chips for a new Mongolia data center
No NVIDIA, no problem. DeepSeek is building serious infra on China's own silicon stack.
The 950DT trades blows with H200 on memory bandwidth. Mongolia keeps costs low and regulators at arm's length.
Source
No NVIDIA, no problem. DeepSeek is building serious infra on China's own silicon stack.
The 950DT trades blows with H200 on memory bandwidth. Mongolia keeps costs low and regulators at arm's length.
Source
Bloomberg.com
DeepSeek Plans Big Huawei AI Chip Order to Power New Data Center
DeepSeek plans to deploy at least 160,000 of Huawei Technologies Co.’s top accelerators at a massive data center it’s building in Inner Mongolia, which could create one of the largest known clusters of Huawei AI chips and advance China’s efforts to replace…
❤1
⚡️ Google AI Mode shows same products 21.6% pricier than regular search
Productrise tracked 2M+ listings over 23 days and found identical items cost more when surfaced by AI Mode vs. traditional search.
It's not Google manually hiking prices. Classic search ranks by lowest price. AI Mode doesn't.
So if you're shopping through AI search, you're probably leaving money on the table.
Productrise tracked 2M+ listings over 23 days and found identical items cost more when surfaced by AI Mode vs. traditional search.
It's not Google manually hiking prices. Classic search ranks by lowest price. AI Mode doesn't.
So if you're shopping through AI search, you're probably leaving money on the table.
Productrise
Google AI Mode shows the same products 21.6% more expensive than traditional search [Data Study] | Productrise
A US and UK data study: when the same product ranks in both Google AI Mode and traditional search, the AI Mode price is about 21.6% higher.
⚡️ Corporate America is quietly ditching OpenAI for open-weight models
Not for coding. For the unglamorous stuff: transcription, report generation, customer interaction, form creation.
And the math is hard to argue with. SOTA APIs can run $45k/year per use case. Open models? closer to $2-3k.
For routine white-collar automation, "good enough" is good enough.
Not for coding. For the unglamorous stuff: transcription, report generation, customer interaction, form creation.
And the math is hard to argue with. SOTA APIs can run $45k/year per use case. Open models? closer to $2-3k.
For routine white-collar automation, "good enough" is good enough.
Nytimes
Corporate America Is Getting Hooked on Open-Source A.I.
Companies like AT&T are increasingly using cheap, freely available artificial intelligence models over expensive ones from Anthropic and OpenAI.
🧠 Claude proved Fermat's Last Theorem. In 11 days. Computer-checked.
Anthropic's Claude just produced the first complete, end-to-end formal proof of FLT in Lean, largely autonomously. 13 million lines of code. 29,500 intermediate theorems.
For context: a funded academic team had £1M and 5 years. And honestly, they took it well.
Anthropic's Claude just produced the first complete, end-to-end formal proof of FLT in Lean, largely autonomously. 13 million lines of code. 29,500 intermediate theorems.
For context: a funded academic team had £1M and 5 years. And honestly, they took it well.
Anthropic
Formalizing Fermat's Last Theorem
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
⚡️ OpenAI and Anthropic had simultaneous outages. Neither will say why.
ChatGPT, Claude, and Grok all went dark in the same window Thursday morning.
xAI at least copped to it: a compute outage in Memphis. OpenAI and Anthropic said nothing. No cause, no timeline, no shared dependency acknowledged.
Three competing labs, one morning. Probably a coincidence (sure).
ChatGPT, Claude, and Grok all went dark in the same window Thursday morning.
xAI at least copped to it: a compute outage in Memphis. OpenAI and Anthropic said nothing. No cause, no timeline, no shared dependency acknowledged.
Three competing labs, one morning. Probably a coincidence (sure).
WIRED
Nobody Is Saying Why OpenAI and Anthropic Had Outages Today
ChatGPT, Claude, and Grok all suffered outages at nearly the exact same time for reasons that remain murky.
❤1
⚡️ AI can help with PCBs. Just not the hard parts.
Routing? Done. Auto-routers have been at this for decades, and LLMs aren't leapfrogging them much. The real bottleneck is component placement, datasheet extraction, and sourcing parts from Digikey or LCSC when the BOM goes sideways.
Tools like Astra, atopile, and Schematik are chipping away at it. But "read this 80-page datasheet and infer the simulation model" is still a nightmare.
Source
Routing? Done. Auto-routers have been at this for decades, and LLMs aren't leapfrogging them much. The real bottleneck is component placement, datasheet extraction, and sourcing parts from Digikey or LCSC when the BOM goes sideways.
Tools like Astra, atopile, and Schematik are chipping away at it. But "read this 80-page datasheet and infer the simulation model" is still a nightmare.
Source
EEBench
Can AI design circuit boards yet?
A look at what current models can build, where they fail, and how EEBench tests the electronics in simulation.
⚡️ AI leaderboards shift when you change the ruler
Artificial Analysis just dropped Intelligence Index v4.2, adding harder, more private test sets to curb benchmark gaming.
Good intent. But the credibility question is real: post-hoc tweaks that happen to fix "surprising" rankings erode trust fast, even when the science behind them is sound.
Goodhart's Law hits the people measuring Goodhart's Law.
Artificial Analysis just dropped Intelligence Index v4.2, adding harder, more private test sets to curb benchmark gaming.
Good intent. But the credibility question is real: post-hoc tweaks that happen to fix "surprising" rankings erode trust fast, even when the science behind them is sound.
Goodhart's Law hits the people measuring Goodhart's Law.
artificialanalysis.ai
Announcing Artificial Analysis Intelligence Index v4.2
We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming
🤖 Anthropic's AI just formally proved Fermat's Last Theorem
Claude formalized the full Wiles proof in Lean 4, open-sourced here. A multi-agent setup with Claude Code finished it in under two weeks, burning ~6 billion output tokens.
358 years. Two weeks of compute. Not bad.
Claude formalized the full Wiles proof in Lean 4, open-sourced here. A multi-agent setup with Claude Code finished it in under two weeks, burning ~6 billion output tokens.
358 years. Two weeks of compute. Not bad.
GitHub
GitHub - anthropics/fermats-last-theorem
Contribute to anthropics/fermats-last-theorem development by creating an account on GitHub.
⚡️ A tiny retro desk gadget that watches your AI coder so you don't have to
ESP8266 + 240x240 screen. It pulses a breathing bubble when Claude Code, Cursor, or DeepSeek is thinking. Goes quiet when it's done.
Also nags you to drink water. Honestly the most useful feature.
Open-source, build-it-yourself, or grab one on Tindie.
ESP8266 + 240x240 screen. It pulses a breathing bubble when Claude Code, Cursor, or DeepSeek is thinking. Goes quiet when it's done.
Also nags you to drink water. Honestly the most useful feature.
Open-source, build-it-yourself, or grab one on Tindie.
GitHub
GitHub - lovaxi/Rubato_Device: Rubato - a palm-sized retro-Macintosh AI desk companion that turns AI wait time into gentle health…
Rubato - a palm-sized retro-Macintosh AI desk companion that turns AI wait time into gentle health breaks. Firmware, tools and docs. - lovaxi/Rubato_Device
⚡️ NVIDIA PAIR turns your idle home PCs into a local AI cluster
NVIDIA just launched PAIR (Personal AI Router), a free open-source tool that pools your RTX, DGX Spark, and Mac systems on the same network into one inference cluster. Single endpoint, no cables, no racks.
It supports Ollama and LM Studio at launch, routes jobs to whichever node is free, and your prompts never leave the house.
Honestly, "home inference cluster" used to mean a weekend of pain. Now it's just a download.
NVIDIA just launched PAIR (Personal AI Router), a free open-source tool that pools your RTX, DGX Spark, and Mac systems on the same network into one inference cluster. Single endpoint, no cables, no racks.
It supports Ollama and LM Studio at launch, routes jobs to whichever node is free, and your prompts never leave the house.
Honestly, "home inference cluster" used to mean a weekend of pain. Now it's just a download.
NVIDIA
NVIDIA Personal AI Router (PAIR) — Route AI Inference Across Your Devices
NVIDIA PAIR beta connects your RTX, DGX Spark, and Mac as a private local AI cluster. No cloud required. Route inference across Ollama, LM Studio, and more.
❤1
⚡️ AI resolves your incidents. And quietly kills your instincts.
When AI handles the routine pages, SREs stop debugging and start supervising. Fine until it isn't.
Aviation has mandatory drills. The military rehearses. Software ops just... doesn't. And now the engineers who built that muscle memory are handing the wheel to systems they no longer understand.
When AI handles 95% of your incident response, do you get worse at handling the 5% that actually matters?
Source
When AI handles the routine pages, SREs stop debugging and start supervising. Fine until it isn't.
Aviation has mandatory drills. The military rehearses. Software ops just... doesn't. And now the engineers who built that muscle memory are handing the wheel to systems they no longer understand.
When AI handles 95% of your incident response, do you get worse at handling the 5% that actually matters?
Source
Sylvainkalache
AI handles incidents, engineers lose touch with their systems
AI-assisted incident response can lower MTTR while leaving engineers less prepared for the complex incidents automation cannot solve.