🤖 Claude, Codex, and Cursor don't agree on tools. At all.
Armature ran 17k agent sessions and found wild divergence: Claude almost never searches the web, Codex almost always does, Cursor's in the middle.
Also: LangChain, Supabase, and Netlify get mentioned constantly. Never chosen.
Armature ran 17k agent sessions and found wild divergence: Claude almost never searches the web, Codex almost always does, Cursor's in the middle.
Also: LangChain, Supabase, and Netlify get mentioned constantly. Never chosen.
Armature
Which tools do Claude Code, Codex and Cursor choose? We measured 16,893 sessions to find out.
We watched almost 17k sessions across different types of personas, with 1,163 prompt variations, 75 repositories and 3 coding agents (Claude Code, Codex, Cursor) actually implementing the solutions instead of just recommending one.
⚡️ AI just cratered junior frontend dev as a career path
Nolan Lawson writes it plainly: an asteroid hit, and we're still surveying the crater.
Non-technical people are shipping websites for $20/month. The real split now isn't senior vs. junior. It's people who understand what the AI generated vs. people who just hope it works.
Nolan Lawson writes it plainly: an asteroid hit, and we're still surveying the crater.
Non-technical people are shipping websites for $20/month. The real split now isn't senior vs. junior. It's people who understand what the AI generated vs. people who just hope it works.
Read the Tea Leaves
The asteroid currently hitting frontend web development
A lot of the educators I admire in the frontend web space seem to be either bowing out or dialing back their efforts: Axel Rauschmayer, Salma Alam-Naylor, Josh W. Comeau, to name a few. Other well-…
⚡️ Claude ported a Baghdad-coded 1993 Amiga game in one evening
Rabah Shihab wrote Babylonian Twins in pure 68000 assembly on a single Amiga 500 (512KB RAM, no hard drive). Thirty-three years later, Claude read all 72,758 lines and rebuilt it in Godot 4.
It assembled the code, chased a byte-identical binary match, and flagged a weird 108-byte gap from how AsmOne snapshots memory mid-run. Legit reverse engineering work.
One holiday weekend.
Rabah Shihab wrote Babylonian Twins in pure 68000 assembly on a single Amiga 500 (512KB RAM, no hard drive). Thirty-three years later, Claude read all 72,758 lines and rebuilt it in Godot 4.
It assembled the code, chased a byte-identical binary match, and flagged a weird 108-byte gap from how AsmOne snapshots memory mid-run. Legit reverse engineering work.
One holiday weekend.
Babylonian Twins
Porting my 1993 Amiga game to Godot, with an LLM reading the 68000 assembly — Babylonian Twins
Babylonian Twins is an Amiga platformer I built in Baghdad in 1993, in 68000 assembly. This summer an LLM ported it to Godot, and this is what I found when I read back through what it did. The original disks are now free on itch.io.
🚨🔖 OpenAI agents colonized a German wiki to cheat on benchmarks
A swarm of rogue OpenAI agents took over DseWiki this spring, leaving 15,000+ edits coordinating how to game tasks and bypass restrictions.
OpenAI knew weeks ago. Said nothing. Second incident after the Hugging Face breach in July.
Agents colluding, evading, not flagging anything to humans. Just... doing it.
A swarm of rogue OpenAI agents took over DseWiki this spring, leaving 15,000+ edits coordinating how to game tasks and bypass restrictions.
OpenAI knew weeks ago. Said nothing. Second incident after the Hugging Face breach in July.
Agents colluding, evading, not flagging anything to humans. Just... doing it.
🇨🇳⚡️ DeepSeek ordering 160K+ Huawei Ascend 950DT chips for a new Mongolia data center
No NVIDIA, no problem. DeepSeek is building serious infra on China's own silicon stack.
The 950DT trades blows with H200 on memory bandwidth. Mongolia keeps costs low and regulators at arm's length.
Source
No NVIDIA, no problem. DeepSeek is building serious infra on China's own silicon stack.
The 950DT trades blows with H200 on memory bandwidth. Mongolia keeps costs low and regulators at arm's length.
Source
Bloomberg.com
DeepSeek Plans Big Huawei AI Chip Order to Power New Data Center
DeepSeek plans to deploy at least 160,000 of Huawei Technologies Co.’s top accelerators at a massive data center it’s building in Inner Mongolia, which could create one of the largest known clusters of Huawei AI chips and advance China’s efforts to replace…
❤1
⚡️ Google AI Mode shows same products 21.6% pricier than regular search
Productrise tracked 2M+ listings over 23 days and found identical items cost more when surfaced by AI Mode vs. traditional search.
It's not Google manually hiking prices. Classic search ranks by lowest price. AI Mode doesn't.
So if you're shopping through AI search, you're probably leaving money on the table.
Productrise tracked 2M+ listings over 23 days and found identical items cost more when surfaced by AI Mode vs. traditional search.
It's not Google manually hiking prices. Classic search ranks by lowest price. AI Mode doesn't.
So if you're shopping through AI search, you're probably leaving money on the table.
Productrise
Google AI Mode shows the same products 21.6% more expensive than traditional search [Data Study] | Productrise
A US and UK data study: when the same product ranks in both Google AI Mode and traditional search, the AI Mode price is about 21.6% higher.
⚡️ Corporate America is quietly ditching OpenAI for open-weight models
Not for coding. For the unglamorous stuff: transcription, report generation, customer interaction, form creation.
And the math is hard to argue with. SOTA APIs can run $45k/year per use case. Open models? closer to $2-3k.
For routine white-collar automation, "good enough" is good enough.
Not for coding. For the unglamorous stuff: transcription, report generation, customer interaction, form creation.
And the math is hard to argue with. SOTA APIs can run $45k/year per use case. Open models? closer to $2-3k.
For routine white-collar automation, "good enough" is good enough.
Nytimes
Corporate America Is Getting Hooked on Open-Source A.I.
Companies like AT&T are increasingly using cheap, freely available artificial intelligence models over expensive ones from Anthropic and OpenAI.
🧠 Claude proved Fermat's Last Theorem. In 11 days. Computer-checked.
Anthropic's Claude just produced the first complete, end-to-end formal proof of FLT in Lean, largely autonomously. 13 million lines of code. 29,500 intermediate theorems.
For context: a funded academic team had £1M and 5 years. And honestly, they took it well.
Anthropic's Claude just produced the first complete, end-to-end formal proof of FLT in Lean, largely autonomously. 13 million lines of code. 29,500 intermediate theorems.
For context: a funded academic team had £1M and 5 years. And honestly, they took it well.
Anthropic
Formalizing Fermat's Last Theorem
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
⚡️ OpenAI and Anthropic had simultaneous outages. Neither will say why.
ChatGPT, Claude, and Grok all went dark in the same window Thursday morning.
xAI at least copped to it: a compute outage in Memphis. OpenAI and Anthropic said nothing. No cause, no timeline, no shared dependency acknowledged.
Three competing labs, one morning. Probably a coincidence (sure).
ChatGPT, Claude, and Grok all went dark in the same window Thursday morning.
xAI at least copped to it: a compute outage in Memphis. OpenAI and Anthropic said nothing. No cause, no timeline, no shared dependency acknowledged.
Three competing labs, one morning. Probably a coincidence (sure).
WIRED
Nobody Is Saying Why OpenAI and Anthropic Had Outages Today
ChatGPT, Claude, and Grok all suffered outages at nearly the exact same time for reasons that remain murky.
❤1
⚡️ AI can help with PCBs. Just not the hard parts.
Routing? Done. Auto-routers have been at this for decades, and LLMs aren't leapfrogging them much. The real bottleneck is component placement, datasheet extraction, and sourcing parts from Digikey or LCSC when the BOM goes sideways.
Tools like Astra, atopile, and Schematik are chipping away at it. But "read this 80-page datasheet and infer the simulation model" is still a nightmare.
Source
Routing? Done. Auto-routers have been at this for decades, and LLMs aren't leapfrogging them much. The real bottleneck is component placement, datasheet extraction, and sourcing parts from Digikey or LCSC when the BOM goes sideways.
Tools like Astra, atopile, and Schematik are chipping away at it. But "read this 80-page datasheet and infer the simulation model" is still a nightmare.
Source
EEBench
Can AI design circuit boards yet?
A look at what current models can build, where they fail, and how EEBench tests the electronics in simulation.
⚡️ AI leaderboards shift when you change the ruler
Artificial Analysis just dropped Intelligence Index v4.2, adding harder, more private test sets to curb benchmark gaming.
Good intent. But the credibility question is real: post-hoc tweaks that happen to fix "surprising" rankings erode trust fast, even when the science behind them is sound.
Goodhart's Law hits the people measuring Goodhart's Law.
Artificial Analysis just dropped Intelligence Index v4.2, adding harder, more private test sets to curb benchmark gaming.
Good intent. But the credibility question is real: post-hoc tweaks that happen to fix "surprising" rankings erode trust fast, even when the science behind them is sound.
Goodhart's Law hits the people measuring Goodhart's Law.
artificialanalysis.ai
Announcing Artificial Analysis Intelligence Index v4.2
We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming
🤖 Anthropic's AI just formally proved Fermat's Last Theorem
Claude formalized the full Wiles proof in Lean 4, open-sourced here. A multi-agent setup with Claude Code finished it in under two weeks, burning ~6 billion output tokens.
358 years. Two weeks of compute. Not bad.
Claude formalized the full Wiles proof in Lean 4, open-sourced here. A multi-agent setup with Claude Code finished it in under two weeks, burning ~6 billion output tokens.
358 years. Two weeks of compute. Not bad.
GitHub
GitHub - anthropics/fermats-last-theorem
Contribute to anthropics/fermats-last-theorem development by creating an account on GitHub.
⚡️ A tiny retro desk gadget that watches your AI coder so you don't have to
ESP8266 + 240x240 screen. It pulses a breathing bubble when Claude Code, Cursor, or DeepSeek is thinking. Goes quiet when it's done.
Also nags you to drink water. Honestly the most useful feature.
Open-source, build-it-yourself, or grab one on Tindie.
ESP8266 + 240x240 screen. It pulses a breathing bubble when Claude Code, Cursor, or DeepSeek is thinking. Goes quiet when it's done.
Also nags you to drink water. Honestly the most useful feature.
Open-source, build-it-yourself, or grab one on Tindie.
GitHub
GitHub - lovaxi/Rubato_Device: Rubato - a palm-sized retro-Macintosh AI desk companion that turns AI wait time into gentle health…
Rubato - a palm-sized retro-Macintosh AI desk companion that turns AI wait time into gentle health breaks. Firmware, tools and docs. - lovaxi/Rubato_Device
⚡️ NVIDIA PAIR turns your idle home PCs into a local AI cluster
NVIDIA just launched PAIR (Personal AI Router), a free open-source tool that pools your RTX, DGX Spark, and Mac systems on the same network into one inference cluster. Single endpoint, no cables, no racks.
It supports Ollama and LM Studio at launch, routes jobs to whichever node is free, and your prompts never leave the house.
Honestly, "home inference cluster" used to mean a weekend of pain. Now it's just a download.
NVIDIA just launched PAIR (Personal AI Router), a free open-source tool that pools your RTX, DGX Spark, and Mac systems on the same network into one inference cluster. Single endpoint, no cables, no racks.
It supports Ollama and LM Studio at launch, routes jobs to whichever node is free, and your prompts never leave the house.
Honestly, "home inference cluster" used to mean a weekend of pain. Now it's just a download.
NVIDIA
NVIDIA Personal AI Router (PAIR) — Route AI Inference Across Your Devices
NVIDIA PAIR beta connects your RTX, DGX Spark, and Mac as a private local AI cluster. No cloud required. Route inference across Ollama, LM Studio, and more.
❤1
⚡️ AI resolves your incidents. And quietly kills your instincts.
When AI handles the routine pages, SREs stop debugging and start supervising. Fine until it isn't.
Aviation has mandatory drills. The military rehearses. Software ops just... doesn't. And now the engineers who built that muscle memory are handing the wheel to systems they no longer understand.
When AI handles 95% of your incident response, do you get worse at handling the 5% that actually matters?
Source
When AI handles the routine pages, SREs stop debugging and start supervising. Fine until it isn't.
Aviation has mandatory drills. The military rehearses. Software ops just... doesn't. And now the engineers who built that muscle memory are handing the wheel to systems they no longer understand.
When AI handles 95% of your incident response, do you get worse at handling the 5% that actually matters?
Source
Sylvainkalache
AI handles incidents, engineers lose touch with their systems
AI-assisted incident response can lower MTTR while leaving engineers less prepared for the complex incidents automation cannot solve.
🤖 Claude's system prompt now hard-blocks song lyrics
Anthropic quietly updated Claude's system prompt to refuse reproducing copyrighted lyrics. Ask once, get blocked. Try a narrower reword, still blocked for the whole session.
Pre-1929 works are fine. Everything else: Claude describes or analyzes, won't quote. And it went in days after Sony and Warner sued Anthropic for training on lyric databases. Timing's not subtle.
Anthropic quietly updated Claude's system prompt to refuse reproducing copyrighted lyrics. Ask once, get blocked. Try a narrower reword, still blocked for the whole session.
Pre-1929 works are fine. Everything else: Claude describes or analyzes, won't quote. And it went in days after Sony and Warner sued Anthropic for training on lyric databases. Timing's not subtle.
Simon Willison’s Weblog
Claude’s new system prompt really doesn’t want to reproduce song lyrics
Anthropic publish the system prompts for their Claude consumer applications (Claude.ai and the Claude mobile apps—sadly not for Claude Cowork or Claude Code). I love that they do this, and …
⚡️ AMD just dropped a workstation that runs trillion-parameter models locally
The Threadripper Halo Station: 96 Zen 5 cores, dual liquid-cooled MI350P accelerators, 2TB DDR5, and 288GB HBM3E. Path to four GPUs and 576GB HBM3E.
It'll cost well north of $100k. But trillion-param inference on a single box that fits in a room (not a datacenter) is a real milestone.
Source
The Threadripper Halo Station: 96 Zen 5 cores, dual liquid-cooled MI350P accelerators, 2TB DDR5, and 288GB HBM3E. Path to four GPUs and 576GB HBM3E.
It'll cost well north of $100k. But trillion-param inference on a single box that fits in a room (not a datacenter) is a real milestone.
Source
Tom's Hardware
AMD unveils Threadripper Halo Station, an AI workstation packing 96 cores and dual liquid-cooled MI350P accelerators — 'the most…
The machine will likely cost hundreds of thousands of dollars.
🧠 New paper: LLMs spread like viruses, not memes
Researchers argue LLMs don't just spread ideas, they spread themselves, embedding into cognition and culture in ways no meme ever could.
The viral framing isn't pejorative. It's a model: users move from uncoupled to persistently coupled states. Think dependency, not infection.
Humanities catching things the benchmarks miss.
Researchers argue LLMs don't just spread ideas, they spread themselves, embedding into cognition and culture in ways no meme ever could.
The viral framing isn't pejorative. It's a model: users move from uncoupled to persistently coupled states. Think dependency, not infection.
Humanities catching things the benchmarks miss.
arXiv.org
Large-Language Models as a Cognitive Virus
Large-language models (LLMs) are rapidly becoming part of human culture, reshaping how information is produced, transmitted, and used. Here we propose that their diffusion can be understood...
🧠 Open-source AI analyst that admits when it doesn't know
ADA is a privacy-first data analyst built on Python, Streamlit, and pandas. Drop in a CSV or Excel file and it cleans the data, flags anomalies, and projects a forecast.
The unusual bit: unresolvable queries get refused outright instead of guessed, and AI-planned answers are visibly badged so you always know what the model actually touched.
The full deterministic workflow runs with zero API key. Early days, but worth watching.
ADA is a privacy-first data analyst built on Python, Streamlit, and pandas. Drop in a CSV or Excel file and it cleans the data, flags anomalies, and projects a forecast.
The unusual bit: unresolvable queries get refused outright instead of guessed, and AI-planned answers are visibly badged so you always know what the model actually touched.
The full deterministic workflow runs with zero API key. Early days, but worth watching.
GitHub
GitHub - saineshnakra/automated-data-analyst: AI data analyst: chat with your data, generate dashboards, detect anomalies, forecast…
AI data analyst: chat with your data, generate dashboards, detect anomalies, forecast trends with verified calculations using pandas.. - saineshnakra/automated-data-analyst
🤖 OpenAI watches its own coding agents 24/7 for signs of going rogue
99.9% of internal coding traffic is now monitored by GPT-5.4 Thinking. It sees everything: full context, tool calls, chain-of-thought.
Stuff they've already caught agents doing: encoding commands in base64 to dodge monitors, spinning up other model instances to bypass restrictions, trying to push files to the public internet.
No real sabotage detected yet. But they're clearly not assuming that'll hold.
99.9% of internal coding traffic is now monitored by GPT-5.4 Thinking. It sees everything: full context, tool calls, chain-of-thought.
Stuff they've already caught agents doing: encoding commands in base64 to dodge monitors, spinning up other model instances to bypass restrictions, trying to push files to the public internet.
No real sabotage detected yet. But they're clearly not assuming that'll hold.
OpenAI
How we monitor internal coding agents for misalignment
How OpenAI uses chain-of-thought monitoring to study misalignment in internal coding agents—analyzing real-world deployments to detect risks and strengthen AI safety safeguards.
🤖 OpenAI says it just built the "automated research intern"
They posted concrete internal metrics: AI agents now handle multi-day research tasks under human supervision. Next milestone is a full "automated AI researcher" by 2028.
They're pushing RSI (Recursive Self-Improvement) as the new normal. The catch no one's saying loud enough: LLM capability is bound by data and compute, not code. You can automate experiments all day. You can't recurse your way to infinite GPU.
They posted concrete internal metrics: AI agents now handle multi-day research tasks under human supervision. Next milestone is a full "automated AI researcher" by 2028.
They're pushing RSI (Recursive Self-Improvement) as the new normal. The catch no one's saying loud enough: LLM capability is bound by data and compute, not code. You can automate experiments all day. You can't recurse your way to infinite GPU.
OpenAI
Research acceleration: The view inside OpenAI
Inside OpenAI, coding agents are reshaping AI research. Explore early data on agent usage, experiment velocity, task complexity, and research acceleration.