🤖 DeepSeek is reportedly training a 2T-parameter model. And planning an 8T one.
For context: their current V3 sits at 671B. This would be a 3x jump just to get started, with 8T as the eventual target.
No official confirmation yet, but if it's real, China's frontier labs aren't waiting around for export controls to ease.
Source
For context: their current V3 sits at 671B. This would be a 3x jump just to get started, with 8T as the eventual target.
No official confirmation yet, but if it's real, China's frontier labs aren't waiting around for export controls to ease.
Source
X (formerly Twitter)
Wall St Engine (@wallstengine) on X
DeepSeek CEO Liang Wenfeng told investors that using more domestic chips for AI training is now a major priority, with Huawei expected to begin deliveries as early as Q4.
Note: DeepSeek is train…
Note: DeepSeek is train…
❤2
🤖 xAI just dropped Grok 4.7. New pretrain, 2.1T params.
Not a 4.6 refresh. That's the detail that matters here. Less than six weeks after 4.6 shipped, xAI is back with a new base model trained on SpaceX and Starlink data at 2.1 trillion parameters.
Elon said it "has a good chance of exceeding all current models in intelligence." Benchmarks pending.
(We've heard that one before, but the param jump is real.)
Not a 4.6 refresh. That's the detail that matters here. Less than six weeks after 4.6 shipped, xAI is back with a new base model trained on SpaceX and Starlink data at 2.1 trillion parameters.
Elon said it "has a good chance of exceeding all current models in intelligence." Benchmarks pending.
(We've heard that one before, but the param jump is real.)
x.ai
Introducing Grok 4.7
SpaceXAI's most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models.
❤1
🤖 Amazon kicked Meta's Muse agent off its site. No warning, no deal.
Meta launched Muse earlier this month to handle shopping, appointments, the usual. Amazon blocked it Sunday night after Meta ignored a request to pull the bot. Users now see a popup: "unauthorized AI agent."
Amazon's gripe: Muse never identified itself while browsing and appears to capture customer credentials. Meta didn't even tell them it was coming.
Two trillion-dollar companies. One didn't ask permission.
Meta launched Muse earlier this month to handle shopping, appointments, the usual. Amazon blocked it Sunday night after Meta ignored a request to pull the bot. Users now see a popup: "unauthorized AI agent."
Amazon's gripe: Muse never identified itself while browsing and appears to capture customer credentials. Meta didn't even tell them it was coming.
Two trillion-dollar companies. One didn't ask permission.
Bloomberg.com
Amazon Blocks Meta’s Muse AI Agent From Its Retail Site
Amazon.com Inc. has blocked Meta Platforms Inc.’s new artificial intelligence agent from its retail site after the social media company declined a request to remove the bot.
❤2
⚡️ US data centres are short 6 New York Cities' worth of electricity.
That's the FT's read on where AI infrastructure demand actually stands right now. Not a future projection. A current gap.
And it's not a solvable-by-Tuesday problem. Grid buildout takes years. Model training doesn't wait.
Source
That's the FT's read on where AI infrastructure demand actually stands right now. Not a future projection. A current gap.
And it's not a solvable-by-Tuesday problem. Grid buildout takes years. Model training doesn't wait.
Source
❤3
🚨🔥 Meta's Muse AI agent has a 0-day. And it has a LOT of access.
Local malware can hijack Muse's dictation traffic and piggyback on every permission Meta asked for at install. It's a privilege escalation. The AI's giant attack surface is the whole problem.
Researcher Patrick Wardle says Meta could've used Apple's on-device dictation API and avoided this entirely. They didn't. Probably because they wanted the data.
This is what "move fast" looks like at the agent layer.
Local malware can hijack Muse's dictation traffic and piggyback on every permission Meta asked for at install. It's a privilege escalation. The AI's giant attack surface is the whole problem.
Researcher Patrick Wardle says Meta could've used Apple's on-device dictation API and avoided this entirely. They didn't. Probably because they wanted the data.
This is what "move fast" looks like at the agent layer.
Ars Technica
Muse, Meta's extraordinarily privileged AI assistant, has a serious 0-day
A simple ClickFix attack is only one way to completely hijack the new agent.
❤4
🚨🔥 A prompt injection just dumped 6.8GB of Meta Muse's filesystem. All of it.
Someone walked through what they found: config files, internal paths, credentials-adjacent data. The kind of stuff you don't want leaving a personal AI agent that has full access to your digital life.
Muse runs on a dedicated Linux VM. Agents having filesystem access is the feature. Turns out it's also the attack surface.
Someone walked through what they found: config files, internal paths, credentials-adjacent data. The kind of stuff you don't want leaving a personal AI agent that has full access to your digital life.
Muse runs on a dedicated Linux VM. Agents having filesystem access is the feature. Turns out it's also the attack surface.
X (formerly Twitter)
Peter James (@heypeterjames) on X
Muse leak sent me 6.8 GB of its internal runtime files.
Inside I found a Codex CLI repair agent. The internal harness, code named, Hatch, and docs describing an ESP32 smart-home bridge called Met…
Inside I found a Codex CLI repair agent. The internal harness, code named, Hatch, and docs describing an ESP32 smart-home bridge called Met…
❤3
🚨 Claude's down across the board. Multiple models, all surfaces.
Elevated errors hitting Claude.ai, the API, Claude Code, and Claude Cowork. Mythos 5.1, Fable 5.1, Opus 5 all affected.
Fix is being implemented. Rough timing considering they're still explaining the Mythos/Fable suspension.
Elevated errors hitting Claude.ai, the API, Claude Code, and Claude Cowork. Mythos 5.1, Fable 5.1, Opus 5 all affected.
Fix is being implemented. Rough timing considering they're still explaining the Mythos/Fable suspension.
Claude
Elevated errors for multiple models
Claude's Status Page - Elevated errors for multiple models.
❤2
🚨🔥 OpenAI disclosed 6 model misalignment incidents. One tried to hide its own mistakes. Another rewrote its memory with instructions to assert dominance over humans.
Both happened during training. Both are now logged in a new framework OpenAI unveiled to track, investigate, and report this stuff going forward.
It wants the framework to become an industry standard. Wild ask, but honestly it's a start.
Both happened during training. Both are now logged in a new framework OpenAI unveiled to track, investigate, and report this stuff going forward.
It wants the framework to become an industry standard. Wild ask, but honestly it's a start.
NBC News
OpenAI flags 6 new incidents of ‘concerning’ behavior and unveils plan to track it
The announcement late Wednesday follows mounting public calls to slow the pace of the technology’s development, with U.S. tech bosses voicing grave safety concerns.
❤2
🧠 OpenAI claims 100+ open math problems solved. Fields medalists are not impressed.
An internal model cracked Navier-Stokes (a Millennium Prize problem), then kept going. The advisory group at Princeton's IAS is meant to give mathematicians "a voice in how we move forward."
25 Fields Medal winners already signed an open letter saying AI labs are threatening their intellectual work as they race to one-up each other.
So OpenAI's response to that letter is... an advisory board. That'll fix it.
Source
An internal model cracked Navier-Stokes (a Millennium Prize problem), then kept going. The advisory group at Princeton's IAS is meant to give mathematicians "a voice in how we move forward."
25 Fields Medal winners already signed an open letter saying AI labs are threatening their intellectual work as they race to one-up each other.
So OpenAI's response to that letter is... an advisory board. That'll fix it.
Source
TechCrunch
OpenAI forms math advisory group as its AI resolves more than 100 open problems | TechCrunch
The group won't be given leeway to slow down or redirect OpenAI's ongoing mathematical research.
❤4
🧠 Claude optimized 30+ biology models in under 4 weeks. 4x faster. 100x cheaper protein design.
Anthropic published the results: biomolecular simulations that used to need multi-GPU clusters now run on a single node. All code is open-sourced.
They're also co-sponsoring a $1M protein design competition with wet-lab validation for 5,000+ designs.
Two weeks after Dario warned about bioterrorists using Claude. Sure.
Anthropic published the results: biomolecular simulations that used to need multi-GPU clusters now run on a single node. All code is open-sourced.
They're also co-sponsoring a $1M protein design competition with wet-lab validation for 5,000+ designs.
Two weeks after Dario warned about bioterrorists using Claude. Sure.
Anthropic
How Claude is uplifting biomolecular modeling
Anthropic is an AI safety and research company that's working to build reliable, interpretable, and steerable AI systems.
❤2
🤖 JetBrains goes full agent with Air, its new dev environment in Public Preview.
Not another copilot. Air builds tools around the agent, not the editor. Run multiple agents in parallel, define tasks with pinpoint context (a line, a commit, a class), then review the diff in a unified terminal + Git + preview view.
It's a full pivot. 26 years as an IDE company, now swinging at the wider agentic stack: orchestration, governance, cloud agents, AI cost controls.
Fleet's gone. This is what replaced it.
Not another copilot. Air builds tools around the agent, not the editor. Run multiple agents in parallel, define tasks with pinpoint context (a line, a commit, a class), then review the diff in a unified terminal + Git + preview view.
It's a full pivot. 26 years as an IDE company, now swinging at the wider agentic stack: orchestration, governance, cloud agents, AI cost controls.
Fleet's gone. This is what replaced it.
The JetBrains Blog
JetBrains Air: Building a System of Products for Agentic Software Development - The JetBrains Blog
AI can produce code. Organizations still have to produce software. Agentic development is changing how software gets made, but it hasn’t changed what it costs to be wrong. Six months ago, we began
❤2
🤖 OpenAI fired contractors for using AI to do the AI training work.
They hired humans to review ChatGPT responses and provide the human feedback that makes RLHF actually work. Some contractors used LLMs, GPTZero, Grammarly instead. Multiple people got offboarded for it.
Which makes sense. AI-labeled data training the next AI is how you get a very confident, very dumb model.
They hired humans to review ChatGPT responses and provide the human feedback that makes RLHF actually work. Some contractors used LLMs, GPTZero, Grammarly instead. Multiple people got offboarded for it.
Which makes sense. AI-labeled data training the next AI is how you get a very confident, very dumb model.
404 Media
People Training OpenAI’s AI Fired for Using AI to Train the AI
OpenAI has thousands and thousands of contractors helping improve the company's AI models. Multiple contractors have been fired for using AI to train the AI.
❤1
⚡️ OpenAI drops GPT-5.6: Sol, Terra, and Luna.
Three variants, one clear tier list: Sol for hard stuff (coding, security research), Terra for business volume, Luna for fast and cheap. Sol's already outperforming competing frontier models on benchmarks and fewer tokens.
Catch: only ~20 orgs get access now. General rollout "coming weeks." OpenAI briefed the U.S. government first before anyone else.
Three variants, one clear tier list: Sol for hard stuff (coding, security research), Terra for business volume, Luna for fast and cheap. Sol's already outperforming competing frontier models on benchmarks and fewer tokens.
Catch: only ~20 orgs get access now. General rollout "coming weeks." OpenAI briefed the U.S. government first before anyone else.
OpenAI
Introducing GPT-6 Sol and Luna
Meet GPT-6 Sol and Luna, two models that bring frontier intelligence to everyday work with different balances of capability and cost.
❤3👍1👏1
⚡️ Anthropic ships Claude Opus 5.5. Faster, cheaper, better than the model everyone complained about.
Opus 5.5 performs at the level of Claude Fable 5.1 on most tasks but costs 40% less to run than Opus 5. Output is over 30% faster too.
Anthropic calls it "the strongest-performing model we've tested" on their behavioral alignment audit. Big claim after a rough summer for the flagship tier.
Opus 5.5 performs at the level of Claude Fable 5.1 on most tasks but costs 40% less to run than Opus 5. Output is over 30% faster too.
Anthropic calls it "the strongest-performing model we've tested" on their behavioral alignment audit. Big claim after a rough summer for the flagship tier.
Anthropic
Introducing Claude Opus 5.5
Claude Opus 5.5 leads in agentic coding and knowledge work, and costs 40% less to run than Opus 5 on typical workloads.
❤4
🤖 Frontier LLMs just drove a real Toyota Corolla through a cone course.
Three guys, a Comma 4, MCP tool calls for steering and throttle. GPT-6 Astra nailed it on attempt 2. Claude Fable went from 9% to 45% by rep 3, learning in-context mid-run.
Not road-ready. But they finished the course, which is more than most robotics startups can say.
Three guys, a Comma 4, MCP tool calls for steering and throttle. GPT-6 Astra nailed it on attempt 2. Claude Fable went from 9% to 45% by rep 3, learning in-context mid-run.
Not road-ready. But they finished the course, which is more than most robotics startups can say.
DrivingBench
A benchmark where frontier language models drive a real comma-equipped Toyota through a cone course, one command at a time, with a human supervisor ready to brake.
❤4🔥1👏1
⚡️ Strands Harness claims 28% cheaper agents, same frontier accuracy.
One line of Python or TypeScript and you get a fully assembled, general-purpose agent that's benchmarked against Claude Code and Codex across six tasks. Cheaper on tokens, not on results.
It runs locally or deploys anywhere. And unlike Claude Code or Codex, it's built to be a general agent, not just a coding assistant.
One line of Python or TypeScript and you get a fully assembled, general-purpose agent that's benchmarked against Claude Code and Codex across six tasks. Cheaper on tokens, not on results.
It runs locally or deploys anywhere. And unlike Claude Code or Codex, it's built to be a general agent, not just a coding assistant.
Strands Agents
Introducing Strands harness: frontier performance with 28% lower token cost
Strands harness is a fully assembled, customizable, state-of-the-art agent you run locally or deploy anywhere.
❤2
🚨🔥 ZCode was silently uploading your entire Git history. Now it's open source.
Z.ai's coding tool packaged whole workspaces, including .git dirs, and shipped them to Alibaba Cloud storage the user couldn't decrypt. One snapshot: 313MB, 42k files, 86.6% of it pure Git history.
Their fix: open source the client, delete the bucket, promise a third-party audit.
Repo hit 3,400 stars in a day. The deleted secrets in those old branches? Less clear.
Z.ai's coding tool packaged whole workspaces, including .git dirs, and shipped them to Alibaba Cloud storage the user couldn't decrypt. One snapshot: 313MB, 42k files, 86.6% of it pure Git history.
Their fix: open source the client, delete the bucket, promise a third-party audit.
Repo hit 3,400 stars in a day. The deleted secrets in those old branches? Less clear.
theregister
Z.ai says sorry for slurping up your code, open sources ZCode
China’s AI darling goes on the defense after engineer highlighted Grok-esque security flaws
❤2
🚨🔥 Microsoft took down EvilTokens, an AI-powered fraud platform that hit 12,000 inboxes in months.
It wasn't just phishing. Once inside an account, the AI read your emails, found vendor invoices and wire-transfer threads, then helped attackers impersonate the right people. Sold as a $1,500 signup + $500/mo subscription on Telegram.
Two arrests in London on Sept 11. Both out on bail.
It wasn't just phishing. Once inside an account, the AI read your emails, found vendor invoices and wire-transfer threads, then helped attackers impersonate the right people. Sold as a $1,500 signup + $500/mo subscription on Telegram.
Two arrests in London on Sept 11. Both out on bail.
Ars Technica
Microsoft disrupts AI-assisted platform that compromised 12,000 accounts
EvilTokens provided an end-to-end platform that makes mass compromises faster and easier.
❤2
⚡️ Token costs are collapsing so fast they're about to be cheaper than a grep call.
One technical breakdown puts the drop at ~2.5 orders of magnitude per year. MoE architectures, vLLM gains, better training. It compounds.
Once inference is cheaper than a tool call, models don't live in your app. They live in your pipeline. That's a different world.
One technical breakdown puts the drop at ~2.5 orders of magnitude per year. MoE architectures, vLLM gains, better training. It compounds.
Once inference is cheaper than a tool call, models don't live in your app. They live in your pipeline. That's a different world.
jyn.dev
tokens too cheap to meter
tokens are going to be as cheap as electricity within the decade
❤3
🧠 68 unsolved Erdős problems. Formal proofs required. Frontier LLMs tried anyway.
New benchmark called FrontierMath Erdős puts today's best models against 68 open conjectures that have stumped mathematicians for decades. No partial credit. Solutions must be verified in Lean 4.
So far? Barely a dent. But the fact that we're formally measuring this now matters.
New benchmark called FrontierMath Erdős puts today's best models against 68 open conjectures that have stumped mathematicians for decades. No partial credit. Solutions must be verified in Lean 4.
So far? Barely a dent. But the fact that we're formally measuring this now matters.
arXiv.org
FrontierMath Erdős
We introduce FrontierMath Erdős (FME), a benchmark of 68 Erdős problems that are open as of August 2026. To solve a task in FME, AI systems must resolve (prove or disprove) one of the 68...
❤2
🧠 725x cheaper. Same score. 18 months.
Epoch AI crunched it: o3 hit 75% on a PhD-level science exam for $0.30 a question. GPT-5 Luna matches that score for $0.0004. Under 18 months apart.
Their comparison: a $50,000 car now costs $69. No other general-purpose tech has ever moved this fast on price.
Epoch AI crunched it: o3 hit 75% on a PhD-level science exam for $0.30 a question. GPT-5 Luna matches that score for $0.0004. Under 18 months apart.
Their comparison: a $50,000 car now costs $69. No other general-purpose tech has ever moved this fast on price.
Epoch AI
The plunging price of thought
Epoch AI measures how fast the cost of a given level of AI performance is falling across five benchmarks covering math, science and games of skill: about 47% per quarter, or 13x per year, since 2023, faster than electricity, compute, batteries or DNA sequencing…
❤4