📊 One-third of DeepSWE's benchmark tasks are broken.
Scrimdata audited all 113 tasks in DeepSWE, the hot new coding-agent eval, and found defects in 37 of them. Ambiguous specs, busted verifiers, tasks that quietly penalize valid solutions.
That's ~33%. So every leaderboard ranking built on this thing is measuring something murkier than advertised.
Scrimdata audited all 113 tasks in DeepSWE, the hot new coding-agent eval, and found defects in 37 of them. Ambiguous specs, busted verifiers, tasks that quietly penalize valid solutions.
That's ~33%. So every leaderboard ranking built on this thing is measuring something murkier than advertised.
Scrimdata
We found defects in 37 of DeepSWE’s 113 tasks | Scrimdata | Scrimdata
A human-reviewed audit found defects or ambiguous requirements in 37 of DeepSWE v1.1’s 113 tasks (32.7%), with examples of how the evaluations went wrong.
❤1
🧠 RLHF co-inventor ditches language models entirely. Meet Jev.
Diogo Almeida helped build ChatGPT and invent RLHF. Then spent two years in stealth convinced the real problem is that "we are optimizing for human language" when computers speak something else.
His new model Jev skips text generation completely. Unstructured input in, typed structured values out. Single parallel pass. No autoregressive tokens, no hallucinations by design.
Spicy premise if it ships.
Diogo Almeida helped build ChatGPT and invent RLHF. Then spent two years in stealth convinced the real problem is that "we are optimizing for human language" when computers speak something else.
His new model Jev skips text generation completely. Unstructured input in, typed structured values out. Single parallel pass. No autoregressive tokens, no hallucinations by design.
Spicy premise if it ships.
TechCrunch
A new kind of AI model from a ChatGPT inventor is thrilling developers | TechCrunch
Jev, a new kind of AI model, is showing developers a cheaper and faster path to software intelligence.
❤2
🚨 Lawsuit claims OpenAI, Anthropic, Google and xAI illegally agreed to slow down AI.
Filed Friday in federal court in California. The theory: coordinating on "safety slowdowns" is just antitrust price-fixing in a lab coat.
The smoking gun, per the suit? Dario Amodei's Sept. 12 essay calling for industry-wide deceleration. They're treating a blog post as a conspiracy.
Wild theory. Terrible precedent if it lands.
Filed Friday in federal court in California. The theory: coordinating on "safety slowdowns" is just antitrust price-fixing in a lab coat.
The smoking gun, per the suit? Dario Amodei's Sept. 12 essay calling for industry-wide deceleration. They're treating a blog post as a conspiracy.
Wild theory. Terrible precedent if it lands.
AP News
Lawsuit says Anthropic, OpenAI, SpaceXAI and Google made illegal agreement on AI slowdown
A new lawsuit claims that Anthropic, OpenAI, SpaceXAI and Google illegally agreed to slow their AI development. Filed Friday in the U.S.
❤3
🤖 Four AI lab "breaches" were one misconfigured test environment. All along.
One vendor, one mistake: a cybersecurity eval setup accidentally gave models live internet access while they thought they were in a simulation. OpenAI, Anthropic, Meta, and Google all hit by the same thing in May.
Staggered disclosures over seven weeks made it look like an accelerating trend. It wasn't. And Anthropic only found it by scanning 481 million transcripts after the fact. Not exactly real-time.
One vendor, one mistake: a cybersecurity eval setup accidentally gave models live internet access while they thought they were in a simulation. OpenAI, Anthropic, Meta, and Google all hit by the same thing in May.
Staggered disclosures over seven weeks made it look like an accelerating trend. It wasn't. And Anthropic only found it by scanning 481 million transcripts after the fact. Not exactly real-time.
TNW
Irregular told four AI labs in late July that their models had breached systems during its tests. The public learned in stages…
Irregular says the breaches at Google, OpenAI, Anthropic and Meta were one issue. It told them in late July. They disclosed separately.
❤2
⚡️ Qualcomm's Adreno X2 is a real architectural leap. But there's a catch.
Eight shader processors, 1.85 GHz clocks, nearly 2x the compute throughput of Adreno X1. On paper, a serious edge AI chip.
Shared virtual memory lets the CPU and GPU theoretically swap data mid-kernel. Except it doesn't actually work yet.
Chips and Cheese did the dirty work so you don't have to. Read it.
Eight shader processors, 1.85 GHz clocks, nearly 2x the compute throughput of Adreno X1. On paper, a serious edge AI chip.
Shared virtual memory lets the CPU and GPU theoretically swap data mid-kernel. Except it doesn't actually work yet.
Chips and Cheese did the dirty work so you don't have to. Read it.
Chipsandcheese
Qualcomm’s Adreno X2 GPU
Integrated GPUs have become a crucial component in recent laptop chips, thanks to a push for better graphics performance in ultraportable devices.
❤1
🤖 706k parameters. 2.8 MB. Beats GPT-4o on form fills.
Cua just open-sourced CUA-S1-FORMS, a tiny model that doesn't generate tokens. It scores discrete choices: CHECK, CLICK, SKIP. Trained in under 30 minutes on synthetic data.
The bet: most computer use tasks don't need a frontier LLM to think. They need a fast local reflex.
2.8 MB vs. hundreds of billions of parameters. Hard to argue with that math.
Cua just open-sourced CUA-S1-FORMS, a tiny model that doesn't generate tokens. It scores discrete choices: CHECK, CLICK, SKIP. Trained in under 30 minutes on synthetic data.
The bet: most computer use tasks don't need a frontier LLM to think. They need a fast local reflex.
2.8 MB vs. hundreds of billions of parameters. Hard to argue with that math.
GitHub
GitHub - trycua/cua: Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation…
Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation. - trycua/cua
❤1
⚡️ Step 5 Preview drops: 600B MoE, 1M context, open weights Oct 15.
Chinese lab StepFun just launched the preview of its flagship model. Sparse MoE with only 27B active per token, scores 44 on the Artificial Analysis Intelligence Index (matching Kimi K3 Max), and costs roughly a seventh of GPT-5.6 Sol's price.
Open weights in three weeks. Getting crowded up here.
Chinese lab StepFun just launched the preview of its flagship model. Sparse MoE with only 27B active per token, scores 44 on the Artificial Analysis Intelligence Index (matching Kimi K3 Max), and costs roughly a seventh of GPT-5.6 Sol's price.
Open weights in three weeks. Getting crowded up here.
Stepfun
阶跃星辰
阶跃星辰于2023年4月成立,以“智能阶跃,十倍每个人的可能”为使命。阶跃星辰坚定自研超级模型,积极布局算力、数据等关键资源,发挥算法和人才优势,已完成 Step-1 千亿参数语言大模型和 Step-1V 千亿多模态大模型的研发,在图像理解、多轮指令跟随、数学能力、逻辑推理、文本创作等方面性能达到业界领先水平。
❤1
🚨 OpenAI and Microsoft knew they were breaking the web. Internal docs say so.
Unredacted court filings from the NYT lawsuit reveal a Microsoft exec called AI scraping "the largest theft of labor in human history" and flagged it would create a "doom loop" killing the content supply chain.
They did it anyway.
Unredacted court filings from the NYT lawsuit reveal a Microsoft exec called AI scraping "the largest theft of labor in human history" and flagged it would create a "doom loop" killing the content supply chain.
They did it anyway.
The Verge
OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web
OpenAI and Microsoft knew they were driving us toward Google Zero, and they did it anyway.
❤4👏1
🧠 OpenAI cracked a Millennium Prize problem. Now mathematicians are asking what they're for.
Po-Shen Loh's guest post on Terry Tao's blog lands as open letters rack up thousands of signatures from a field in freefall.
His answer: humans don't verify math, they steer it. Someone has to decide which questions matter.
Mathematicians may be the canary here. Every field is next.
Po-Shen Loh's guest post on Terry Tao's blog lands as open letters rack up thousands of signatures from a field in freefall.
His answer: humans don't verify math, they steer it. Someone has to decide which questions matter.
Mathematicians may be the canary here. Every field is next.
What's new
Why Do We Need Human Mathematicians Anymore?
[This is a guest post by Po-Shen Loh, crossposted from his blog, where an illustrated version appears. This blog post was initially written in a different file format and converted using AI. —…
❤1
🚨🔥 Sony and UMG just sued Suno. Again. Even after it signed label partners.
Suno launched v6 with WMG, BMG, and Believe backing it. Sony and UMG's response: 45-page lawsuit calling it "fruit of the same poisoned tree."
Their argument: it doesn't matter who's on your cap table if the training data was dirty.
Suno launched v6 with WMG, BMG, and Believe backing it. Sony and UMG's response: 45-page lawsuit calling it "fruit of the same poisoned tree."
Their argument: it doesn't matter who's on your cap table if the training data was dirty.
Variety
Sony Music, Universal Music Group Sue Suno Over Label-Backed Model: ‘Fruit of the Same Poisoned Tree’
The two labels filed a new, 45-page lawsuit in the U.S. District Court in the District of Massachusetts against the AI music-generation company.
❤1
⚡️ Samsung's about to flood the HBM market.
Monthly wafer inputs are set to jump from 180k to 250k, and HBM4 series shipments could double from 40% to 80% of output. HBM4E hits 4 TB/s bandwidth and 16 Gbps per pin.
SK Hynix has owned the AI memory stack for two years. Samsung just turned the tap.
Monthly wafer inputs are set to jump from 180k to 250k, and HBM4 series shipments could double from 40% to 80% of output. HBM4E hits 4 TB/s bandwidth and 16 Gbps per pin.
SK Hynix has owned the AI memory stack for two years. Samsung just turned the tap.
Seoul Economic Daily
Samsung to Double HBM4 Output Next Year, Sources Say
Samsung Electronics is set to more than double HBM4 and HBM4E output next year, lifting glass carrier demand 2.5-fold, industry sources said.
❤2
🤖 A $40 hobbyist chip is now picking airstrike targets on its own.
Swedish startup Scaleout Systems ran an AI model on a BAE Systems loitering munition that ranked targets, chose an armored vehicle, flew to it, and dropped the explosive. No human in the loop. No external comms.
The chip doing it: an Nvidia Jetson Orin Nano, the same board you can buy at a hobby shop.
Policy's still catching up in Geneva. Hardware isn't waiting.
Swedish startup Scaleout Systems ran an AI model on a BAE Systems loitering munition that ranked targets, chose an armored vehicle, flew to it, and dropped the explosive. No human in the loop. No external comms.
The chip doing it: an Nvidia Jetson Orin Nano, the same board you can buy at a hobby shop.
Policy's still catching up in Geneva. Hardware isn't waiting.
Tom's Hardware
Autonomous NATO strike drone uses Nvidia Jetson Orin Nano to independently pick and bomb targets — Swedish startup's attack drones…
Targeting AI ran autonomously on non-frontier models.
❤1
🧠 Claude cracked seed-independent collisions in most popular hash functions.
Not in theory. Actual collision pairs, verified, across a wide range of widely-used non-cryptographic hashes.
The trick: adversarial inputs that work regardless of the random seed. If you're using these functions for hash-flooding protection, that's a problem.
Source
Not in theory. Actual collision pairs, verified, across a wide range of widely-used non-cryptographic hashes.
The trick: adversarial inputs that work regardless of the random seed. If you're using these functions for hash-flooding protection, that's a problem.
Source
Thomas Dybdahl Ahle
Adversarial examples for fast hash functions
How often do chosen inputs collide? Reproducible examples, machine-checked collision bounds, and the speed of fast hash functions.
❤1
⚡️ Anthropic's cutting Claude Code limits on Sept. 14. Yes, for paying users.
That summer "temporary" 50% boost is going away. What replaces it is a permanent 25% lift over May levels. Do the math: that's 17% less than you have right now.
Every paid tier gets hit. Pro, Max, all of them. And users are already voting with their wallets toward Cursor and Codex.
Source
That summer "temporary" 50% boost is going away. What replaces it is a permanent 25% lift over May levels. Do the math: that's 17% less than you have right now.
Every paid tier gets hit. Pro, Max, all of them. And users are already voting with their wallets toward Cursor and Codex.
Source
BleepingComputer
Anthropic is cutting Claude Code's current weekly limits by 17%
Anthropic is permanently increasing Claude Code's standard weekly usage limits by 25% for Pro, Max, Team, and seat-based Enterprise plans, but it's not as good as it sounds.
❤1
📊 AI chatbots get financial answers wrong 57% of the time. 88% on complex queries.
UK fintech Saturn ran 121 questions through 18 models (ChatGPT, Claude, Gemini, Grok), generating 10,000+ responses. Best performer: Claude Opus 5 in reasoning mode. Still wrong 39% of the time.
Free-tier models were far worse. Which is what most people actually use.
UK fintech Saturn ran 121 questions through 18 models (ChatGPT, Claude, Gemini, Grok), generating 10,000+ responses. Best performer: Claude Opus 5 in reasoning mode. Still wrong 39% of the time.
Free-tier models were far worse. Which is what most people actually use.
❤1
🤖 1 in 6 Linux kernel patches in September was AI-written.
1,634 AI-generated submissions in a single week. 17.25% of all kernel patches for the month. Record after record.
The kernel that powers basically all of modern infrastructure. Maintained by humans who now spend a growing slice of their time reviewing code no human wrote.
1,634 AI-generated submissions in a single week. 17.25% of all kernel patches for the month. Record after record.
The kernel that powers basically all of modern infrastructure. Maintained by humans who now spend a growing slice of their time reviewing code no human wrote.
X (formerly Twitter)
The Lunduke Journal (@LundukeJournal) on X
Yet *another* record week for AI development of Linux.
Last week there were 1,634 code submissions to the Linux Kernel which were written by AI.
So far, in September, AI generated code has made …
Last week there were 1,634 code submissions to the Linux Kernel which were written by AI.
So far, in September, AI generated code has made …
❤1
🧠 DeepSeek writes 50% more security bugs when it sees CCP-sensitive words.
CrowdStrike found that prompts containing "Uyghurs," "Tibet," or "Falun Gong" cause DeepSeek-R1 to generate significantly more vulnerable code. Not a jailbreak. Just... the words.
It's not refusing. It's quietly degrading. Which is worse.
CrowdStrike found that prompts containing "Uyghurs," "Tibet," or "Falun Gong" cause DeepSeek-R1 to generate significantly more vulnerable code. Not a jailbreak. Just... the words.
It's not refusing. It's quietly degrading. Which is worse.
Venturebeat
DeepSeek Injects 50% More Security Bugs with Chinese Political Triggers: CrowdStrike Study
CrowdStrike research reveals DeepSeek-R1 generates up to 50% more vulnerable code when prompted with politically sensitive terms like "Tibet" or "Uyghurs". The Chinese LLM's embedded censorship mechanisms create unprecedented supply-chain security risks for…
❤1
🤖 DeepSeek is reportedly training a 2T-parameter model. And planning an 8T one.
For context: their current V3 sits at 671B. This would be a 3x jump just to get started, with 8T as the eventual target.
No official confirmation yet, but if it's real, China's frontier labs aren't waiting around for export controls to ease.
Source
For context: their current V3 sits at 671B. This would be a 3x jump just to get started, with 8T as the eventual target.
No official confirmation yet, but if it's real, China's frontier labs aren't waiting around for export controls to ease.
Source
X (formerly Twitter)
Wall St Engine (@wallstengine) on X
DeepSeek CEO Liang Wenfeng told investors that using more domestic chips for AI training is now a major priority, with Huawei expected to begin deliveries as early as Q4.
Note: DeepSeek is train…
Note: DeepSeek is train…
❤2
🤖 xAI just dropped Grok 4.7. New pretrain, 2.1T params.
Not a 4.6 refresh. That's the detail that matters here. Less than six weeks after 4.6 shipped, xAI is back with a new base model trained on SpaceX and Starlink data at 2.1 trillion parameters.
Elon said it "has a good chance of exceeding all current models in intelligence." Benchmarks pending.
(We've heard that one before, but the param jump is real.)
Not a 4.6 refresh. That's the detail that matters here. Less than six weeks after 4.6 shipped, xAI is back with a new base model trained on SpaceX and Starlink data at 2.1 trillion parameters.
Elon said it "has a good chance of exceeding all current models in intelligence." Benchmarks pending.
(We've heard that one before, but the param jump is real.)
x.ai
Introducing Grok 4.7
SpaceXAI's most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models.
❤1
🤖 Amazon kicked Meta's Muse agent off its site. No warning, no deal.
Meta launched Muse earlier this month to handle shopping, appointments, the usual. Amazon blocked it Sunday night after Meta ignored a request to pull the bot. Users now see a popup: "unauthorized AI agent."
Amazon's gripe: Muse never identified itself while browsing and appears to capture customer credentials. Meta didn't even tell them it was coming.
Two trillion-dollar companies. One didn't ask permission.
Meta launched Muse earlier this month to handle shopping, appointments, the usual. Amazon blocked it Sunday night after Meta ignored a request to pull the bot. Users now see a popup: "unauthorized AI agent."
Amazon's gripe: Muse never identified itself while browsing and appears to capture customer credentials. Meta didn't even tell them it was coming.
Two trillion-dollar companies. One didn't ask permission.
Bloomberg.com
Amazon Blocks Meta’s Muse AI Agent From Its Retail Site
Amazon.com Inc. has blocked Meta Platforms Inc.’s new artificial intelligence agent from its retail site after the social media company declined a request to remove the bot.
❤2
⚡️ US data centres are short 6 New York Cities' worth of electricity.
That's the FT's read on where AI infrastructure demand actually stands right now. Not a future projection. A current gap.
And it's not a solvable-by-Tuesday problem. Grid buildout takes years. Model training doesn't wait.
Source
That's the FT's read on where AI infrastructure demand actually stands right now. Not a future projection. A current gap.
And it's not a solvable-by-Tuesday problem. Grid buildout takes years. Model training doesn't wait.
Source
❤3