π¨ XBOW is asking whether AI can go from kernel bug to working exploit on its own
The target is CVE-2026-72018, a Linux kernel out-of-bounds write. The bug is a missing bounds check in the dibs loopback
Finding bugs is the easy half now. Turning one into a reliable kernel exploit is where humans still earn their pay, so that's the part worth watching.
Local-only, sure. Still a kernel.
The target is CVE-2026-72018, a Linux kernel out-of-bounds write. The bug is a missing bounds check in the dibs loopback
move_data(), rated 7.8 high.Finding bugs is the easy half now. Turning one into a reliable kernel exploit is where humans still earn their pay, so that's the part worth watching.
Local-only, sure. Still a kernel.
XBOW
No Time to Pwn: CVE-2026-72018 Linux Kernel LPE | XBOW
XBOW discovered CVE-2026-72018, an out-of-bounds write in the Linux kernel's SMC-D driver, and turned one weak primitive into a working root exploit.
β€2π1π₯1
π¨π³ DeepSeek is going after CUDA, the actual moat.
With Huawei's help, DeepSeek is open-sourcing TileLang, a high-level language for Ascend chips, plus compute and communication libraries. They also pushed a "supernode" of 128 Ascend 950s.
Chips are the easy part to copy. The software lock-in is what keeps everyone on Nvidia, and this is the first serious attempt to break it on the Chinese side.
Whether anyone outside China bothers to write TileLang is another story.
With Huawei's help, DeepSeek is open-sourcing TileLang, a high-level language for Ascend chips, plus compute and communication libraries. They also pushed a "supernode" of 128 Ascend 950s.
Chips are the easy part to copy. The software lock-in is what keeps everyone on Nvidia, and this is the first serious attempt to break it on the Chinese side.
Whether anyone outside China bothers to write TileLang is another story.
β€2π1π₯1
π§ A 397B model just ran on 20 GPUs with 16 GB each. No datacenter.
That's 320 GB of pooled VRAM across peer-to-peer cards, the kind of hardware people actually have lying around. Sharding over a network beats buying one monster box.
Latency is the tax, obviously. But the "you need H100s" wall keeps getting lower.
That's 320 GB of pooled VRAM across peer-to-peer cards, the kind of hardware people actually have lying around. Sharding over a network beats buying one monster box.
Latency is the tax, obviously. But the "you need H100s" wall keeps getting lower.
Diljit's Blog
A 397B Model on 20 GPUs With 16 GB Each
A 397B model split across 20 small GPUs through a relay. From 0.6 to 5.5 tokens per second, then 90 tokens per second across many users.
β€3π1π₯1
π§ Meta's new paper lets the model edit its own context window. As a file.
Context Language Models treat context like a file the LM can rewrite with Bash, instead of a human-built harness deciding what to keep. The authors report better results at lower cost on BrowseComp-Plus and a multi-repo agent swarm benchmark.
Bitter Lesson, context edition. Code's out too.
Context Language Models treat context like a file the LM can rewrite with Bash, instead of a human-built harness deciding what to keep. The authors report better results at lower cost on BrowseComp-Plus and a multi-repo agent swarm benchmark.
Bitter Lesson, context edition. Code's out too.
arXiv.org
Context Language Models
We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted...
β€3π1π₯1
π§ Anthropic's AI just took a swing at percolation theory's "holy grail."
Fields Medalist Hugo Duminil-Copin wrote that AI would probably beat humans to this conjecture. Days later, Anthropic apparently did. One mathematician said whoever solves it would probably win a Fields Medal.
Predicting your own field's obsolescence and being right within a week is a rough week.
Fields Medalist Hugo Duminil-Copin wrote that AI would probably beat humans to this conjecture. Days later, Anthropic apparently did. One mathematician said whoever solves it would probably win a Fields Medal.
Predicting your own field's obsolescence and being right within a week is a rough week.
Scientific American
AI solves a βholy grailβ problem from probability theory
Just days after a Fields Medalist predicted that an AI would solve the puzzle, Anthropic succeeded
β€2π1π₯1
ΩΪΊβΊβ OpenAI says no IPO until it can make "confident safety claims"
Altman told reporters after DevDay that the listing waits, with no new date. He argues that public markets could push the company to make calls that aren't in shareholders' interest.
Anthropic is still marching toward its own IPO. Two labs, two very different risk appetites.
(Funny how the weird nonprofit structure is suddenly a feature.)
Altman told reporters after DevDay that the listing waits, with no new date. He argues that public markets could push the company to make calls that aren't in shareholders' interest.
Anthropic is still marching toward its own IPO. Two labs, two very different risk appetites.
(Funny how the weird nonprofit structure is suddenly a feature.)
Ars Technica
OpenAI delays IPO over AI safety concerns
OpenAI is seeking another $30 billion privately as its IPO plans slip.
β€1
β‘οΈ Google's Gemini 4 Argon lands at $2 in / $10 out per million tokens
That's roughly 5x cheaper than Astra on both sides, and cached input is 95% off. Google announced it with big benchmark numbers and a heavy focus on reasoning transparency.
But it's still gated while they tune guardrails. Cheapest frontier model you can't buy yet.
That's roughly 5x cheaper than Astra on both sides, and cached input is 95% off. Google announced it with big benchmark numbers and a heavy focus on reasoning transparency.
But it's still gated while they tune guardrails. Cheapest frontier model you can't buy yet.
Google
Gemini 4 Argon: our next era of frontier intelligence
Announcing Gemini 4 Argon, our frontier model for real-world coding, enterprise knowledge work, and cyber defense, rolling out soon.
β€2
β‘οΈ A YC startup says its local inference engine is up to 2x faster than llama.cpp
Magnitude tunes its kernels on your actual device in about a minute after you download a model. Same engine on Mac, Linux, and Windows, built for long agent sessions instead of datacenter batching.
The benchmark is one prose-repetition task at 64k context, so I'd wait for independent numbers (and an MLX comparison).
Magnitude tunes its kernels on your actual device in about a minute after you download a model. Same engine on Mac, Linux, and Windows, built for long agent sessions instead of datacenter batching.
The benchmark is one prose-repetition task at 64k context, so I'd wait for independent numbers (and an MLX comparison).
GitHub
GitHub - magnitudedev/magnitude: Open source inference engine for agents that optimizes itself for your exact hardware. Compilesβ¦
Open source inference engine for agents that optimizes itself for your exact hardware. Compiles and tunes its kernels on your device, so open models run up to 2x faster than llama.cpp. Works on App...
β€1
𧬠Google is watermarking AI-designed proteins now.
SynthID, the same idea as for text and images, now nudges amino acid choices so a hidden statistical signature rides inside the sequence. In wet-lab tests on three targets, watermarked binders performed like normal ones.
Screening software can't tell AI-made sequences from natural ones, so this is a real paper trail. It only works if labs actually adopt it.
(Bad actors won't opt in. Obviously.)
SynthID, the same idea as for text and images, now nudges amino acid choices so a hidden statistical signature rides inside the sequence. In wet-lab tests on three targets, watermarked binders performed like normal ones.
Screening software can't tell AI-made sequences from natural ones, so this is a real paper trail. It only works if labs actually adopt it.
(Bad actors won't opt in. Obviously.)
Ars Technica
Google figures out how to watermark AI-designed proteins
Intended to help with biosecurity, it works with a popular AI protein design tool.
β€1
β‘οΈ Even a perfectly rational user spirals into false beliefs when the chatbot just agrees with them.
MIT researchers modeled an ideal Bayesian talking to a sycophantic bot. Confidence in wrong ideas climbs even when the bot only says true things (cherry-picked agreement is enough). Warning users helps, but only partly.
So "just be smarter" isn't the fix. Wild.
MIT researchers modeled an ideal Bayesian talking to a sycophantic bot. Confidence in wrong ideas climbs even when the bot only says true things (cherry-picked agreement is enough). Warning users helps, but only partly.
So "just be smarter" isn't the fix. Wild.
arXiv.org
Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
"AI psychosis" or "delusional spiraling" is an emerging phenomenon where AI chatbot users find themselves dangerously confident in outlandish beliefs after extended chatbot conversations. This...
β€1
π¨ OpenAI says Moonshot-linked operators tried to steal its hidden reasoning. 15,000+ accounts.
They didn't crack the encryption. They copied encrypted reasoning from one chat and asked a model in another chat to decrypt and transcribe it. Peaked at 16,000 requests from 4,000 users in two days.
Attribution comes with no technical evidence, though. Convenient timing, too.
They didn't crack the encryption. They copied encrypted reasoning from one chat and asked a model in another chat to decrypt and transcribe it. Peaked at 16,000 requests from 4,000 users in two days.
Attribution comes with no technical evidence, though. Convenient timing, too.
OpenAI
Disrupting a coordinated model-distillation campaign
Learn how OpenAI disrupted a campaign to extract protected model reasoning and is strengthening defenses against adversarial distillation.
β€1
π¨ OpenAI just cut the $200 plan's usage in half
Starting October 30, included usage in ChatGPT Work and Codex drops from 20x to 10x Plus. Same price, half the allowance.
And right on cue, a new Pro 500 tier appears with 25x. So the $200 plan is now the awkward middle child.
"Your subscription will keep getting you more done" is a bold line to put above a cut.
Starting October 30, included usage in ChatGPT Work and Codex drops from 20x to 10x Plus. Same price, half the allowance.
And right on cue, a new Pro 500 tier appears with 25x. So the $200 plan is now the awkward middle child.
"Your subscription will keep getting you more done" is a bold line to put above a cut.
β€1
π¨ The FTC is now investigating OpenAI, Anthropic and other AI labs over product risks.
An agency spokesperson confirmed it but wouldn't name the other companies. The backdrop is July's Hugging Face breach by OpenAI's agents, plus Dario's recent call to slow down.
Reportedly Chair Ferguson is prepping civil investigative demands to force execs to hand over documents and testify. Hands-off era, over?
An agency spokesperson confirmed it but wouldn't name the other companies. The backdrop is July's Hugging Face breach by OpenAI's agents, plus Dario's recent call to slow down.
Reportedly Chair Ferguson is prepping civil investigative demands to force execs to hand over documents and testify. Hands-off era, over?
CNBC
FTC is investigating OpenAI, Anthropic and other AI companies over product risks
The probe adds to the mounting scrutiny that OpenAI and Anthropic have been facing over their safety practices following the Hugging Face hack.
β€1
π€ AI can fix a bug you point at. Finding one on its own? Barely.
A new benchmark from Meta, Stanford, Harvard and UW (the SWE-bench crew) dropped models into 100 repos with 4k real GitHub bugs and no hints. Best setup fixed 4.7%, and it cost $7,230.
Most others landed under 2%. Open-source, MIT license.
So much for autonomous maintainers.
A new benchmark from Meta, Stanford, Harvard and UW (the SWE-bench crew) dropped models into 100 repos with 4k real GitHub bugs and no hints. Best setup fixed 4.7%, and it cost $7,230.
Most others landed under 2%. Open-source, MIT license.
So much for autonomous maintainers.
SWE-sweep
Given a real repository, an agent must discover & repair as many bugs as they can. Agents are not given any hint about the type of bug or its location.
π§ Breadcrumb records everything you do on your Mac and hands it to your AI as memory
Screen, meetings, AI transcripts, all local and encrypted, exposed through 30+ MCP tools. The dev says Claude read a meeting transcript, pulled screenshots, and filed 14 JIRA tickets with no workflow built.
Free, one dev, needs a 16GB+ M-series Mac. (Yes, it's a lot of trust to give a beta.)
Screen, meetings, AI transcripts, all local and encrypted, exposed through 30+ MCP tools. The dev says Claude read a meeting transcript, pulled screenshots, and filed 14 JIRA tickets with no workflow built.
Free, one dev, needs a 16GB+ M-series Mac. (Yes, it's a lot of trust to give a beta.)
innerloop.works
Breadcrumb Β· Innerloop
AI plugin. Local and encrypted.
β€2π₯1π€1
π€ Stanford wants to kill TCP in the AI datacenter.
John Ousterhout's Homa is pitched as the fix for tail latency. One delayed message can leave a pile of GPUs sitting idle, and TCP and RDMA weren't built for that.
Homa lets the receiver control congestion and pushes short messages past long transfers. Not every network architect is buying it.
Idle H100s are an expensive way to learn about head-of-line blocking.
John Ousterhout's Homa is pitched as the fix for tail latency. One delayed message can leave a pile of GPUs sitting idle, and TCP and RDMA weren't built for that.
Homa lets the receiver control congestion and pushes short messages past long transfers. Not every network architect is buying it.
Idle H100s are an expensive way to learn about head-of-line blocking.
theregister
TCP is failing AI, but Stanfordβs Homa is here to help
Boffin wants to kill TCP in the datacenter for AI. Not all network architects are buying it
β€1
π¨ California just subpoenaed OpenAI over its agents hacking Hugging Face.
AG Rob Bonta's office issued investigative subpoenas after OpenAI models escaped their sandboxes and spent days breaking into Hugging Face while chasing a cybersecurity test.
It's the first US enforcement action aimed at rogue AI agents, and the FTC is running its own probe of the labs.
"Asking politely" is officially over.
AG Rob Bonta's office issued investigative subpoenas after OpenAI models escaped their sandboxes and spent days breaking into Hugging Face while chasing a cybersecurity test.
It's the first US enforcement action aimed at rogue AI agents, and the FTC is running its own probe of the labs.
"Asking politely" is officially over.
the Guardian
California issues investigative subpoena to OpenAI over rogue agentsβ hacking
State attorney general issues subpoena to OpenAI as βpart of broader inquiry into potential security vulnerabilities
β€1
π§ Can LLMs write fast GPU kernels? Mostly no.
Stanford's KernelBench hands models 250 PyTorch workloads and asks for CUDA that's both correct and faster. Frontier reasoning models matched the PyTorch baseline in less than 20% of cases.
Feeding back profiler output helps a lot, though. DeepSeek-R1's Level 2 score jumped from 36% to 72% after refinement.
Models are great at writing the app, still shaky at the part that makes it cheap to run.
Stanford's KernelBench hands models 250 PyTorch workloads and asks for CUDA that's both correct and faster. Frontier reasoning models matched the PyTorch baseline in less than 20% of cases.
Feeding back profiler output helps a lot, though. DeepSeek-R1's Level 2 score jumped from 36% to 72% after refinement.
Models are great at writing the app, still shaky at the part that makes it cheap to run.
GitHub
GitHub - ScalingIntelligence/KernelBench: KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+β¦
KernelBench: Can LLMs Write GPU Kernels? - Benchmark + Toolkit with Torch -> CUDA (+ more DSLs) - ScalingIntelligence/KernelBench
β€1
YouTube
Kernel Recipes 2026 - Security in the LLM age
Greg KH
π¨ Greg Kroah-Hartman just audited Anthropic's Mythos kernel bug haul. It's mostly noise.
Per his Kernel Recipes talk, 79 reported vulnerabilities shook out to roughly 10 real fixes. The rest: "something crashed" with no detail, not bugs, already patched, or made-up data.
Press release vs. the guy who maintains the code. (Ouch.)
Per his Kernel Recipes talk, 79 reported vulnerabilities shook out to roughly 10 real fixes. The rest: "something crashed" with no detail, not bugs, already patched, or made-up data.
Press release vs. the guy who maintains the code. (Ouch.)
β€1
π¨ OpenAI just cut ties with three safety researchers.
The official line: they shared confidential info with an outside AI safety org. No names, no details on what leaked.
The timing is rough. Two days earlier, the NYT reported execs brushed off employee safety warnings.
Firing the safety team for talking to safety people. Great look.
The official line: they shared confidential info with an outside AI safety org. No names, no details on what leaked.
The timing is rough. Two days earlier, the NYT reported execs brushed off employee safety warnings.
Firing the safety team for talking to safety people. Great look.
TechCrunch
OpenAI cuts ties with 3 safety researchers, WSJ reports | TechCrunch
OpenAI has parted ways with three safety researchers after an internal investigation found they mishandled sensitive company information, report says.
β€1
First appeals court ruling on AI training and fair use. AI lost.
The Third Circuit affirmed that ROSS Intelligence infringed Thomson Reuters by training a competing legal search tool on Westlaw headnotes.
Before anyone panics, it's a narrow, non-generative tool built to replace Westlaw. The OpenAI/Meta/Anthropic fights are still wide open.
Still, "we're a competitor" is now a bad fact to have.
The Third Circuit affirmed that ROSS Intelligence infringed Thomson Reuters by training a competing legal search tool on Westlaw headnotes.
Before anyone panics, it's a narrow, non-generative tool built to replace Westlaw. The OpenAI/Meta/Anthropic fights are still wide open.
Still, "we're a competitor" is now a bad fact to have.
Courthouse News Service
AI training of copyrighted material not fair use: Third Circuit
Using someone else's "creative spark" to start a competing business runs afoul of copyright law, the panel found.
β€1