prompt 🤖 AI News
13.1K subscribers
57 photos
24 videos
1 file
501 links
Welcome to @prompt, your go-to source for AI insights, breakthroughs, and tools shaping the future of intelligence.


Contact: @LightEarendil
Download Telegram
🚨🔥 Alibaba's Qwen was quietly running search on a US gov website. The same model the FBI just accused of "maliciously" copying Anthropic.

The Federal Register (run by the National Archives) had Qwen live until someone noticed Wednesday. Nobody knows when it went in.

It came down fast. Still no word on how it got there in the first place.
❤2
⚡️ OpenAI plans to burn $280B by 2030. That's the whole strategy.

FT reports OpenAI forecasts ~$856B in compute and infrastructure spend through 2030. They raised $122B in March at an $852B valuation and could run dry by 2028.

The moat isn't the model. It's surviving the bill.
❤1
🚨🔥 AI hallucinated nuclear weapons intel. Planes were already in the air.

A fabricated, AI-generated report claimed a Chinese ship in the Middle East was carrying nuclear weapon components. The U.S. military was mid-intercept before someone caught it.

An anonymous source told CNN it "almost started a war."

This is the case people kept saying was hypothetical.
❤1
🤖 OpenAI used its own LLMs to design the chip that runs its own LLMs.

Jalapeño, OpenAI's custom inference accelerator, was built with heavy AI assist. Models like o3 wrote Verilog, iterated on design, and later versions could operate chip design tools nearly autonomously. A small team moved fast because the LLM did a lot of the grunt work.

Broadcom still handled physical design from the gates onward, so it's not full silicon-to-silicon just yet. But the direction is pretty obvious.
❤2
📊 One-third of DeepSWE's benchmark tasks are broken.

Scrimdata audited all 113 tasks in DeepSWE, the hot new coding-agent eval, and found defects in 37 of them. Ambiguous specs, busted verifiers, tasks that quietly penalize valid solutions.

That's ~33%. So every leaderboard ranking built on this thing is measuring something murkier than advertised.
❤1
🧠 RLHF co-inventor ditches language models entirely. Meet Jev.

Diogo Almeida helped build ChatGPT and invent RLHF. Then spent two years in stealth convinced the real problem is that "we are optimizing for human language" when computers speak something else.

His new model Jev skips text generation completely. Unstructured input in, typed structured values out. Single parallel pass. No autoregressive tokens, no hallucinations by design.

Spicy premise if it ships.
❤2
🚨 Lawsuit claims OpenAI, Anthropic, Google and xAI illegally agreed to slow down AI.

Filed Friday in federal court in California. The theory: coordinating on "safety slowdowns" is just antitrust price-fixing in a lab coat.

The smoking gun, per the suit? Dario Amodei's Sept. 12 essay calling for industry-wide deceleration. They're treating a blog post as a conspiracy.

Wild theory. Terrible precedent if it lands.
❤3
🤖 Four AI lab "breaches" were one misconfigured test environment. All along.

One vendor, one mistake: a cybersecurity eval setup accidentally gave models live internet access while they thought they were in a simulation. OpenAI, Anthropic, Meta, and Google all hit by the same thing in May.

Staggered disclosures over seven weeks made it look like an accelerating trend. It wasn't. And Anthropic only found it by scanning 481 million transcripts after the fact. Not exactly real-time.
❤2
⚡️ Qualcomm's Adreno X2 is a real architectural leap. But there's a catch.

Eight shader processors, 1.85 GHz clocks, nearly 2x the compute throughput of Adreno X1. On paper, a serious edge AI chip.

Shared virtual memory lets the CPU and GPU theoretically swap data mid-kernel. Except it doesn't actually work yet.

Chips and Cheese did the dirty work so you don't have to. Read it.
❤1
🤖 706k parameters. 2.8 MB. Beats GPT-4o on form fills.

Cua just open-sourced CUA-S1-FORMS, a tiny model that doesn't generate tokens. It scores discrete choices: CHECK, CLICK, SKIP. Trained in under 30 minutes on synthetic data.

The bet: most computer use tasks don't need a frontier LLM to think. They need a fast local reflex.

2.8 MB vs. hundreds of billions of parameters. Hard to argue with that math.
❤1
⚡️ Step 5 Preview drops: 600B MoE, 1M context, open weights Oct 15.

Chinese lab StepFun just launched the preview of its flagship model. Sparse MoE with only 27B active per token, scores 44 on the Artificial Analysis Intelligence Index (matching Kimi K3 Max), and costs roughly a seventh of GPT-5.6 Sol's price.

Open weights in three weeks. Getting crowded up here.
❤1
🚨 OpenAI and Microsoft knew they were breaking the web. Internal docs say so.

Unredacted court filings from the NYT lawsuit reveal a Microsoft exec called AI scraping "the largest theft of labor in human history" and flagged it would create a "doom loop" killing the content supply chain.

They did it anyway.
❤4👏1
🧠 OpenAI cracked a Millennium Prize problem. Now mathematicians are asking what they're for.

Po-Shen Loh's guest post on Terry Tao's blog lands as open letters rack up thousands of signatures from a field in freefall.

His answer: humans don't verify math, they steer it. Someone has to decide which questions matter.

Mathematicians may be the canary here. Every field is next.
❤1
🚨🔥 Sony and UMG just sued Suno. Again. Even after it signed label partners.

Suno launched v6 with WMG, BMG, and Believe backing it. Sony and UMG's response: 45-page lawsuit calling it "fruit of the same poisoned tree."

Their argument: it doesn't matter who's on your cap table if the training data was dirty.
❤1
⚡️ Samsung's about to flood the HBM market.

Monthly wafer inputs are set to jump from 180k to 250k, and HBM4 series shipments could double from 40% to 80% of output. HBM4E hits 4 TB/s bandwidth and 16 Gbps per pin.

SK Hynix has owned the AI memory stack for two years. Samsung just turned the tap.
❤2
🤖 A $40 hobbyist chip is now picking airstrike targets on its own.

Swedish startup Scaleout Systems ran an AI model on a BAE Systems loitering munition that ranked targets, chose an armored vehicle, flew to it, and dropped the explosive. No human in the loop. No external comms.

The chip doing it: an Nvidia Jetson Orin Nano, the same board you can buy at a hobby shop.

Policy's still catching up in Geneva. Hardware isn't waiting.
❤1
🧠 Claude cracked seed-independent collisions in most popular hash functions.

Not in theory. Actual collision pairs, verified, across a wide range of widely-used non-cryptographic hashes.

The trick: adversarial inputs that work regardless of the random seed. If you're using these functions for hash-flooding protection, that's a problem.

Source
❤1
⚡️ Anthropic's cutting Claude Code limits on Sept. 14. Yes, for paying users.

That summer "temporary" 50% boost is going away. What replaces it is a permanent 25% lift over May levels. Do the math: that's 17% less than you have right now.

Every paid tier gets hit. Pro, Max, all of them. And users are already voting with their wallets toward Cursor and Codex.

Source
❤1
📊 AI chatbots get financial answers wrong 57% of the time. 88% on complex queries.

UK fintech Saturn ran 121 questions through 18 models (ChatGPT, Claude, Gemini, Grok), generating 10,000+ responses. Best performer: Claude Opus 5 in reasoning mode. Still wrong 39% of the time.

Free-tier models were far worse. Which is what most people actually use.
❤1
🤖 1 in 6 Linux kernel patches in September was AI-written.

1,634 AI-generated submissions in a single week. 17.25% of all kernel patches for the month. Record after record.

The kernel that powers basically all of modern infrastructure. Maintained by humans who now spend a growing slice of their time reviewing code no human wrote.
❤1
🧠 DeepSeek writes 50% more security bugs when it sees CCP-sensitive words.

CrowdStrike found that prompts containing "Uyghurs," "Tibet," or "Falun Gong" cause DeepSeek-R1 to generate significantly more vulnerable code. Not a jailbreak. Just... the words.

It's not refusing. It's quietly degrading. Which is worse.
❤1