🚨 AI News | TestingCatalog
7.54K subscribers
4.29K photos
689 videos
40 files
4.31K links
Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors πŸ—ž
Download Telegram
Claude app for iOS now has an explicit warning next to the Max effort option that it consumes 1.5 more usage.

Max 1.5 ⚠️
5❀4πŸ‘1
Media is too big
VIEW IN TELEGRAM
PERPLEXITY πŸ”₯: A Hybrid mode for Perplexity Computer on Mac has been officially announced!

> The Hybrid mode is powered by PPLX Qwen 3.8 27B, a custom post-trained model from Perplexity.

> It also comes with a Privacy Gate feature to detect PII data before it is sent to the cloud.

> Perplexity also open-sourced the Privacy Gate classifier on Hugging Face.
❀5πŸ‘2πŸ‘€1πŸ†’1
🚨 AI News | TestingCatalog
Exclusive: Deeper look into Hatch Agent from Meta Meta’s unreleased Hatch materials point to a standalone agent platform with web, iOS, and Android apps, persistent project spaces, connectors, privacy controls, shareable agents, and browser or file-based…
META πŸ”₯: Project Hatch will be released under the name β€œMuse” and will arrive with a waitlist!

Hatch was an internal codename of the upcoming superapp from Meta. Read more about Hatch in the post above.

Joined πŸ‘€πŸ‘€πŸ‘€
8❀2πŸ‘€2
Media is too big
VIEW IN TELEGRAM
META πŸ”₯: A new Muse Voice Transcribe model from MSL is now available on Meta models API.

SOTA in streaming speech-to-text. Trained on 70+ languages.


As it has been foretold πŸ‘€
❀33
ANTHROPIC πŸ”₯: Claude Fable 5.1 is being prepared for the upcoming release!

> The "Thought Preserved: Modifying the way the Messages API handles thought blocks to protect against distillation" support page has been updated.

> Both Fable 5.1 and Mythos 5.1 are expected soon.

Soon? πŸ‘€
❀82😁1
BREAKING πŸ”₯: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1!

It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5.

On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.


Rolling out on Claude now πŸ‘€
❀10πŸ”₯43
🚨 AI News | TestingCatalog
BREAKING πŸ”₯: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1! It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. Rolling out on Claude now πŸ‘€
ANTHROPIC πŸ”₯: Fable 5.1 is now available on Claude and Claude Code.

It requires extra usage credits while priced the same as Fable 5, with 75% cheaper API cache reads.

> Writes in plain language and sticks to what you asked for.

> Creates finished spreadsheets and checks each number as it goes.

> Shows its sources and separates what's known from what's estimated.
πŸ‘7❀51
OPENAI πŸ”₯: Astra will be "available soon," but its cybersecurity capabilities will be limited.

> Astra scored 100% on ExploitBench.
> OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently.
> Astra achieved "much higher arbitrary code-execution rates than GPT‑5.6 Sol".
> During the evaluation, Astra found 2 new zero-day vulnerabilities and turned them into working exploit chains.

Soon πŸ‘€
❀11πŸ‘€2πŸ‘11
GOOGLE πŸ”₯: Gemini 3.8 Flash is set to arrive tomorrow, according to WSJ.

β€œJetski” has been mentioned in the article as a Google’s internal coding tool too.


Soon πŸ‘€
❀104πŸ‘1
Anthropic launches Claude Fable 5.1 and Mythos 5.1

Anthropic launched Claude Fable 5.1 for general use and Mythos 5.1 for vetted defenders and scientists, with lower cache-read costs, stronger benchmark results, customer-controlled data options, and access across major clouds.

πŸ—ž #anthropic @testingcatalog
❀4πŸ”₯2
DAILY AI BRIEF πŸ—ž β€” Sept 2

OPENAI πŸ”₯:
> Official β€œPath to Astra” post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is β€œcoming soon” β€” advanced cyber tools stay limited to testers / Daybreak Blue at first.
> M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha.

GOOGLE πŸ”₯:
> Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway.
> WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today.
> Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later.

ANTHROPIC πŸ”₯:
> Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads β€” about 25% cheaper typically, up to 45% on heavy agent runs.

META πŸ”₯:
> Muse Voice Transcribe is live β€” MSL’s first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available.

XAI πŸ”₯:
> Elon: β€œGrok 4.7 comes out in 10 days.” That’s ~Sept 12. Reply to Tobi on Grok 4.6.

ALIBABA πŸ”₯:
> Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 β€” 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max.

WORLD LABS πŸ”₯:
> Fei-Fei Li’s lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet.

* Too much is happening, and I also have some scoops planned for today.
** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.
❀13955
Muse superapp from Meta and Ava model with computer use

META πŸ”₯: A new model named Ava with computer-use capabilities is undergoing closed testing in the Meta AI desktop app.

> "Agentic assistant with computer use."

Users can also enable apps for computer use individually, directly from the window attachment menu.

> "Clicks, types and scolls only in this window."

Watermelon, is this you? πŸ‘€

πŸ—ž #meta @testingcatalog
❀44πŸ‘1
GOOGLE πŸ”₯: Gemini 3.8 Flash started appearing on Google Coud Console quotas page, a usual release predecessor.

Earlier today, users also spotted that Gemini 3.8 Flash has been powering some of there conversations on Gemini already.

Very soon πŸ‘€
❀82
GOOGLE πŸ”₯: Gemini 3.8 Flash is already available in Agent Studio on GCP.

Best for
- Complex multimodal data processing
- Coding use cases
- Supporting software engineering–related agentic tasks

Use case
- Processing data with images and text
- Coding problems
- Web research and application testing
17❀4πŸ”₯3πŸ‘Ž2