๐Ÿšจ AI News | TestingCatalog
7.54K subscribers
4.29K photos
689 videos
40 files
4.31K links
Latest AI News on AI Agents, Model Releases, Tools, Leaks, and Rumors ๐Ÿ—ž
Download Telegram
ANTHROPIC ๐Ÿ”ฅ: Claude Fable 5.1 is being prepared for the upcoming release!

> The "Thought Preserved: Modifying the way the Messages API handles thought blocks to protect against distillation" support page has been updated.

> Both Fable 5.1 and Mythos 5.1 are expected soon.

Soon? ๐Ÿ‘€
โค82๐Ÿ˜1
BREAKING ๐Ÿ”ฅ: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1!

It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5.

On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.


Rolling out on Claude now ๐Ÿ‘€
โค10๐Ÿ”ฅ43
๐Ÿšจ AI News | TestingCatalog
BREAKING ๐Ÿ”ฅ: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1! It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. Rolling out on Claude now ๐Ÿ‘€
ANTHROPIC ๐Ÿ”ฅ: Fable 5.1 is now available on Claude and Claude Code.

It requires extra usage credits while priced the same as Fable 5, with 75% cheaper API cache reads.

> Writes in plain language and sticks to what you asked for.

> Creates finished spreadsheets and checks each number as it goes.

> Shows its sources and separates what's known from what's estimated.
๐Ÿ‘7โค51
OPENAI ๐Ÿ”ฅ: Astra will be "available soon," but its cybersecurity capabilities will be limited.

> Astra scored 100% on ExploitBench.
> OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently.
> Astra achieved "much higher arbitrary code-execution rates than GPTโ€‘5.6 Sol".
> During the evaluation, Astra found 2 new zero-day vulnerabilities and turned them into working exploit chains.

Soon ๐Ÿ‘€
โค11๐Ÿ‘€2๐Ÿ‘11
GOOGLE ๐Ÿ”ฅ: Gemini 3.8 Flash is set to arrive tomorrow, according to WSJ.

โ€œJetskiโ€ has been mentioned in the article as a Googleโ€™s internal coding tool too.


Soon ๐Ÿ‘€
โค104๐Ÿ‘1
Anthropic launches Claude Fable 5.1 and Mythos 5.1

Anthropic launched Claude Fable 5.1 for general use and Mythos 5.1 for vetted defenders and scientists, with lower cache-read costs, stronger benchmark results, customer-controlled data options, and access across major clouds.

๐Ÿ—ž #anthropic @testingcatalog
โค4๐Ÿ”ฅ2
DAILY AI BRIEF ๐Ÿ—ž โ€” Sept 2

OPENAI ๐Ÿ”ฅ:
> Official โ€œPath to Astraโ€ post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is โ€œcoming soonโ€ โ€” advanced cyber tools stay limited to testers / Daybreak Blue at first.
> M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha.

GOOGLE ๐Ÿ”ฅ:
> Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway.
> WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today.
> Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later.

ANTHROPIC ๐Ÿ”ฅ:
> Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads โ€” about 25% cheaper typically, up to 45% on heavy agent runs.

META ๐Ÿ”ฅ:
> Muse Voice Transcribe is live โ€” MSLโ€™s first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available.

XAI ๐Ÿ”ฅ:
> Elon: โ€œGrok 4.7 comes out in 10 days.โ€ Thatโ€™s ~Sept 12. Reply to Tobi on Grok 4.6.

ALIBABA ๐Ÿ”ฅ:
> Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 โ€” 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max.

WORLD LABS ๐Ÿ”ฅ:
> Fei-Fei Liโ€™s lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet.

* Too much is happening, and I also have some scoops planned for today.
** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.
โค13955
Muse superapp from Meta and Ava model with computer use

META ๐Ÿ”ฅ: A new model named Ava with computer-use capabilities is undergoing closed testing in the Meta AI desktop app.

> "Agentic assistant with computer use."

Users can also enable apps for computer use individually, directly from the window attachment menu.

> "Clicks, types and scolls only in this window."

Watermelon, is this you? ๐Ÿ‘€

๐Ÿ—ž #meta @testingcatalog
โค44๐Ÿ‘1
GOOGLE ๐Ÿ”ฅ: Gemini 3.8 Flash started appearing on Google Coud Console quotas page, a usual release predecessor.

Earlier today, users also spotted that Gemini 3.8 Flash has been powering some of there conversations on Gemini already.

Very soon ๐Ÿ‘€
โค82
GOOGLE ๐Ÿ”ฅ: Gemini 3.8 Flash is already available in Agent Studio on GCP.

Best for
- Complex multimodal data processing
- Coding use cases
- Supporting software engineeringโ€“related agentic tasks

Use case
- Processing data with images and text
- Coding problems
- Web research and application testing
17โค4๐Ÿ”ฅ3๐Ÿ‘Ž2
GOOGLE ๐Ÿ”ฅ: Gemini 3.8 Flash is rolling out on Gemini, Google AI Studio and APIs.

Gemini 3.8 Flash scores 71% on DeepSWE 1.1, compared to 74% for Claude Opus 5, at a much lower price.

> Input price
$0.75 through December 31, 2026.
$1.50 starting January 1, 2027.

> Output price (including thinking tokens)
$3.75 through December 31, 2026.
$7.50 starting January 1, 2027.

This is big ๐Ÿ‘€
๐Ÿ”ฅ85โค4
This media is not supported in your browser
VIEW IN TELEGRAM
SPACEXAI ๐Ÿ”ฅ: Grok Bot is now available on Android platform!

Bot testing time ๐Ÿ‘€
โค5๐Ÿ‘41
OPENAI ๐Ÿ”ฅ: GPT-6-Astra model slug has been spotted on the APIs.

If we will actually get it tomorrow, it would be a huge week.

Routing first ๐Ÿ‘€
โค1610๐Ÿ‘1
META ๐Ÿ”ฅ: Muse Spark 1.3 has been officially announced, and it scored above GPT-5.6 and Opus 5 on DeepSWE 1.1!

> Muse Spark 1.3 is rolling out on Meta model APIs.

> "Watermelon", the next big model upgrade from Meta, and the Muse Spark open-weight version are coming soon!

The competition is getting hotter ๐Ÿ‘€
12โค6๐Ÿ˜ด21