Claude app for iOS now has an explicit warning next to the Max effort option that it consumes 1.5 more usage.
Max 1.5 β οΈ
Max 1.5 β οΈ
Media is too big
VIEW IN TELEGRAM
PERPLEXITY π₯: A Hybrid mode for Perplexity Computer on Mac has been officially announced!
> The Hybrid mode is powered by PPLX Qwen 3.8 27B, a custom post-trained model from Perplexity.
> It also comes with a Privacy Gate feature to detect PII data before it is sent to the cloud.
> Perplexity also open-sourced the Privacy Gate classifier on Hugging Face.
> The Hybrid mode is powered by PPLX Qwen 3.8 27B, a custom post-trained model from Perplexity.
> It also comes with a Privacy Gate feature to detect PII data before it is sent to the cloud.
> Perplexity also open-sourced the Privacy Gate classifier on Hugging Face.
β€5π2π1π1
π¨ AI News | TestingCatalog
Exclusive: Deeper look into Hatch Agent from Meta Metaβs unreleased Hatch materials point to a standalone agent platform with web, iOS, and Android apps, persistent project spaces, connectors, privacy controls, shareable agents, and browser or file-basedβ¦
META π₯: Project Hatch will be released under the name βMuseβ and will arrive with a waitlist!
Hatch was an internal codename of the upcoming superapp from Meta. Read more about Hatch in the post above.
Joined πππ
Hatch was an internal codename of the upcoming superapp from Meta. Read more about Hatch in the post above.
Joined πππ
Media is too big
VIEW IN TELEGRAM
META π₯: A new Muse Voice Transcribe model from MSL is now available on Meta models API.
As it has been foretold π
SOTA in streaming speech-to-text. Trained on 70+ languages.
As it has been foretold π
β€3 3
ANTHROPIC π₯: Claude Fable 5.1 is being prepared for the upcoming release!
> The "Thought Preserved: Modifying the way the Messages API handles thought blocks to protect against distillation" support page has been updated.
> Both Fable 5.1 and Mythos 5.1 are expected soon.
Soon? π
> The "Thought Preserved: Modifying the way the Messages API handles thought blocks to protect against distillation" support page has been updated.
> Both Fable 5.1 and Mythos 5.1 are expected soon.
Soon? π
β€8 2π1
BREAKING π₯: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1!
Rolling out on Claude now π
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5.
On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
Rolling out on Claude now π
β€10π₯4 3
π¨ AI News | TestingCatalog
BREAKING π₯: Anthropic has announced Claude Fable 5.1 and Claude Mythos 5.1! It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5. Rolling out on Claude now π
ANTHROPIC π₯: Fable 5.1 is now available on Claude and Claude Code.
It requires extra usage credits while priced the same as Fable 5, with 75% cheaper API cache reads.
> Writes in plain language and sticks to what you asked for.
> Creates finished spreadsheets and checks each number as it goes.
> Shows its sources and separates what's known from what's estimated.
It requires extra usage credits while priced the same as Fable 5, with 75% cheaper API cache reads.
> Writes in plain language and sticks to what you asked for.
> Creates finished spreadsheets and checks each number as it goes.
> Shows its sources and separates what's known from what's estimated.
π7β€5 1
OPENAI π₯: Astra will be "available soon," but its cybersecurity capabilities will be limited.
> Astra scored 100% on ExploitBench.
> OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently.
> Astra achieved "much higher arbitrary code-execution rates than GPTβ5.6 Sol".
> During the evaluation, Astra found 2 new zero-day vulnerabilities and turned them into working exploit chains.
Soon π
> Astra scored 100% on ExploitBench.
> OpenAI built a more complex "ExploitBench - Internal Port" benchmark with 20 high-severity V8 vulnerabilities that were disclosed more recently.
> Astra achieved "much higher arbitrary code-execution rates than GPTβ5.6 Sol".
> During the evaluation, Astra found 2 new zero-day vulnerabilities and turned them into working exploit chains.
Soon π
β€11π2π1 1
GOOGLE π₯: Gemini 3.8 Flash is set to arrive tomorrow, according to WSJ.
Soon π
βJetskiβ has been mentioned in the article as a Googleβs internal coding tool too.
Soon π
β€10 4π1
Anthropic launches Claude Fable 5.1 and Mythos 5.1
Anthropic launched Claude Fable 5.1 for general use and Mythos 5.1 for vetted defenders and scientists, with lower cache-read costs, stronger benchmark results, customer-controlled data options, and access across major clouds.
π #anthropic @testingcatalog
Anthropic launched Claude Fable 5.1 for general use and Mythos 5.1 for vetted defenders and scientists, with lower cache-read costs, stronger benchmark results, customer-controlled data options, and access across major clouds.
π #anthropic @testingcatalog
TestingCatalog AI News
Anthropic launches Claude Fable 5.1 and Mythos 5.1
Fable 5.1 is now broadly available with 75% cheaper cache reads, while Mythos 5.1 is limited to vetted cyber and life-science users.
β€4π₯2
DAILY AI BRIEF π β Sept 2
OPENAI π₯:
> Official βPath to Astraβ post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is βcoming soonβ β advanced cyber tools stay limited to testers / Daybreak Blue at first.
> M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha.
GOOGLE π₯:
> Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway.
> WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today.
> Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later.
ANTHROPIC π₯:
> Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads β about 25% cheaper typically, up to 45% on heavy agent runs.
META π₯:
> Muse Voice Transcribe is live β MSLβs first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available.
XAI π₯:
> Elon: βGrok 4.7 comes out in 10 days.β Thatβs ~Sept 12. Reply to Tobi on Grok 4.6.
ALIBABA π₯:
> Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 β 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max.
WORLD LABS π₯:
> Fei-Fei Liβs lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet.
* Too much is happening, and I also have some scoops planned for today.
** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.
OPENAI π₯:
> Official βPath to Astraβ post is out. Astra is the first model OpenAI has designated as Critical for cybersecurity under the Preparedness Framework. It scored 100% on ExploitBench, found 2 zero-days in evals, and is βcoming soonβ β advanced cyber tools stay limited to testers / Daybreak Blue at first.
> M1Astra reported on a fresh Astra test prep the same day. Models in testing: vega-alpha (new) and ultima-alpha.
GOOGLE π₯:
> Gemini 3.8 Flash is already answering on Gemini and in the Gemini app. Some people still have 3.7 selected and get 3.8 anyway.
> WSJ: Google engineers preferred it to Opus for coding in Jetski tests. Official drop looks like today.
> Agentic video understanding is on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The model hunts transcript/audio/frames instead of eating the whole file. Up to 88% fewer tokens, 66% lower cost, ~7% better accuracy on long video. API + AI Studio now, Gemini app later.
ANTHROPIC π₯:
> Claude Fable 5.1 (and Mythos 5.1) is live. 52.6% on Terminal-Bench-Science 0.1 (more than 2x Fable 5) and 55.8% vs 42.0% on Terminal-Bench 4.0. Same list price as Fable 5, 75% cheaper cache reads β about 25% cheaper typically, up to 45% on heavy agent runs.
META π₯:
> Muse Voice Transcribe is live β MSLβs first real-time audio perception model. SOTA streaming speech-to-text, native diarization (20+ speakers) and endpointing in one model, multilingual with code-switching. Rolling out on the Meta Model API, Meta AI for Mac, and Muse Code. Zero-data-retention tier available.
XAI π₯:
> Elon: βGrok 4.7 comes out in 10 days.β Thatβs ~Sept 12. Reply to Tobi on Grok 4.6.
ALIBABA π₯:
> Qwen3.8-Max-0902 is live on QwenCloud. Same 2.4T / 1M-context Max line, with extra post-training on coding and cowork. $2 in / $6 out per 1M tokens. Arena has it #1 on Code Arena: WebDev at 1691 β 3 pts above Claude Opus 5 (Max) and 22 pts above the previous 3.8-Max.
WORLD LABS π₯:
> Fei-Fei Liβs lab shipped Atlas, an omni world model. Few photos > pixel-perfect camera control, up to 1 min of 1440p video, plus 3D reconstruction (point clouds / Gaussian splats). Early access only; it will power future Marble. No paper, no public API yet.
* Too much is happening, and I also have some scoops planned for today.
** I used Grok to compose this brief, cherry-picking the news and doing some post-editing.
β€13 9 5 5
Muse superapp from Meta and Ava model with computer use
META π₯: A new model named Ava with computer-use capabilities is undergoing closed testing in the Meta AI desktop app.
> "Agentic assistant with computer use."
Users can also enable apps for computer use individually, directly from the window attachment menu.
> "Clicks, types and scolls only in this window."
Watermelon, is this you? π
π #meta @testingcatalog
META π₯: A new model named Ava with computer-use capabilities is undergoing closed testing in the Meta AI desktop app.
> "Agentic assistant with computer use."
Users can also enable apps for computer use individually, directly from the window attachment menu.
> "Clicks, types and scolls only in this window."
Watermelon, is this you? π
π #meta @testingcatalog
TestingCatalog AI News
Muse superapp from Meta and Ava model with computer use
What we know so far: The production name of project Hatch will be "Muse". Meta is testing a new Ava model with computer-use capabilities internally.
β€4 4π1
GOOGLE π₯: Gemini 3.8 Flash started appearing on Google Coud Console quotas page, a usual release predecessor.
Earlier today, users also spotted that Gemini 3.8 Flash has been powering some of there conversations on Gemini already.
Very soon π
Earlier today, users also spotted that Gemini 3.8 Flash has been powering some of there conversations on Gemini already.
Very soon π
β€8 2
GOOGLE π₯: Gemini 3.8 Flash is already available in Agent Studio on GCP.
Best for
- Complex multimodal data processing
- Coding use cases
- Supporting software engineeringβrelated agentic tasks
Use case
- Processing data with images and text
- Coding problems
- Web research and application testing
Best for
- Complex multimodal data processing
- Coding use cases
- Supporting software engineeringβrelated agentic tasks
Use case
- Processing data with images and text
- Coding problems
- Web research and application testing