OPENAI π₯: ChatGPT Auto-review is now free for all uses signed via ChatGPT account.
The need to approve baby steps w/o providing a full access is still a challenge. Letβs see if that would make it less districting.
This feature runs a separate background agent that supervises actions taken by the main agent to make sure they are aligned with the user intent.
The need to approve baby steps w/o providing a full access is still a challenge. Letβs see if that would make it less districting.
β€10 4
Google is making Markdown files natively supported in Google Drive and Google Docs.
Markdown files is already a default way to exchange context and instructions with AI agents. This change will simplify so much!
Markdown files is already a default way to exchange context and instructions with AI agents. This change will simplify so much!
β€14 9π3
MISTRAL π₯: A new big model from Mistral is about to drop, according to Reuters.
Since we expect it to outperform Chinese models, there is a big Chunk of possibilities that it will be open weight. Right?
A big Chunk π
βThe model weβre actually announcing today is actually above the Chinese models on certain aspects, including cyber. So the narrative that Europe cannot compete is something that is not true,β
Since we expect it to outperform Chinese models, there is a big Chunk of possibilities that it will be open weight. Right?
A big Chunk π
β€5π₯5
This media is not supported in your browser
VIEW IN TELEGRAM
BREAKING π₯: Mistral announced Mistral Large 4 "Le Chonk", a new 1T-parameter open-weight model!
> 49B active parameters, native multimodality.
> Rolling out via APIs today; open-weight release is planned for the end of October.
> SOTA on "critical workloads", including cyber defense.
Le Chaton Fat "Le Chonk" is here π
> 49B active parameters, native multimodality.
> Rolling out via APIs today; open-weight release is planned for the end of October.
> SOTA on "critical workloads", including cyber defense.
Le Chaton Fat "Le Chonk" is here π
β€8π5
Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.3, according to VentureBeat.
Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).
Le Chonk also scores 15% on Harveyβs Legal Agent Benchmark (Legal tasks, SOTA open-weight).
We need a tech report now π
Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight).
Le Chonk also scores 15% on Harveyβs Legal Agent Benchmark (Legal tasks, SOTA open-weight).
We need a tech report now π
β€11π2πΏ1
Mistral launches Large 4 preview with 1 T parameters
Mistral Large 4 jumps to the 6th spot on the Open-Weight Artificial Analysis Index.
Le Chonk also scores 82 on the CyberGym-E2E (AA) benchmark. Similar performance appears across other Cyber benchmarks.
Needs testing now π
Overall, it seems like a decent release. The biggest question is how quickly Mistral can ship new model upgrades to stay relevant.
π #mistral @testingcatalog
Mistral Large 4 jumps to the 6th spot on the Open-Weight Artificial Analysis Index.
Le Chonk also scores 82 on the CyberGym-E2E (AA) benchmark. Similar performance appears across other Cyber benchmarks.
Needs testing now π
Overall, it seems like a decent release. The biggest question is how quickly Mistral can ship new model upgrades to stay relevant.
π #mistral @testingcatalog
TestingCatalog AI News
Mistral launches Large 4 preview with 1 T parameters
Mistral opened the Large 4 preview API with 49 billion active parameters and plans to release its model weights by the end of the month.
β€7π3
GOOGLE π₯: Nano Banana 2.1 is rolling out on Google AI Studio, APIs, and Gemini!
> "High-quality image generation and editing model with enhanced factuality, visual fidelity, and multi-turn consistency at Flash scale."
> Image output is priced at $30 per 1,000,000 tokens.
> Output images at 1K (1024x1024px) consume 1120 tokens and cost $0.034 per image.
> Output images at 2K (2048x2048px) consume 1680 tokens and cost $0.050 per image.
> Output images at 4K (4096x4096px) consume 2520 tokens and cost $0.076 per image.
Based on the pricing, it looks like it is a Pro model!?
> "High-quality image generation and editing model with enhanced factuality, visual fidelity, and multi-turn consistency at Flash scale."
> Image output is priced at $30 per 1,000,000 tokens.
> Output images at 1K (1024x1024px) consume 1120 tokens and cost $0.034 per image.
> Output images at 2K (2048x2048px) consume 1680 tokens and cost $0.050 per image.
> Output images at 4K (4096x4096px) consume 2520 tokens and cost $0.076 per image.
Based on the pricing, it looks like it is a Pro model!?
π¨ AI News | TestingCatalog
Figure plans to release Hark AI assistant this week Figure is planning to release its own proactive AI assistant this week! Here is what we know so far π > Hark will likely run on Opus 5.5 for general tasks, with Gemini 3.1 Flash Image, Qwen3.6-27B, andβ¦
Hark Pro is now available on web, iOS, and Android and a $100 Pro plan is now available for free.
Hark is a proactive AI agent from Figure, with separate screens for widgets, chats and projects.
App UIUX is very hot π₯
Hark is a proactive AI agent from Figure, with separate screens for widgets, chats and projects.
App UIUX is very hot π₯
β€6π3π2
Media is too big
VIEW IN TELEGRAM
Sesame has announced its own eyewear product, launching in 2027!
> "Our collection will be Made in Japan, available in 2027."
As we reported earlier, Sesame is working on new payment plans, browser use support, scheduled tasks, connectors, and loads of agentic features for its AI assistant app.
One of the most exciting products for me to monitor, besides major labs.
Soon π
> "Our collection will be Made in Japan, available in 2027."
As we reported earlier, Sesame is working on new payment plans, browser use support, scheduled tasks, connectors, and loads of agentic features for its AI assistant app.
One of the most exciting products for me to monitor, besides major labs.
Soon π
β€4π3
Google released EmbeddingGemma 2 open weight model under Apache 2 license!
Embedded testing time π
βOur lightweight, multimodal embedding model maps text, code, images, video, and audio into a single, unified embedding space.β
740M parameter form factor with modular encoders and 8K context window.
Embedded testing time π
β€8 5
This media is not supported in your browser
VIEW IN TELEGRAM
ANTHROPIC π₯: Claude is now available in Google Docs, Sheets, and Slides! Besides that, these file types now open inside Claude as well.
A small change but a big quality-of-life improvement!
A small change but a big quality-of-life improvement!
SESAME π₯: A huge upgrade arrived for the Sesame voice assistant, along with an updated, more expressive model, image generation, connectors, computer use, and more!
Sesame can now operate coding agents like Devin, Codex, or Claude Code.
ALL IN VOICE! πππ
Sesame can now operate coding agents like Devin, Codex, or Claude Code.
ALL IN VOICE! πππ
β€5π₯3π1π1
OpenAI published a repository with novel solutions to a range of math problems solved by their next frontier model.
The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model.
On average, each result used three hours of ChatGPT Pro thinking compute with that model.
Over the course of the evaluation, the model was posed approximately 4,000 problems.
π¨ AI News | TestingCatalog
DAILY AI BRIEF π β Oct 6 OPENAI π₯: * GPT-6 Astra and GPT-6.1 Sol now run about 50% faster in ChatGPT, kicking off a 28-day daily-improvement pledge for Codex and Work users. * API customers can now opt in to textGrain text watermarking, with ChatGPT and Codexβ¦
DAILY AI BRIEF π β Oct 7
MISTRAL π₯:
* Mistral released Mistral Large 4 "Le Chonk", a 1T-parameter (49B active) multimodal model, with open weights planned in about three weeks.
GOOGLE π₯:
* Nano Banana 2.1 is rolling out across the Gemini app, AI Mode, AI Studio, the Gemini API, Flow and Stitch.
* Google released EmbeddingGemma 2, a 740M-parameter open multimodal embedding model under Apache 2.0.
* Google AI Edge launched Foresight, an on-device note-taking app for Mac.
OPENAI π₯:
* OpenAI launched the Decisions API in beta with gpt-6-luna, returning typed answers 10x faster than the Responses API.
* Auto-review in Codex is now free for everyone signed in with ChatGPT, day two of the 28-day pledge.
* OpenAI published new math results from an unreleased internal frontier model, with many proofs formalized in Lean.
* OpenAI cut its API usage tiers from five to three: Build, Launch and Grow.
ANTHROPIC π₯:
* Claude for Google Workspace is in public beta, adding a Claude sidebar to Docs, Sheets and Slides.
* Anthropic's expanded Claude for Startups program gives qualifying startups a free year of Claude Team and $1,000 in API credits.
* Claude Code 2.1.292 lets Claude run sub-agents at a chosen effort level.
HARK π₯:
* Hark, Brett Adcock's new startup, launched Hark Pro widely as a free computer-use AI assistant, with a paid tier for heavy users.
SESAME π₯:
* Sesame's voice assistant got a new voice model, a dedicated computer for each agent, and Google app and MCP connections.
MICROSOFT π₯:
* GitHub stacked pull requests are now generally available.
MISTRAL π₯:
* Mistral released Mistral Large 4 "Le Chonk", a 1T-parameter (49B active) multimodal model, with open weights planned in about three weeks.
GOOGLE π₯:
* Nano Banana 2.1 is rolling out across the Gemini app, AI Mode, AI Studio, the Gemini API, Flow and Stitch.
* Google released EmbeddingGemma 2, a 740M-parameter open multimodal embedding model under Apache 2.0.
* Google AI Edge launched Foresight, an on-device note-taking app for Mac.
OPENAI π₯:
* OpenAI launched the Decisions API in beta with gpt-6-luna, returning typed answers 10x faster than the Responses API.
* Auto-review in Codex is now free for everyone signed in with ChatGPT, day two of the 28-day pledge.
* OpenAI published new math results from an unreleased internal frontier model, with many proofs formalized in Lean.
* OpenAI cut its API usage tiers from five to three: Build, Launch and Grow.
ANTHROPIC π₯:
* Claude for Google Workspace is in public beta, adding a Claude sidebar to Docs, Sheets and Slides.
* Anthropic's expanded Claude for Startups program gives qualifying startups a free year of Claude Team and $1,000 in API credits.
* Claude Code 2.1.292 lets Claude run sub-agents at a chosen effort level.
HARK π₯:
* Hark, Brett Adcock's new startup, launched Hark Pro widely as a free computer-use AI assistant, with a paid tier for heavy users.
SESAME π₯:
* Sesame's voice assistant got a new voice model, a dedicated computer for each agent, and Google app and MCP connections.
MICROSOFT π₯:
* GitHub stacked pull requests are now generally available.
β€6π1