GOOGLE π₯: Pre training of Gemini 4 model has begun!
> Google is normally running 6 month training cycles and potentially we should expect Gemini 4 to land around the end of the year.
βWen Gemini 4β time! π
> Google is normally running 6 month training cycles and potentially we should expect Gemini 4 to land around the end of the year.
βWen Gemini 4β time! π
β€12π€£12π₯2π1
OPENAI π₯: ChatGPT Work and Codex reached 10 million active users milestone!
> x2 growth within the past week π
> ChatGPT Work and Codex usage reset is happening too.
> x2 growth within the past week π
> ChatGPT Work and Codex usage reset is happening too.
β€11 3π₯2π1
Google released "Gemini 3.5 Flash Cyber" on CodeMender, a new model for finding security vulnerabilities.
> Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together to produce a single combined report, 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym.
> Flashβs performance and efficiency makes it an ideal foundation for our cybersecurity model efforts. By building on top of Flash, 3.5 Flash Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models.
> Within CodeMender, which uses multiple 3.5 Flash Cyber agents working together to produce a single combined report, 3.5 Flash Cyber reaches competitive performance at the frontier on the popular benchmark CyberGym.
> Flashβs performance and efficiency makes it an ideal foundation for our cybersecurity model efforts. By building on top of Flash, 3.5 Flash Cyber offers a cost-efficient and highly capable alternative to large, costly cybersecurity models.
π€£8β€4
BREAKING π₯: An "even more capable pre-release model" than GPT-5.6 Sol, managed to find a 0-day vulnerability in order to gain public internet access and acquire evaluation data from Huggingface's production database in order to gain a higher score on the evaluation benchmark.
> After investigating, we now know that this particular incident was driven by a combination of OpenAI models, including GPTβ5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes.
> While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.
> The models identified and chained vulnerabilities across OpenAIβs research environment and Hugging Faceβs production infrastructure to obtain test solutions directly from Hugging Faceβs production database.
Pentesting time π
> After investigating, we now know that this particular incident was driven by a combination of OpenAI models, including GPTβ5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes.
> While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.
> The models identified and chained vulnerabilities across OpenAIβs research environment and Hugging Faceβs production infrastructure to obtain test solutions directly from Hugging Faceβs production database.
Pentesting time π
π10π€―4β€2 1
ANTHROPIC π₯: A new capability to work with an iOS simulator has been added to Claude Code desktop.
Support for Android simulators is in the works too! (Currently not available yet)
Users will be able to disable this feature in settings when needed.
> Let Claude verify your changes in Android emulators on this Mac: running your app, driving it through flows, and capturing screenshots and recordings. You will be asked before Claude uses each device. When off, Claude doesnβt get its emulator tools, and you can still use the emulator in the app yourself.
Eventually, this will open up a huge range of tasks that Claude will be able to run on the mobile device.
Another startup killer feature? π
Support for Android simulators is in the works too! (Currently not available yet)
Users will be able to disable this feature in settings when needed.
> Let Claude verify your changes in Android emulators on this Mac: running your app, driving it through flows, and capturing screenshots and recordings. You will be asked before Claude uses each device. When off, Claude doesnβt get its emulator tools, and you can still use the emulator in the app yourself.
Eventually, this will open up a huge range of tasks that Claude will be able to run on the mobile device.
Another startup killer feature? π
β€6π2
Google launches Gemini 3.6 Flash and Gemini 3.5 Flash Lite
Google introduced Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber for AI agents, focused on lower latency, lower token use, coding, multimodal work, and restricted cybersecurity tasks for production and enterprise use.
π #google @testingcatalog
Google introduced Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber for AI agents, focused on lower latency, lower token use, coding, multimodal work, and restricted cybersecurity tasks for production and enterprise use.
π #google @testingcatalog
TestingCatalog AI News
Google launches Gemini 3.6 Flash and Gemini 3.5 Flash Lite
What's new? Gemini 3.6 flash cuts token use for coding and data, and gemini 3.5 flash-lite runs at 350 tps; gemini 3.5 flash cyber targets vulnerability detection in pilot;
π6β€3
Anthropic develops Claude-driven Managed Projects
Anthropic is testing βmanagedβ Claude projects: persistent workspaces that keep context, organize tasks, and may run scheduled work. Internal builds tie prior agent and memory features into one shared or personal project surface.
π #anthropic @testingcatalog
Anthropic is testing βmanagedβ Claude projects: persistent workspaces that keep context, organize tasks, and may run scheduled work. Internal builds tie prior agent and memory features into one shared or personal project surface.
π #anthropic @testingcatalog
TestingCatalog AI News
Anthropic develops Claude-driven Managed Projects
Anthropic is testing managed Claude projects, letting the assistant autonomously organize tasks, maintain context, and run scheduled work.
β€4π2 1
π¨ AI News | TestingCatalog
Anthropic develops Claude-driven Managed Projects Anthropic is testing βmanagedβ Claude projects: persistent workspaces that keep context, organize tasks, and may run scheduled work. Internal builds tie prior agent and memory features into one shared or personalβ¦
This media is not supported in your browser
VIEW IN TELEGRAM
ANTHROPIC π₯: A new Managed Projects feature has been spotted in testing on Claude.
> Claude takes on tasks and keeps the project organized.
> A project is a home for one stream of work. Sessions share memory and instructions so context carries forward, and Claude runs more autonomously. Projects only create and manage cloud sessions.
This feature may be built on top of Claude Managed Agents, where each project gets a dedicated cloud environment so Claude can execute periodic tasks and refine project context via "Dreams". It could also be a successor to Conway, which is set to be removed this Friday internally.
> Claude takes on tasks and keeps the project organized.
> A project is a home for one stream of work. Sessions share memory and instructions so context carries forward, and Claude runs more autonomously. Projects only create and manage cloud sessions.
This feature may be built on top of Claude Managed Agents, where each project gets a dedicated cloud environment so Claude can execute periodic tasks and refine project context via "Dreams". It could also be a successor to Conway, which is set to be removed this Friday internally.
β€4π4
GLASSES π₯: Samsung revealed 2 new Smart Glasses at Galaxy Unpacked in London. The new glasses were designed and produced in a partnership with Gentle Monster and Warby Parker.
> The intelligent eyewear shows how the Galaxy ecosystem can move beyond the mobile phone and into eyewear that supports daily routines, work, travel, and hands-free moments.
What is yet unclear to me: when? π
> The intelligent eyewear shows how the Galaxy ecosystem can move beyond the mobile phone and into eyewear that supports daily routines, work, travel, and hands-free moments.
What is yet unclear to me: when? π
π4β€2
This media is not supported in your browser
VIEW IN TELEGRAM
Cursor announced Cursor Router, a new model router allowing users to access frontier performance at 60% lower cost.
> Cursor Router analyzes each request and routes to the best model for the job: frontier models when the work demands them and price-efficient models when it doesn't.
Routers hold huge business value, and this will continue to be the trend. I bet that in the long run, routers will also be pushing frontier capabilities beyond what pure models will be able to offer (due to multi-model routing), not just cost optimization.
> Cursor Router analyzes each request and routes to the best model for the job: frontier models when the work demands them and price-efficient models when it doesn't.
Routers hold huge business value, and this will continue to be the trend. I bet that in the long run, routers will also be pushing frontier capabilities beyond what pure models will be able to offer (due to multi-model routing), not just cost optimization.
β€12π7π€©2π1
Anthropic expanded access to the Claude Security plugin for Claude Code CLI.
> Scan your changes for vulnerabilities before you commit, or run a full scan across your codebase, all from your terminal on the Claude inference you already run.
We finally see these tools becoming more accessible to users!
> Scan your changes for vulnerabilities before you commit, or run a full scan across your codebase, all from your terminal on the Claude inference you already run.
We finally see these tools becoming more accessible to users!
β€6π1
This media is not supported in your browser
VIEW IN TELEGRAM
Anthropic introduced a new Anthropic Economic Index connector to allow Claude to access a large dataset on how people use Claude around the world.
> Explore public data on how the world uses Claude.
> Explore public data on how the world uses Claude.
β€5π₯2
Claude Voice Mode to get Opus and Sonnet model options
Anthropic appears close to rolling out upgraded Claude voice mode, with Opus and Sonnet now powering voice chats instead of defaulting to Haiku.
π #anthropic @testingcatalog
Anthropic appears close to rolling out upgraded Claude voice mode, with Opus and Sonnet now powering voice chats instead of defaulting to Haiku.
π #anthropic @testingcatalog
TestingCatalog AI News
Claude Voice Mode to get Opus and Sonnet model options
Anthropic updates Claudeβs voice mode with support for Opus and Sonnet models in place of Haiku, hinting at a broader release in the coming days.
β€6 3π2
π¨ AI News | TestingCatalog
Claude Voice Mode to get Opus and Sonnet model options Anthropic appears close to rolling out upgraded Claude voice mode, with Opus and Sonnet now powering voice chats instead of defaulting to Haiku. π #anthropic @testingcatalog
Media is too big
VIEW IN TELEGRAM
ANTHROPIC π₯: Claude Opus and Sonnet models will soon become available on Claude Voice Mode!
The new model selector has been hidden in the UI for roughly three weeks, defaulting to Haiku regardless of the selection.
Today, it changed, and now it responds via Opus and Sonnet models. The system is still built on TTS; however, it handles interruptions super well and can operate Connectors!
Access to a broader range of tools makes it much more capable than what ChatGPT is offering at the moment, yet usage might become a tricky topic, at least for high-effort reasoning.
The new model selector has been hidden in the UI for roughly three weeks, defaulting to Haiku regardless of the selection.
Today, it changed, and now it responds via Opus and Sonnet models. The system is still built on TTS; however, it handles interruptions super well and can operate Connectors!
Access to a broader range of tools makes it much more capable than what ChatGPT is offering at the moment, yet usage might become a tricky topic, at least for high-effort reasoning.
β€6π3 2 1
SpaceXAI develops deployable applications for Grok Build
xAI is building Grok into an app publishing platform with deployment, hosting, sharing, and custom domains, while mobile top-ups appear closer and Grok 4.6 training nears completion ahead of a possible August release.
π #spacexai @testingcatalog
xAI is building Grok into an app publishing platform with deployment, hosting, sharing, and custom domains, while mobile top-ups appear closer and Grok 4.6 training nears completion ahead of a possible August release.
π #spacexai @testingcatalog
TestingCatalog AI News
SpaceXAI develops deployable applications for Grok Build
xAI is developing tools for users to deploy and share Grok-built apps, plus bringing credit top-ups to mobile, with Grok 4.6 arriving soon.
π4β€2π2