Bagel Labs launches WorldDiT world model for robotics
Bagel Labs released WorldDiT, an open robotics world model that jointly predicts robot actions and future scenes with one diffusion backbone, delivering strong LIBERO performance at sub-billion scale for on-robot deployment.
π #ai @testingcatalog
Bagel Labs released WorldDiT, an open robotics world model that jointly predicts robot actions and future scenes with one diffusion backbone, delivering strong LIBERO performance at sub-billion scale for on-robot deployment.
π #ai @testingcatalog
TestingCatalog AI News
Bagel Labs launches WorldDiT world model for robotics
WorldDiT learns robot actions and future scene states in one shared model, with Bagel reporting frontier LIBERO results under 1B params.
SpaceXAI launches Grok Voice Think Fast 2.0 on Agent Builder
xAI launched Grok Voice Think Fast 2.0, a speech model for voice agents with higher benchmark scores, faster first audio, stronger transcription across 24 languages, and lower latency. It costs $0.08 per audio minute.
π #spacexai @testingcatalog
xAI launched Grok Voice Think Fast 2.0, a speech model for voice agents with higher benchmark scores, faster first audio, stronger transcription across 24 languages, and lower latency. It costs $0.08 per audio minute.
π #spacexai @testingcatalog
TestingCatalog AI News
SpaceXAI launches Grok Voice Think Fast 2.0 on Agent Builder
Grok Voice Think Fast 2.0 is available at $0.08 per audio minute, with grok-voice-latest switching to the new model on August 5.
β€2π₯2π1
GOOGLE π₯: Lyria 3.5 has been released on Google Flow Music! Besides that, Flow Music now has covers, lip-sync videos, and an iOS app.
> Meet Lyria 3.5. Experience dynamic vocals, richer musicality, and advanced creative controls with our new flagship model.
> Reimagine your music with Covers. Transform songs into a completely new style while keeping the original structure intact.
> Take the studio with you. Download the Google Flow Music iOS app to create, listen, and share from anywhere.
> Direct lip-synced music videos. Use Gemini Omni Flash and new lip-syncing capabilities to create stunning visuals.
> Meet Lyria 3.5. Experience dynamic vocals, richer musicality, and advanced creative controls with our new flagship model.
> Reimagine your music with Covers. Transform songs into a completely new style while keeping the original structure intact.
> Take the studio with you. Download the Google Flow Music iOS app to create, listen, and share from anywhere.
> Direct lip-synced music videos. Use Gemini Omni Flash and new lip-syncing capabilities to create stunning visuals.
β€9π5
This media is not supported in your browser
VIEW IN TELEGRAM
CURSOR π₯: The latest version of Cursor for iOS is now compatible with iPad! Besides that, the app got a new Inbox and a complete PR review experience.
I need more devices for testing π
I need more devices for testing π
β€4π4π1
Revolut π€ OpenAI
Revolut introduced ChatGPT Go subscription plan as a benefit for their customers.
> Starting from 3 months free on Standard, you can unlock up to 12 months of ChatGPT Go included at no extra cost, depending on your plan.
ChatGPT user base is about to grow quite a lot soon!
Revolut introduced ChatGPT Go subscription plan as a benefit for their customers.
> Starting from 3 months free on Standard, you can unlock up to 12 months of ChatGPT Go included at no extra cost, depending on your plan.
ChatGPT user base is about to grow quite a lot soon!
β€8π3π₯3
Google introduced Gemini Robotics ER 2, a new embodied reasoning model!
Benchmarks π
> Success/failure detection: Now operates on raw video feeds rather than static snapshots to catch mid-execution failures like spills, slips, or misalignments.
> General instrument reading: Extends beyond circular dials and sight glasses to include digital displays, linear scales, rulers, and liquid thermometers. We tested it across 10 different types of instruments.
> Enhanced spatial VQA: Improves Visual Question Answering throughGeminiβs advancements in multi-modal understanding.
Benchmarks π
> Success/failure detection: Now operates on raw video feeds rather than static snapshots to catch mid-execution failures like spills, slips, or misalignments.
> General instrument reading: Extends beyond circular dials and sight glasses to include digital displays, linear scales, rulers, and liquid thermometers. We tested it across 10 different types of instruments.
> Enhanced spatial VQA: Improves Visual Question Answering throughGeminiβs advancements in multi-modal understanding.
β€6π₯1π€1
OPENAI π₯: GPT-5.6 Luna prices got reduced by 80% along with a 20% cut for GPT-5.6 Terra!
GPT-5.6 Sol got a new faster option on the API with a 2.5x speed boost at 2x price.
Many models got overshadowed π
GPT-5.6 Sol got a new faster option on the API with a 2.5x speed boost at 2x price.
Many models got overshadowed π
β€9π3π¦2π€©1
Media is too big
VIEW IN TELEGRAM
PERPLEXITY π₯: Spaces got upgraded to Projects, a new type of workspaces for collaboration with Perplexity Computer, powered by a shared file system and self-improving Brain memory!
> Between tasks, Brain reviews the Project's files and sessions and updates what it knows, so each task starts with full context from previous work.
> Between tasks, Brain reviews the Project's files and sessions and updates what it knows, so each task starts with full context from previous work.
β€5π5
Dreamina Seedance 2.5 is now available on Dreamina AI for paid plans in Southeast Asia, the Middle East, Africa, Europe, and South America.
No US π
- Native 30s videos
- A new interactive editing experience
- Long video mode (up to 3 minutes)
- Dreamina AI plugins for Maya and Blender
- Up to 50 multimodal references
- More true-to-life lighting and shadows
No US π
- Native 30s videos
- A new interactive editing experience
- Long video mode (up to 3 minutes)
- Dreamina AI plugins for Maya and Blender
- Up to 50 multimodal references
- More true-to-life lighting and shadows
β€9π₯2π΄1
MiniMax H3 is now available on HailuoAI & MiniMax APIs!
> All-in-One Reference - Creating Videos Using Text, Image, Audio, and Video Inputs
> Precise editing controls - Improve and iterate by adhering to precise guidelines.
> Built for all creative scenarios - from movies and commercials to games, brands, and e-commerce
> 2K video from $0.081/sec, 768p video coming soon, from $0.047/sec
> All-in-One Reference - Creating Videos Using Text, Image, Audio, and Video Inputs
> Precise editing controls - Improve and iterate by adhering to precise guidelines.
> Built for all creative scenarios - from movies and commercials to games, brands, and e-commerce
> 2K video from $0.081/sec, 768p video coming soon, from $0.047/sec
β€6π₯3π΄2
This media is not supported in your browser
VIEW IN TELEGRAM
OPENAI π₯: The built-in web browser in the ChatGPT app is becoming more mature. Now it supports URL suggestions during typing.
Besides that, the ChatGPT Chrome extension can now reference open tabs, highlight to ask, and more!
I may need to try it as a default π
Besides that, the ChatGPT Chrome extension can now reference open tabs, highlight to ask, and more!
I may need to try it as a default π
β€6π₯4π1
No more AI Studio for mobile? π
But instead, we will be getting something else! Any guesses on what this could be?
> Weβve decided to take an entirely different approach: one where apps emerge naturally, in the course of your everyday conversations with Gemini.
> Weβre partnering with the Gemini app team to make that a reality on mobile and desktop.
But instead, we will be getting something else! Any guesses on what this could be?
> Weβve decided to take an entirely different approach: one where apps emerge naturally, in the course of your everyday conversations with Gemini.
> Weβre partnering with the Gemini app team to make that a reality on mobile and desktop.
β€11β5
Gemini for macOS adds new "Speak to Window" feature
Google is rolling out Gemini voice on macOS for English users worldwide. A long press of Fn enables polished dictation and on-screen reasoning to summarize files, rewrite text, and create or edit images without leaving the current window.
π #google @testingcatalog
Google is rolling out Gemini voice on macOS for English users worldwide. A long press of Fn enables polished dictation and on-screen reasoning to summarize files, rewrite text, and create or edit images without leaving the current window.
π #google @testingcatalog
TestingCatalog AI News
Gemini for macOS adds new "Speak to Window" feature
New Gemini voice tools are rolling out globally in English on macOS, with Fn-key dictation and optional reasoning based on on-screen context.
β€5π1
Google started rolling out Gemini Spark to Google AI Pro users outside the U.S.
> Spark is your personal AI agent that works in the background 24/7 to get things done under your direction, handling the heavy lifting so you can focus on what matters.
Letβs see how long will that take. The next big upgrade there would be to get support for real MCPs (which is in the works at least).
> Spark is your personal AI agent that works in the background 24/7 to get things done under your direction, handling the heavy lifting so you can focus on what matters.
Letβs see how long will that take. The next big upgrade there would be to get support for real MCPs (which is in the works at least).
π6β€5