Wow, now this is interesting!
CapCut is partnering with Google.
Soon, users will be able to edit images and videos directly inside the Gemini app using CapCutβs editing tools.
In other words, Gemini is basically getting a full editing timeline.
Whatβs especially interesting is that the announcement says nothing about Flow. Google, of course, loves building overlapping ecosystems and products...
Itβs also notable that Google still seems committed to the traditional content workflow, where editing remains a core step.
Meanwhile, Higgsfield is betting on an agentic approach β editing happens during generation itself, with AI deciding how the final clip should be assembled.
A very interesting move, especially since Dreamina and CapCut already support Seedance, Nanobanana, and other models. With the Gemini Omni API arriving, Googleβs own models will likely appear there too.
CapCut is partnering with Google.
Soon, users will be able to edit images and videos directly inside the Gemini app using CapCutβs editing tools.
In other words, Gemini is basically getting a full editing timeline.
Whatβs especially interesting is that the announcement says nothing about Flow. Google, of course, loves building overlapping ecosystems and products...
Itβs also notable that Google still seems committed to the traditional content workflow, where editing remains a core step.
Meanwhile, Higgsfield is betting on an agentic approach β editing happens during generation itself, with AI deciding how the final clip should be assembled.
A very interesting move, especially since Dreamina and CapCut already support Seedance, Nanobanana, and other models. With the Gemini Omni API arriving, Googleβs own models will likely appear there too.
π838β€31π21π₯17
This media is not supported in your browser
VIEW IN TELEGRAM
Gemini Omni: translation, lip-syncing, and dubbing with accurate lip matching.
Omni can translate audio even if the prompt doesnβt include either the original or the translated text:
β’ matches lip movements for the new language
β’ keeps the background music unchanged
β’ adjusts the video length when necessary (!)
For example, the Japanese and Spanish phrases during the close-up shot with the cream are longer, so Omni simply extends the clip by a couple of seconds on its own.
Omni can translate audio even if the prompt doesnβt include either the original or the translated text:
β’ matches lip movements for the new language
β’ keeps the background music unchanged
β’ adjusts the video length when necessary (!)
For example, the Japanese and Spanish phrases during the close-up shot with the cream are longer, so Omni simply extends the clip by a couple of seconds on its own.
β€127π107π92π₯90
A short prompt guide for Gemini Omni from Google:
Use real-world knowledge. Gemini Omni understands history, science, and culture, so you donβt need to over-explain. Use cultural references, eras, or scientific terms directly.
Control text rendering. Specify typography, placement, animation style, and visual effects like double exposure, synced with the video action.
Direct the camera. Think like a cinematographer: describe camera type, framing, movement, and shot style.
Edit precisely. Ask for targeted changes, such as replacing a background or caption, without rewriting the whole prompt. Omni can preserve the video structure while adjusting movement, emotion, or object interaction.
https://x.com/GoogleAI/status/2059381218660270435
Use real-world knowledge. Gemini Omni understands history, science, and culture, so you donβt need to over-explain. Use cultural references, eras, or scientific terms directly.
Control text rendering. Specify typography, placement, animation style, and visual effects like double exposure, synced with the video action.
Direct the camera. Think like a cinematographer: describe camera type, framing, movement, and shot style.
Edit precisely. Ask for targeted changes, such as replacing a background or caption, without rewriting the whole prompt. Omni can preserve the video structure while adjusting movement, emotion, or object interaction.
https://x.com/GoogleAI/status/2059381218660270435
π331π₯130π110β€104
This media is not supported in your browser
VIEW IN TELEGRAM
A metaverse worth having.
At Google I/O, they combined Street View data from Google Maps (collected over the last 20 years, mind you) with Project Genie.
Now, in this world-generation system, you can specify any point on the map and use a prompt to describe how the real world should be transformed. In the example above, a flooded bridge in San Francisco becomes part of an underwater world.
In other words, you can create virtually any "skin" for any real-world location β and not just in 2D, but in full 3D.
Instead of location scouts, we'll soon have location prompters.
Now that's the kind of metaverse we need.
For now, it only works with locations in the US.
At Google I/O, they combined Street View data from Google Maps (collected over the last 20 years, mind you) with Project Genie.
Now, in this world-generation system, you can specify any point on the map and use a prompt to describe how the real world should be transformed. In the example above, a flooded bridge in San Francisco becomes part of an underwater world.
In other words, you can create virtually any "skin" for any real-world location β and not just in 2D, but in full 3D.
Instead of location scouts, we'll soon have location prompters.
Now that's the kind of metaverse we need.
For now, it only works with locations in the US.
π289β€97π₯93π81
Gemini Omni vs Grok vs Seedance
Omni is the one with 10 seconds.
And youβll recognize Seedance by the dynamics and the edit cut at the very end.
Omni is the one with 10 seconds.
And youβll recognize Seedance by the dynamics and the edit cut at the very end.
π990π340β€327π₯310
This media is not supported in your browser
VIEW IN TELEGRAM
Ideogram 4 has gone open source!
According to current benchmarks and arena rankings, it's one of the strongest open-source image generators available.
"Ideogram 4 is Ideogram's first open-weights model, trained from scratch. It features a structured JSON prompt format, best-in-class multilingual text rendering, deep language understanding, color palette control, and native 2K generation with aspect ratios up to 6:1."
With just 9.3B parameters, it should run on modest GPUs (Qwen-Image: 20B, FLUX.2 [dev]: 32B).
For the geeks:
β’ Flow-matching text-to-image model built on a fully single-stream DiT architecture. β’ Text and image tokens are processed together by the same 34-layer transformer. β’ Uses Qwen3-VL-8B-Instruct, providing richer visual understanding than CLIP or T5.
Trained on JSON annotations and includes a prompt enhancer and guide.
Some content moderation is built in (Hive AI). The license is non-commercial.
GitHub repo (weights, code, docs, guides): https://github.com/ideogram-oss/ideogram4
According to current benchmarks and arena rankings, it's one of the strongest open-source image generators available.
"Ideogram 4 is Ideogram's first open-weights model, trained from scratch. It features a structured JSON prompt format, best-in-class multilingual text rendering, deep language understanding, color palette control, and native 2K generation with aspect ratios up to 6:1."
With just 9.3B parameters, it should run on modest GPUs (Qwen-Image: 20B, FLUX.2 [dev]: 32B).
For the geeks:
β’ Flow-matching text-to-image model built on a fully single-stream DiT architecture. β’ Text and image tokens are processed together by the same 34-layer transformer. β’ Uses Qwen3-VL-8B-Instruct, providing richer visual understanding than CLIP or T5.
Trained on JSON annotations and includes a prompt enhancer and guide.
Some content moderation is built in (Hive AI). The license is non-commercial.
GitHub repo (weights, code, docs, guides): https://github.com/ideogram-oss/ideogram4
π57β€55π₯49π38
Media is too big
VIEW IN TELEGRAM
3D generators are coming for the sacred territory β organic modeling, specifically head modeling.
This is Rodin 2.5. They offer 12K textures and generation at 10 million polygons.
The input is several head images from different angles β either generated images or photos.
https://hyper3d.ai/
This is Rodin 2.5. They offer 12K textures and generation at 10 million polygons.
The input is several head images from different angles β either generated images or photos.
https://hyper3d.ai/
π78π78β€75π₯65
Are Game devs in trouble?
Fable 5 is the biggest step up I've felt in our models since Opus 4.5 back in November.
We just one shot GTA 6 and it works. No miskates.
One prompt. 15 minutes. The whole game.
Rockstar took 12 years. Fable took a quarter of an hour.
Mythos was deemed too dangerous to release, this is the nerfed version. And it still does THIS.
Game devs, are we in trouble?
Fable 5 is the biggest step up I've felt in our models since Opus 4.5 back in November.
We just one shot GTA 6 and it works. No miskates.
One prompt. 15 minutes. The whole game.
Rockstar took 12 years. Fable took a quarter of an hour.
Mythos was deemed too dangerous to release, this is the nerfed version. And it still does THIS.
Game devs, are we in trouble?
π83β€67π67π₯59
This media is not supported in your browser
VIEW IN TELEGRAM
Krea 2: Procedural Images
Krea now has real-time sliders: intensity, complexity, and movement.
Basically: prompt intensity, noise/details, and object movement within the frame.
This definitely makes the generation process more interactive and speeds up image selection.
It also reminds me of procedural textures. If youβve worked with them, youβll remember all those endless knobs for tweaking patterns, noise, and textures.
If you havenβt, think of the effect preview in Photoshop: you see the result immediately.
I wonder how many more sliders they could add: color grading, 3D depth, typography, layer separation, backgroundβ¦
The idea is to give regular users as many clear and understandable controls as possible β instead of all that CFG Scale, Sampling Method, and so on β and keep them inside the generation interface.
Krea now has real-time sliders: intensity, complexity, and movement.
Basically: prompt intensity, noise/details, and object movement within the frame.
This definitely makes the generation process more interactive and speeds up image selection.
It also reminds me of procedural textures. If youβve worked with them, youβll remember all those endless knobs for tweaking patterns, noise, and textures.
If you havenβt, think of the effect preview in Photoshop: you see the result immediately.
I wonder how many more sliders they could add: color grading, 3D depth, typography, layer separation, backgroundβ¦
The idea is to give regular users as many clear and understandable controls as possible β instead of all that CFG Scale, Sampling Method, and so on β and keep them inside the generation interface.
π183π₯181π172β€164
AI video is getting wild.
On June 15, Dreamina is releasing Seedance 2.0 Mini, and this is huge:
Same performance level as Seedance 2.0, but much cheaper.
That means creators can generate more AI drama shots, test more ideas, and build cinematic videos without burning insane credits.
AI video is not slowing down.
On June 15, Dreamina is releasing Seedance 2.0 Mini, and this is huge:
Same performance level as Seedance 2.0, but much cheaper.
That means creators can generate more AI drama shots, test more ideas, and build cinematic videos without burning insane credits.
AI video is not slowing down.
π₯209β€203π202π197
Reverse Ideogram 4
Interesting tool: you give it any image as input, and it outputs a JSON prompt for Ideogram 4, complete with bounding boxes and labels for them.
You can edit specific parts of images locally.
Thereβs code here:
https://github.com/cocktailpeanut/image-to-prompt
Under the hood, it uses Microsoftβs Florence-2 model for image recognition.
And if you donβt want to mess with running it locally, thereβs a Space where you can try it online:
https://huggingface.co/spaces/cocktailpeanut/image-to-prompt
Interesting tool: you give it any image as input, and it outputs a JSON prompt for Ideogram 4, complete with bounding boxes and labels for them.
You can edit specific parts of images locally.
Thereβs code here:
https://github.com/cocktailpeanut/image-to-prompt
Under the hood, it uses Microsoftβs Florence-2 model for image recognition.
And if you donβt want to mess with running it locally, thereβs a Space where you can try it online:
https://huggingface.co/spaces/cocktailpeanut/image-to-prompt
β€603π587π578π₯564
Media is too big
VIEW IN TELEGRAM
Grok Imagine Video 1.5
Itβs out of Preview.
The interesting part: 720p, 15 seconds β but the SuperGrok subscription page says β30-second videos.β
Thatβs a marketing trick. The API clearly says 15 seconds. The β30 secondsβ refers to the Extend from Frame feature, which is available there.
Thereβs also Grok Imagine Video 1.5 Fast β it generates 720p videos in 25 seconds.
https://x.ai/news/grok-imagine-video-1-5
Itβs out of Preview.
The interesting part: 720p, 15 seconds β but the SuperGrok subscription page says β30-second videos.β
Thatβs a marketing trick. The API clearly says 15 seconds. The β30 secondsβ refers to the Extend from Frame feature, which is available there.
Thereβs also Grok Imagine Video 1.5 Fast β it generates 720p videos in 25 seconds.
https://x.ai/news/grok-imagine-video-1-5
β€85π₯73π70π70
Who will win the World Cup before it happens?
Soccer Buddy uses AI-style prediction logic, 80+ match parameters, 10,000 simulations, and deep pro stats to find stronger football angles before kickoff.
xG, BTTS, overs, corners, cards, shots, possession, home and away splits.
AI + stats = smarter soccer predictions.
https://zcodesystem.com/soccerbuddy/?WC2026
Soccer Buddy uses AI-style prediction logic, 80+ match parameters, 10,000 simulations, and deep pro stats to find stronger football angles before kickoff.
xG, BTTS, overs, corners, cards, shots, possession, home and away splits.
AI + stats = smarter soccer predictions.
https://zcodesystem.com/soccerbuddy/?WC2026
π195β€184π₯165π153
Seedance 2.5
ByteDance announced the new version at Volcano Engine FORCE 2026.
The interesting part: native 30-second video generation in a single run. Previous public versions were limited to 15 seconds, so this is a major jump.
Another big upgrade: up to 50 multimodal references β images, videos, and audio β to control style, characters, motion, and editing.
Seedance 2.0 also received an upgrade and now supports native 4K video generation.
The model is currently in enterprise beta and is expected to launch publicly in early July.
ByteDance announced the new version at Volcano Engine FORCE 2026.
The interesting part: native 30-second video generation in a single run. Previous public versions were limited to 15 seconds, so this is a major jump.
Another big upgrade: up to 50 multimodal references β images, videos, and audio β to control style, characters, motion, and editing.
Seedance 2.0 also received an upgrade and now supports native 4K video generation.
The model is currently in enterprise beta and is expected to launch publicly in early July.
π₯89β€79π70π63
π€ AI company Anthropic is expanding its presence in Europe
Anthropic, the company behind the Claude AI assistant, has hired the head of artificial intelligence from the French telecom company Orange. This move is part of Anthropicβs major expansion into the European market.
It seems that competition between leading AI companies is no longer just about building the best models β it is also about attracting the worldβs top experts.
Anthropic, the company behind the Claude AI assistant, has hired the head of artificial intelligence from the French telecom company Orange. This move is part of Anthropicβs major expansion into the European market.
It seems that competition between leading AI companies is no longer just about building the best models β it is also about attracting the worldβs top experts.
π228π₯224π222β€187
This media is not supported in your browser
VIEW IN TELEGRAM
Comfy MCP
Comfy Org has launched the beta version of its MCP server for ComfyUI.
MCP allows AI agents to interact with and manage ComfyUI more easily:
β’ Run workflows using natural-language descriptions
β’ Search for models, nodes, and templates across the ecosystem
β’ Import workflows directly from a URL
β’ Rerun saved workflows with new inputs
β’ Access hundreds of ready-made workflows with automatic updates
Documentation
They are also launching the beta version of Comfy CLI and the Comfy Skill repository for AI agents.
Comfy Org has launched the beta version of its MCP server for ComfyUI.
MCP allows AI agents to interact with and manage ComfyUI more easily:
β’ Run workflows using natural-language descriptions
β’ Search for models, nodes, and templates across the ecosystem
β’ Import workflows directly from a URL
β’ Rerun saved workflows with new inputs
β’ Access hundreds of ready-made workflows with automatic updates
Documentation
They are also launching the beta version of Comfy CLI and the Comfy Skill repository for AI agents.
π₯232β€227π224π216