This media is not supported in your browser
VIEW IN TELEGRAM
Gemini Omni Flash.
You donโt need to spell out every step for it โ itโs enough to formulate the concept of the video.
It independently finds the theory, descriptions of objects and details, visualizes them, and adds text.
Education will never be the same again.
P.S. Directing educational videos is a different kind of storytelling. You donโt need fights and chase scenes. You need concepts. And they already exist within it.
You donโt need to spell out every step for it โ itโs enough to formulate the concept of the video.
It independently finds the theory, descriptions of objects and details, visualizes them, and adds text.
Education will never be the same again.
P.S. Directing educational videos is a different kind of storytelling. You donโt need fights and chase scenes. You need concepts. And they already exist within it.
๐384โค152๐120๐ฅ117
Wow, now this is interesting!
CapCut is partnering with Google.
Soon, users will be able to edit images and videos directly inside the Gemini app using CapCutโs editing tools.
In other words, Gemini is basically getting a full editing timeline.
Whatโs especially interesting is that the announcement says nothing about Flow. Google, of course, loves building overlapping ecosystems and products...
Itโs also notable that Google still seems committed to the traditional content workflow, where editing remains a core step.
Meanwhile, Higgsfield is betting on an agentic approach โ editing happens during generation itself, with AI deciding how the final clip should be assembled.
A very interesting move, especially since Dreamina and CapCut already support Seedance, Nanobanana, and other models. With the Gemini Omni API arriving, Googleโs own models will likely appear there too.
CapCut is partnering with Google.
Soon, users will be able to edit images and videos directly inside the Gemini app using CapCutโs editing tools.
In other words, Gemini is basically getting a full editing timeline.
Whatโs especially interesting is that the announcement says nothing about Flow. Google, of course, loves building overlapping ecosystems and products...
Itโs also notable that Google still seems committed to the traditional content workflow, where editing remains a core step.
Meanwhile, Higgsfield is betting on an agentic approach โ editing happens during generation itself, with AI deciding how the final clip should be assembled.
A very interesting move, especially since Dreamina and CapCut already support Seedance, Nanobanana, and other models. With the Gemini Omni API arriving, Googleโs own models will likely appear there too.
๐838โค31๐21๐ฅ17
This media is not supported in your browser
VIEW IN TELEGRAM
Gemini Omni: translation, lip-syncing, and dubbing with accurate lip matching.
Omni can translate audio even if the prompt doesnโt include either the original or the translated text:
โข matches lip movements for the new language
โข keeps the background music unchanged
โข adjusts the video length when necessary (!)
For example, the Japanese and Spanish phrases during the close-up shot with the cream are longer, so Omni simply extends the clip by a couple of seconds on its own.
Omni can translate audio even if the prompt doesnโt include either the original or the translated text:
โข matches lip movements for the new language
โข keeps the background music unchanged
โข adjusts the video length when necessary (!)
For example, the Japanese and Spanish phrases during the close-up shot with the cream are longer, so Omni simply extends the clip by a couple of seconds on its own.
โค127๐107๐92๐ฅ90
A short prompt guide for Gemini Omni from Google:
Use real-world knowledge. Gemini Omni understands history, science, and culture, so you donโt need to over-explain. Use cultural references, eras, or scientific terms directly.
Control text rendering. Specify typography, placement, animation style, and visual effects like double exposure, synced with the video action.
Direct the camera. Think like a cinematographer: describe camera type, framing, movement, and shot style.
Edit precisely. Ask for targeted changes, such as replacing a background or caption, without rewriting the whole prompt. Omni can preserve the video structure while adjusting movement, emotion, or object interaction.
https://x.com/GoogleAI/status/2059381218660270435
Use real-world knowledge. Gemini Omni understands history, science, and culture, so you donโt need to over-explain. Use cultural references, eras, or scientific terms directly.
Control text rendering. Specify typography, placement, animation style, and visual effects like double exposure, synced with the video action.
Direct the camera. Think like a cinematographer: describe camera type, framing, movement, and shot style.
Edit precisely. Ask for targeted changes, such as replacing a background or caption, without rewriting the whole prompt. Omni can preserve the video structure while adjusting movement, emotion, or object interaction.
https://x.com/GoogleAI/status/2059381218660270435
๐331๐ฅ130๐110โค104
This media is not supported in your browser
VIEW IN TELEGRAM
A metaverse worth having.
At Google I/O, they combined Street View data from Google Maps (collected over the last 20 years, mind you) with Project Genie.
Now, in this world-generation system, you can specify any point on the map and use a prompt to describe how the real world should be transformed. In the example above, a flooded bridge in San Francisco becomes part of an underwater world.
In other words, you can create virtually any "skin" for any real-world location โ and not just in 2D, but in full 3D.
Instead of location scouts, we'll soon have location prompters.
Now that's the kind of metaverse we need.
For now, it only works with locations in the US.
At Google I/O, they combined Street View data from Google Maps (collected over the last 20 years, mind you) with Project Genie.
Now, in this world-generation system, you can specify any point on the map and use a prompt to describe how the real world should be transformed. In the example above, a flooded bridge in San Francisco becomes part of an underwater world.
In other words, you can create virtually any "skin" for any real-world location โ and not just in 2D, but in full 3D.
Instead of location scouts, we'll soon have location prompters.
Now that's the kind of metaverse we need.
For now, it only works with locations in the US.
๐289โค97๐ฅ93๐81
Gemini Omni vs Grok vs Seedance
Omni is the one with 10 seconds.
And youโll recognize Seedance by the dynamics and the edit cut at the very end.
Omni is the one with 10 seconds.
And youโll recognize Seedance by the dynamics and the edit cut at the very end.
๐990๐340โค327๐ฅ310
This media is not supported in your browser
VIEW IN TELEGRAM
Ideogram 4 has gone open source!
According to current benchmarks and arena rankings, it's one of the strongest open-source image generators available.
"Ideogram 4 is Ideogram's first open-weights model, trained from scratch. It features a structured JSON prompt format, best-in-class multilingual text rendering, deep language understanding, color palette control, and native 2K generation with aspect ratios up to 6:1."
With just 9.3B parameters, it should run on modest GPUs (Qwen-Image: 20B, FLUX.2 [dev]: 32B).
For the geeks:
โข Flow-matching text-to-image model built on a fully single-stream DiT architecture. โข Text and image tokens are processed together by the same 34-layer transformer. โข Uses Qwen3-VL-8B-Instruct, providing richer visual understanding than CLIP or T5.
Trained on JSON annotations and includes a prompt enhancer and guide.
Some content moderation is built in (Hive AI). The license is non-commercial.
GitHub repo (weights, code, docs, guides): https://github.com/ideogram-oss/ideogram4
According to current benchmarks and arena rankings, it's one of the strongest open-source image generators available.
"Ideogram 4 is Ideogram's first open-weights model, trained from scratch. It features a structured JSON prompt format, best-in-class multilingual text rendering, deep language understanding, color palette control, and native 2K generation with aspect ratios up to 6:1."
With just 9.3B parameters, it should run on modest GPUs (Qwen-Image: 20B, FLUX.2 [dev]: 32B).
For the geeks:
โข Flow-matching text-to-image model built on a fully single-stream DiT architecture. โข Text and image tokens are processed together by the same 34-layer transformer. โข Uses Qwen3-VL-8B-Instruct, providing richer visual understanding than CLIP or T5.
Trained on JSON annotations and includes a prompt enhancer and guide.
Some content moderation is built in (Hive AI). The license is non-commercial.
GitHub repo (weights, code, docs, guides): https://github.com/ideogram-oss/ideogram4
๐57โค55๐ฅ49๐38
Media is too big
VIEW IN TELEGRAM
3D generators are coming for the sacred territory โ organic modeling, specifically head modeling.
This is Rodin 2.5. They offer 12K textures and generation at 10 million polygons.
The input is several head images from different angles โ either generated images or photos.
https://hyper3d.ai/
This is Rodin 2.5. They offer 12K textures and generation at 10 million polygons.
The input is several head images from different angles โ either generated images or photos.
https://hyper3d.ai/
๐78๐78โค75๐ฅ65
Are Game devs in trouble?
Fable 5 is the biggest step up I've felt in our models since Opus 4.5 back in November.
We just one shot GTA 6 and it works. No miskates.
One prompt. 15 minutes. The whole game.
Rockstar took 12 years. Fable took a quarter of an hour.
Mythos was deemed too dangerous to release, this is the nerfed version. And it still does THIS.
Game devs, are we in trouble?
Fable 5 is the biggest step up I've felt in our models since Opus 4.5 back in November.
We just one shot GTA 6 and it works. No miskates.
One prompt. 15 minutes. The whole game.
Rockstar took 12 years. Fable took a quarter of an hour.
Mythos was deemed too dangerous to release, this is the nerfed version. And it still does THIS.
Game devs, are we in trouble?
๐83โค67๐67๐ฅ59
This media is not supported in your browser
VIEW IN TELEGRAM
Krea 2: Procedural Images
Krea now has real-time sliders: intensity, complexity, and movement.
Basically: prompt intensity, noise/details, and object movement within the frame.
This definitely makes the generation process more interactive and speeds up image selection.
It also reminds me of procedural textures. If youโve worked with them, youโll remember all those endless knobs for tweaking patterns, noise, and textures.
If you havenโt, think of the effect preview in Photoshop: you see the result immediately.
I wonder how many more sliders they could add: color grading, 3D depth, typography, layer separation, backgroundโฆ
The idea is to give regular users as many clear and understandable controls as possible โ instead of all that CFG Scale, Sampling Method, and so on โ and keep them inside the generation interface.
Krea now has real-time sliders: intensity, complexity, and movement.
Basically: prompt intensity, noise/details, and object movement within the frame.
This definitely makes the generation process more interactive and speeds up image selection.
It also reminds me of procedural textures. If youโve worked with them, youโll remember all those endless knobs for tweaking patterns, noise, and textures.
If you havenโt, think of the effect preview in Photoshop: you see the result immediately.
I wonder how many more sliders they could add: color grading, 3D depth, typography, layer separation, backgroundโฆ
The idea is to give regular users as many clear and understandable controls as possible โ instead of all that CFG Scale, Sampling Method, and so on โ and keep them inside the generation interface.
๐183๐ฅ181๐172โค164
AI video is getting wild.
On June 15, Dreamina is releasing Seedance 2.0 Mini, and this is huge:
Same performance level as Seedance 2.0, but much cheaper.
That means creators can generate more AI drama shots, test more ideas, and build cinematic videos without burning insane credits.
AI video is not slowing down.
On June 15, Dreamina is releasing Seedance 2.0 Mini, and this is huge:
Same performance level as Seedance 2.0, but much cheaper.
That means creators can generate more AI drama shots, test more ideas, and build cinematic videos without burning insane credits.
AI video is not slowing down.
๐ฅ209โค203๐202๐197
Reverse Ideogram 4
Interesting tool: you give it any image as input, and it outputs a JSON prompt for Ideogram 4, complete with bounding boxes and labels for them.
You can edit specific parts of images locally.
Thereโs code here:
https://github.com/cocktailpeanut/image-to-prompt
Under the hood, it uses Microsoftโs Florence-2 model for image recognition.
And if you donโt want to mess with running it locally, thereโs a Space where you can try it online:
https://huggingface.co/spaces/cocktailpeanut/image-to-prompt
Interesting tool: you give it any image as input, and it outputs a JSON prompt for Ideogram 4, complete with bounding boxes and labels for them.
You can edit specific parts of images locally.
Thereโs code here:
https://github.com/cocktailpeanut/image-to-prompt
Under the hood, it uses Microsoftโs Florence-2 model for image recognition.
And if you donโt want to mess with running it locally, thereโs a Space where you can try it online:
https://huggingface.co/spaces/cocktailpeanut/image-to-prompt
โค603๐587๐578๐ฅ564
Media is too big
VIEW IN TELEGRAM
Grok Imagine Video 1.5
Itโs out of Preview.
The interesting part: 720p, 15 seconds โ but the SuperGrok subscription page says โ30-second videos.โ
Thatโs a marketing trick. The API clearly says 15 seconds. The โ30 secondsโ refers to the Extend from Frame feature, which is available there.
Thereโs also Grok Imagine Video 1.5 Fast โ it generates 720p videos in 25 seconds.
https://x.ai/news/grok-imagine-video-1-5
Itโs out of Preview.
The interesting part: 720p, 15 seconds โ but the SuperGrok subscription page says โ30-second videos.โ
Thatโs a marketing trick. The API clearly says 15 seconds. The โ30 secondsโ refers to the Extend from Frame feature, which is available there.
Thereโs also Grok Imagine Video 1.5 Fast โ it generates 720p videos in 25 seconds.
https://x.ai/news/grok-imagine-video-1-5
โค85๐ฅ73๐70๐70
Who will win the World Cup before it happens?
Soccer Buddy uses AI-style prediction logic, 80+ match parameters, 10,000 simulations, and deep pro stats to find stronger football angles before kickoff.
xG, BTTS, overs, corners, cards, shots, possession, home and away splits.
AI + stats = smarter soccer predictions.
https://zcodesystem.com/soccerbuddy/?WC2026
Soccer Buddy uses AI-style prediction logic, 80+ match parameters, 10,000 simulations, and deep pro stats to find stronger football angles before kickoff.
xG, BTTS, overs, corners, cards, shots, possession, home and away splits.
AI + stats = smarter soccer predictions.
https://zcodesystem.com/soccerbuddy/?WC2026
๐195โค184๐ฅ165๐153
Seedance 2.5
ByteDance announced the new version at Volcano Engine FORCE 2026.
The interesting part: native 30-second video generation in a single run. Previous public versions were limited to 15 seconds, so this is a major jump.
Another big upgrade: up to 50 multimodal references โ images, videos, and audio โ to control style, characters, motion, and editing.
Seedance 2.0 also received an upgrade and now supports native 4K video generation.
The model is currently in enterprise beta and is expected to launch publicly in early July.
ByteDance announced the new version at Volcano Engine FORCE 2026.
The interesting part: native 30-second video generation in a single run. Previous public versions were limited to 15 seconds, so this is a major jump.
Another big upgrade: up to 50 multimodal references โ images, videos, and audio โ to control style, characters, motion, and editing.
Seedance 2.0 also received an upgrade and now supports native 4K video generation.
The model is currently in enterprise beta and is expected to launch publicly in early July.
๐ฅ89โค79๐70๐63
๐ค AI company Anthropic is expanding its presence in Europe
Anthropic, the company behind the Claude AI assistant, has hired the head of artificial intelligence from the French telecom company Orange. This move is part of Anthropicโs major expansion into the European market.
It seems that competition between leading AI companies is no longer just about building the best models โ it is also about attracting the worldโs top experts.
Anthropic, the company behind the Claude AI assistant, has hired the head of artificial intelligence from the French telecom company Orange. This move is part of Anthropicโs major expansion into the European market.
It seems that competition between leading AI companies is no longer just about building the best models โ it is also about attracting the worldโs top experts.
๐228๐ฅ224๐222โค187