Media is too big
VIEW IN TELEGRAM
Super Computer by Higgsfield
This is basically the long-awaited โmake it look goodโ button.
Super Computer is an agent that takes your ideas and decides for itself HOW to execute them.
* It chooses which models to use for the task (Seedance, Kling, Veo, Nano Banana, GPT-Image, Soul, and basically anything with an API โ and it has access to pretty much everything).
* It analyzes not only the type of content (ads, fashion, animation, viral videos), but also the CREATOR themselves โ what they do, what kind of content they make, what performs best for them (check the link, itโs honestly a little scary).
* It analyzes social platforms (TikTok, Instagram, YouTube, Meta Ads Library), digs through trends, competitors, and similar videos.
* It can code and search the web: searching, scraping, parsing, analyzing.
* It connects with Notion, Gmail, Google Drive, Figma, Slack, and more.
* It can work inside Telegram as a classic AI agent.
* It can use Soul ID and train models on specific faces automatically.
This is basically the long-awaited โmake it look goodโ button.
Super Computer is an agent that takes your ideas and decides for itself HOW to execute them.
* It chooses which models to use for the task (Seedance, Kling, Veo, Nano Banana, GPT-Image, Soul, and basically anything with an API โ and it has access to pretty much everything).
* It analyzes not only the type of content (ads, fashion, animation, viral videos), but also the CREATOR themselves โ what they do, what kind of content they make, what performs best for them (check the link, itโs honestly a little scary).
* It analyzes social platforms (TikTok, Instagram, YouTube, Meta Ads Library), digs through trends, competitors, and similar videos.
* It can code and search the web: searching, scraping, parsing, analyzing.
* It connects with Notion, Gmail, Google Drive, Figma, Slack, and more.
* It can work inside Telegram as a classic AI agent.
* It can use Soul ID and train models on specific faces automatically.
๐271๐97โค91๐ฅ80
They said you canโt vibe-code a real game with AI in browser.
Three JS hold my beer. Try my free castle builder: Build by day, bitten by night. Play it, roast it, tell me what breaks.
Try here for free here in browser https://vortex.channel/citybuilder/
Soon on Steam.
Three JS hold my beer. Try my free castle builder: Build by day, bitten by night. Play it, roast it, tell me what breaks.
Try here for free here in browser https://vortex.channel/citybuilder/
Soon on Steam.
๐151๐70๐ฅ61โค53
This is WILD.
Seedance 2.1 and Seedance 2.0 Mini are reportedly coming soon.
Seedance 2.1 could bring around 20% better generation quality than Seedance 2.0.
Seedance 2.0 Mini is rumored to be lighter, faster, better than Seedance 2.0 Fast, and only around $0.073/sec.
The real question now:
Can ByteDance actually challenge Veo 4 and Googleโs Omni model?
Seedance 2.1 and Seedance 2.0 Mini are reportedly coming soon.
Seedance 2.1 could bring around 20% better generation quality than Seedance 2.0.
Seedance 2.0 Mini is rumored to be lighter, faster, better than Seedance 2.0 Fast, and only around $0.073/sec.
The real question now:
Can ByteDance actually challenge Veo 4 and Googleโs Omni model?
๐362๐ฅ131๐119โค116
This media is not supported in your browser
VIEW IN TELEGRAM
Google just dropped Gemini Omni ๐ฎ and this is WILD.
Think Nano Banana, but for VIDEO.
One model. Any input. Text, image, clips, ideasโฆ straight into generated video.
Rolling out across Gemini App, Flow and even YouTube. API support coming next.
And Google is not playing small here. This sounds like a full creative engine, not just another text-to-video toy.
The AI video war just went nuclear. Veo, Sora, Seedance, Klingโฆ everyone just got a new boss fight.
But the real question isโฆ will Gemini Omni beat Seedance 2.1, which is also about to drop with a rumored 20% quality boost over Seedance 2.0?
Itโs so over for old video workflows.
Think Nano Banana, but for VIDEO.
One model. Any input. Text, image, clips, ideasโฆ straight into generated video.
Rolling out across Gemini App, Flow and even YouTube. API support coming next.
And Google is not playing small here. This sounds like a full creative engine, not just another text-to-video toy.
The AI video war just went nuclear. Veo, Sora, Seedance, Klingโฆ everyone just got a new boss fight.
But the real question isโฆ will Gemini Omni beat Seedance 2.1, which is also about to drop with a rumored 20% quality boost over Seedance 2.0?
Itโs so over for old video workflows.
๐181๐12โค6๐ฅ6
This media is not supported in your browser
VIEW IN TELEGRAM
Gemini Omni Flash.
You donโt need to spell out every step for it โ itโs enough to formulate the concept of the video.
It independently finds the theory, descriptions of objects and details, visualizes them, and adds text.
Education will never be the same again.
P.S. Directing educational videos is a different kind of storytelling. You donโt need fights and chase scenes. You need concepts. And they already exist within it.
You donโt need to spell out every step for it โ itโs enough to formulate the concept of the video.
It independently finds the theory, descriptions of objects and details, visualizes them, and adds text.
Education will never be the same again.
P.S. Directing educational videos is a different kind of storytelling. You donโt need fights and chase scenes. You need concepts. And they already exist within it.
๐384โค152๐120๐ฅ117
Wow, now this is interesting!
CapCut is partnering with Google.
Soon, users will be able to edit images and videos directly inside the Gemini app using CapCutโs editing tools.
In other words, Gemini is basically getting a full editing timeline.
Whatโs especially interesting is that the announcement says nothing about Flow. Google, of course, loves building overlapping ecosystems and products...
Itโs also notable that Google still seems committed to the traditional content workflow, where editing remains a core step.
Meanwhile, Higgsfield is betting on an agentic approach โ editing happens during generation itself, with AI deciding how the final clip should be assembled.
A very interesting move, especially since Dreamina and CapCut already support Seedance, Nanobanana, and other models. With the Gemini Omni API arriving, Googleโs own models will likely appear there too.
CapCut is partnering with Google.
Soon, users will be able to edit images and videos directly inside the Gemini app using CapCutโs editing tools.
In other words, Gemini is basically getting a full editing timeline.
Whatโs especially interesting is that the announcement says nothing about Flow. Google, of course, loves building overlapping ecosystems and products...
Itโs also notable that Google still seems committed to the traditional content workflow, where editing remains a core step.
Meanwhile, Higgsfield is betting on an agentic approach โ editing happens during generation itself, with AI deciding how the final clip should be assembled.
A very interesting move, especially since Dreamina and CapCut already support Seedance, Nanobanana, and other models. With the Gemini Omni API arriving, Googleโs own models will likely appear there too.
๐838โค31๐21๐ฅ17
This media is not supported in your browser
VIEW IN TELEGRAM
Gemini Omni: translation, lip-syncing, and dubbing with accurate lip matching.
Omni can translate audio even if the prompt doesnโt include either the original or the translated text:
โข matches lip movements for the new language
โข keeps the background music unchanged
โข adjusts the video length when necessary (!)
For example, the Japanese and Spanish phrases during the close-up shot with the cream are longer, so Omni simply extends the clip by a couple of seconds on its own.
Omni can translate audio even if the prompt doesnโt include either the original or the translated text:
โข matches lip movements for the new language
โข keeps the background music unchanged
โข adjusts the video length when necessary (!)
For example, the Japanese and Spanish phrases during the close-up shot with the cream are longer, so Omni simply extends the clip by a couple of seconds on its own.
โค127๐107๐92๐ฅ90
A short prompt guide for Gemini Omni from Google:
Use real-world knowledge. Gemini Omni understands history, science, and culture, so you donโt need to over-explain. Use cultural references, eras, or scientific terms directly.
Control text rendering. Specify typography, placement, animation style, and visual effects like double exposure, synced with the video action.
Direct the camera. Think like a cinematographer: describe camera type, framing, movement, and shot style.
Edit precisely. Ask for targeted changes, such as replacing a background or caption, without rewriting the whole prompt. Omni can preserve the video structure while adjusting movement, emotion, or object interaction.
https://x.com/GoogleAI/status/2059381218660270435
Use real-world knowledge. Gemini Omni understands history, science, and culture, so you donโt need to over-explain. Use cultural references, eras, or scientific terms directly.
Control text rendering. Specify typography, placement, animation style, and visual effects like double exposure, synced with the video action.
Direct the camera. Think like a cinematographer: describe camera type, framing, movement, and shot style.
Edit precisely. Ask for targeted changes, such as replacing a background or caption, without rewriting the whole prompt. Omni can preserve the video structure while adjusting movement, emotion, or object interaction.
https://x.com/GoogleAI/status/2059381218660270435
๐331๐ฅ130๐110โค104
This media is not supported in your browser
VIEW IN TELEGRAM
A metaverse worth having.
At Google I/O, they combined Street View data from Google Maps (collected over the last 20 years, mind you) with Project Genie.
Now, in this world-generation system, you can specify any point on the map and use a prompt to describe how the real world should be transformed. In the example above, a flooded bridge in San Francisco becomes part of an underwater world.
In other words, you can create virtually any "skin" for any real-world location โ and not just in 2D, but in full 3D.
Instead of location scouts, we'll soon have location prompters.
Now that's the kind of metaverse we need.
For now, it only works with locations in the US.
At Google I/O, they combined Street View data from Google Maps (collected over the last 20 years, mind you) with Project Genie.
Now, in this world-generation system, you can specify any point on the map and use a prompt to describe how the real world should be transformed. In the example above, a flooded bridge in San Francisco becomes part of an underwater world.
In other words, you can create virtually any "skin" for any real-world location โ and not just in 2D, but in full 3D.
Instead of location scouts, we'll soon have location prompters.
Now that's the kind of metaverse we need.
For now, it only works with locations in the US.
๐289โค97๐ฅ93๐81
Gemini Omni vs Grok vs Seedance
Omni is the one with 10 seconds.
And youโll recognize Seedance by the dynamics and the edit cut at the very end.
Omni is the one with 10 seconds.
And youโll recognize Seedance by the dynamics and the edit cut at the very end.
๐990๐340โค327๐ฅ310
This media is not supported in your browser
VIEW IN TELEGRAM
Ideogram 4 has gone open source!
According to current benchmarks and arena rankings, it's one of the strongest open-source image generators available.
"Ideogram 4 is Ideogram's first open-weights model, trained from scratch. It features a structured JSON prompt format, best-in-class multilingual text rendering, deep language understanding, color palette control, and native 2K generation with aspect ratios up to 6:1."
With just 9.3B parameters, it should run on modest GPUs (Qwen-Image: 20B, FLUX.2 [dev]: 32B).
For the geeks:
โข Flow-matching text-to-image model built on a fully single-stream DiT architecture. โข Text and image tokens are processed together by the same 34-layer transformer. โข Uses Qwen3-VL-8B-Instruct, providing richer visual understanding than CLIP or T5.
Trained on JSON annotations and includes a prompt enhancer and guide.
Some content moderation is built in (Hive AI). The license is non-commercial.
GitHub repo (weights, code, docs, guides): https://github.com/ideogram-oss/ideogram4
According to current benchmarks and arena rankings, it's one of the strongest open-source image generators available.
"Ideogram 4 is Ideogram's first open-weights model, trained from scratch. It features a structured JSON prompt format, best-in-class multilingual text rendering, deep language understanding, color palette control, and native 2K generation with aspect ratios up to 6:1."
With just 9.3B parameters, it should run on modest GPUs (Qwen-Image: 20B, FLUX.2 [dev]: 32B).
For the geeks:
โข Flow-matching text-to-image model built on a fully single-stream DiT architecture. โข Text and image tokens are processed together by the same 34-layer transformer. โข Uses Qwen3-VL-8B-Instruct, providing richer visual understanding than CLIP or T5.
Trained on JSON annotations and includes a prompt enhancer and guide.
Some content moderation is built in (Hive AI). The license is non-commercial.
GitHub repo (weights, code, docs, guides): https://github.com/ideogram-oss/ideogram4
๐57โค55๐ฅ49๐38
Media is too big
VIEW IN TELEGRAM
3D generators are coming for the sacred territory โ organic modeling, specifically head modeling.
This is Rodin 2.5. They offer 12K textures and generation at 10 million polygons.
The input is several head images from different angles โ either generated images or photos.
https://hyper3d.ai/
This is Rodin 2.5. They offer 12K textures and generation at 10 million polygons.
The input is several head images from different angles โ either generated images or photos.
https://hyper3d.ai/
๐78๐78โค75๐ฅ65
Are Game devs in trouble?
Fable 5 is the biggest step up I've felt in our models since Opus 4.5 back in November.
We just one shot GTA 6 and it works. No miskates.
One prompt. 15 minutes. The whole game.
Rockstar took 12 years. Fable took a quarter of an hour.
Mythos was deemed too dangerous to release, this is the nerfed version. And it still does THIS.
Game devs, are we in trouble?
Fable 5 is the biggest step up I've felt in our models since Opus 4.5 back in November.
We just one shot GTA 6 and it works. No miskates.
One prompt. 15 minutes. The whole game.
Rockstar took 12 years. Fable took a quarter of an hour.
Mythos was deemed too dangerous to release, this is the nerfed version. And it still does THIS.
Game devs, are we in trouble?
๐83โค67๐67๐ฅ59
This media is not supported in your browser
VIEW IN TELEGRAM
Krea 2: Procedural Images
Krea now has real-time sliders: intensity, complexity, and movement.
Basically: prompt intensity, noise/details, and object movement within the frame.
This definitely makes the generation process more interactive and speeds up image selection.
It also reminds me of procedural textures. If youโve worked with them, youโll remember all those endless knobs for tweaking patterns, noise, and textures.
If you havenโt, think of the effect preview in Photoshop: you see the result immediately.
I wonder how many more sliders they could add: color grading, 3D depth, typography, layer separation, backgroundโฆ
The idea is to give regular users as many clear and understandable controls as possible โ instead of all that CFG Scale, Sampling Method, and so on โ and keep them inside the generation interface.
Krea now has real-time sliders: intensity, complexity, and movement.
Basically: prompt intensity, noise/details, and object movement within the frame.
This definitely makes the generation process more interactive and speeds up image selection.
It also reminds me of procedural textures. If youโve worked with them, youโll remember all those endless knobs for tweaking patterns, noise, and textures.
If you havenโt, think of the effect preview in Photoshop: you see the result immediately.
I wonder how many more sliders they could add: color grading, 3D depth, typography, layer separation, backgroundโฆ
The idea is to give regular users as many clear and understandable controls as possible โ instead of all that CFG Scale, Sampling Method, and so on โ and keep them inside the generation interface.
๐183๐ฅ181๐172โค164
AI video is getting wild.
On June 15, Dreamina is releasing Seedance 2.0 Mini, and this is huge:
Same performance level as Seedance 2.0, but much cheaper.
That means creators can generate more AI drama shots, test more ideas, and build cinematic videos without burning insane credits.
AI video is not slowing down.
On June 15, Dreamina is releasing Seedance 2.0 Mini, and this is huge:
Same performance level as Seedance 2.0, but much cheaper.
That means creators can generate more AI drama shots, test more ideas, and build cinematic videos without burning insane credits.
AI video is not slowing down.
๐ฅ209โค203๐202๐197
Reverse Ideogram 4
Interesting tool: you give it any image as input, and it outputs a JSON prompt for Ideogram 4, complete with bounding boxes and labels for them.
You can edit specific parts of images locally.
Thereโs code here:
https://github.com/cocktailpeanut/image-to-prompt
Under the hood, it uses Microsoftโs Florence-2 model for image recognition.
And if you donโt want to mess with running it locally, thereโs a Space where you can try it online:
https://huggingface.co/spaces/cocktailpeanut/image-to-prompt
Interesting tool: you give it any image as input, and it outputs a JSON prompt for Ideogram 4, complete with bounding boxes and labels for them.
You can edit specific parts of images locally.
Thereโs code here:
https://github.com/cocktailpeanut/image-to-prompt
Under the hood, it uses Microsoftโs Florence-2 model for image recognition.
And if you donโt want to mess with running it locally, thereโs a Space where you can try it online:
https://huggingface.co/spaces/cocktailpeanut/image-to-prompt
โค603๐587๐578๐ฅ564