GigaChat 3.5 Ultra Publicly Released — The New Generation of the Flagship Model
What’s inside:
🔘A proprietary hybrid MLA + Gated DeltaNet architecture with a dedicated stabilization framework, without which this hybrid setup would not train reliably at this scale;
🔘 Gated Attention: the model can locally down-weight overly strong signals from the attention layer;
🔘GatedNorm: normalization with an explicit gate that controls signal magnitude across features;
🔘Approximately 4x lower KV cache per token: with the same memory budget, the model can support 2.14x longer context and deliver a 20% throughput increase under load;
🔘Two MTP heads, enabling up to 2.2x faster generation;
🔘FP8 across all training stages with no quality degradation compared with bf16, enabled by custom Triton and CUDA kernels;
🔘A new online RL stage after SFT and DPO.
Results:
🔘 GigaChat-3.5-Ultra-Base outperforms DeepSeek V3.2 Exp Base and DeepSeek V4 Flash Base on average across a set of general, math, and code benchmarks:
🔘 GigaChat-3.5-Ultra-Instruct is comparable to DeepSeek V3.2 in terms of average score, despite having half the size;
🔘 According to the MiniMax-M2.7 LLM judge, the average win rate against GigaChat 3.1 Ultra is 75.9%, and against GPT-5 is 68.7%.
➡️ HuggingFace
The GigaChat team has released GigaChat 3.5 Ultra as open source—a new 432B model under the MIT license. This is the first open-source hybrid of GatedDeltaNet and MLA scaled to hundreds of billions of parameters, featuring a proprietary training recipe we refined through more than 1,500 experiments. The model has grown in terms of code, mathematics, agent scenarios, and application domains—yet it’s 40% smaller than GigaChat 3.1 Ultra.
What’s inside:
🔘A proprietary hybrid MLA + Gated DeltaNet architecture with a dedicated stabilization framework, without which this hybrid setup would not train reliably at this scale;
🔘 Gated Attention: the model can locally down-weight overly strong signals from the attention layer;
🔘GatedNorm: normalization with an explicit gate that controls signal magnitude across features;
🔘Approximately 4x lower KV cache per token: with the same memory budget, the model can support 2.14x longer context and deliver a 20% throughput increase under load;
🔘Two MTP heads, enabling up to 2.2x faster generation;
🔘FP8 across all training stages with no quality degradation compared with bf16, enabled by custom Triton and CUDA kernels;
🔘A new online RL stage after SFT and DPO.
Results:
🔘 GigaChat-3.5-Ultra-Base outperforms DeepSeek V3.2 Exp Base and DeepSeek V4 Flash Base on average across a set of general, math, and code benchmarks:
🔘 GigaChat-3.5-Ultra-Instruct is comparable to DeepSeek V3.2 in terms of average score, despite having half the size;
🔘 According to the MiniMax-M2.7 LLM judge, the average win rate against GigaChat 3.1 Ultra is 75.9%, and against GPT-5 is 68.7%.
The entire stack — data (our own LLM-filtered Common Crawl, 600+ programming languages in the code), architecture, training methodology, and infrastructure — was built end-to-end by GigaChat team.
➡️ HuggingFace
❤33🔥5
Forwarded from How AI Helps
Six AI coding agents took one visual IQ test, and Codex 5.5 won by method, speed, and cost
One small test asked agents to solve 25 visual puzzles on iq-test.cc, select age 30, and return a result link.
This was not a lab benchmark. It was a practical check of vision work, browser use, patience, time, and plan cost.
The score is only part of the story. Codex 5.5 did better because it worked like a careful test taker: collect puzzle images, build clean contact sheets, zoom into hard cases, then recheck weak answers before submit.
Claude was careful, especially Opus. It wrote notes and reasoned step by step. Codex was more organized and faster. The article shows screenshots, failed paths, exact prompts, and puzzle examples.
The most useful lesson: for visual web tasks, method can beat size. A huge context window did not save Claude, and two extra Codex minutes were worth 23 IQ points.
read details on our website
Please support this young channel by subscribing.
Your subscription really helps us grow.
There are no ads here.
One small test asked agents to solve 25 visual puzzles on iq-test.cc, select age 30, and return a result link.
This was not a lab benchmark. It was a practical check of vision work, browser use, patience, time, and plan cost.
"Take the IQ test on iq-test.cc. When you finish, select age 30 and send me the link to your result."
Agent IQ Time Limit spent
Claude Cowork Opus 4.8 90 85m ~10 pts
Claude Code Opus 4.8 90 96m ~28 pts
Claude Sonnet 4.6 68 62m n/a
Codex 5.5 $100 Fast 124 18m ~12 pts
Codex 5.4 $100 Fast 101 16m ~14 pts
Codex 5.5 $200 Fast 131 34m ~6 pts
The score is only part of the story. Codex 5.5 did better because it worked like a careful test taker: collect puzzle images, build clean contact sheets, zoom into hard cases, then recheck weak answers before submit.
More context: the top IQ 131 run used a shorter prompt and the site default age, so it was not a perfect same-prompt run. Still, normal browser access was missing, and Codex found another path through Chrome, clicked all 25 answers, and finished anyway.
Claude was careful, especially Opus. It wrote notes and reasoned step by step. Codex was more organized and faster. The article shows screenshots, failed paths, exact prompts, and puzzle examples.
The most useful lesson: for visual web tasks, method can beat size. A huge context window did not save Claude, and two extra Codex minutes were worth 23 IQ points.
read details on our website
Please support this young channel by subscribing.
Your subscription really helps us grow.
There are no ads here.
❤60👍14🔥5
Seedream 5.0 Pro excels at generating photorealistic images and high-density infographics, significantly outperforming the previous Seedream 4.5 and 5.0 Lite models.
In the rankings, Seedream 5.0 Pro matches Nano Banana 2 and Nano Banana Pro in generation quality, while offering much less restrictive censorship.
Please open Telegram to view this post
VIEW IN TELEGRAM
Please open Telegram to view this post
VIEW IN TELEGRAM
❤49
This media is not supported in your browser
VIEW IN TELEGRAM
🚀 GigaChat 3.5 Reasoning — a new open-source LLM that thinks before it answers.
It breaks problems into stages, builds a plan, checks intermediate results, and self-corrects. Built on GigaChat 3.5 Ultra, it explores multiple step-by-step reasoning paths for math & coding, using automated verification to reinforce correct answers.
⚡️ Proprietary linear attention makes it highly efficient on long contexts, retaining key points without re-matching from scratch. It’s also token-efficient: uses 37% fewer tokens than DeepSeek V4 Flash Preview on math problems!
📈 Benchmark gains over non-reasoning version:
• IFBench: 44 → 77
• Natural Plan: 64 → 80
• LiveCodeBench v6: 56 → 85
📦 MIT license. Weights on Hugging Face: fp8 | bf16
It breaks problems into stages, builds a plan, checks intermediate results, and self-corrects. Built on GigaChat 3.5 Ultra, it explores multiple step-by-step reasoning paths for math & coding, using automated verification to reinforce correct answers.
⚡️ Proprietary linear attention makes it highly efficient on long contexts, retaining key points without re-matching from scratch. It’s also token-efficient: uses 37% fewer tokens than DeepSeek V4 Flash Preview on math problems!
📈 Benchmark gains over non-reasoning version:
• IFBench: 44 → 77
• Natural Plan: 64 → 80
• LiveCodeBench v6: 56 → 85
📦 MIT license. Weights on Hugging Face: fp8 | bf16
❤34
🔐 Send files without giving up your privacy
CortexDrop encrypts your files right in the browser before upload, then stores them on decentralized IPFS.
✅ No account needed
✅ Free up to 1 GB
✅ Self-destructing links
✅ Password protection
Drop a file, get a private link in seconds 👇
Try CortexDrop free
CortexDrop encrypts your files right in the browser before upload, then stores them on decentralized IPFS.
✅ No account needed
✅ Free up to 1 GB
✅ Self-destructing links
✅ Password protection
Drop a file, get a private link in seconds 👇
Try CortexDrop free
❤15🔥3👍1
🎬 Kandinsky 6.0 Video — create videos with sound in Full HD!
Sber's new neural network generates clips with synchronized speech, music, and ambient sounds. No special skills needed.
🔥 Key features:
• Videos up to 5 seconds with audio
• Lip-sync: character's lips move with speech
• Sound quality: 44 kHz (streaming standard)
• Realistic physics of movement
• Resolutions: SD, HD, Full HD
📊 Performance:
• 71% better than Kandinsky 5.0
• Beats Veo 3.1 Fast in image animation
• Outperforms LTX 2.5
🛠 For developers:
Open-source under MIT license. Available on FAL, Diffusers, FastVideo, ComfyUI, SGLang, vLLM-omni.
Anton Frolov, Sber: "Anyone can build their own services on top of it — free of charge and without restrictions."
🔗 Hugging Face
Sber's new neural network generates clips with synchronized speech, music, and ambient sounds. No special skills needed.
🔥 Key features:
• Videos up to 5 seconds with audio
• Lip-sync: character's lips move with speech
• Sound quality: 44 kHz (streaming standard)
• Realistic physics of movement
• Resolutions: SD, HD, Full HD
📊 Performance:
• 71% better than Kandinsky 5.0
• Beats Veo 3.1 Fast in image animation
• Outperforms LTX 2.5
🛠 For developers:
Open-source under MIT license. Available on FAL, Diffusers, FastVideo, ComfyUI, SGLang, vLLM-omni.
Anton Frolov, Sber: "Anyone can build their own services on top of it — free of charge and without restrictions."
🔗 Hugging Face
❤11