GenAI monitor
552 subscribers
4.4K links
AI frontier model updates & open source LLM releases
Download Telegram
πŸ†• [HF Models] google - gemma-4-12B-it-qat-w4a16-ct

https://huggingface.co/google/gemma-4-12B-it-qat-w4a16-ct


πŸ†• [HF Models] google - gemma-4-12B-it-qat-q4_0-gguf

https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf


πŸ†• [HF Models] google - gemma-4-12B-it-qat-q4_0-unquantized-assistant

https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-unquantized-assistant


πŸ†• [HF Models] google - gemma-4-31B-it-qat-w4a16-ct

https://huggingface.co/google/gemma-4-31B-it-qat-w4a16-ct


πŸ”“ [HF Models] google - gemma-4-E2B-it-qat-q4_0-unquantized-assistant

https://huggingface.co/google/gemma-4-E2B-it-qat-q4_0-unquantized-assistant


πŸ”“ [HF Models] google - gemma-4-E4B-it-qat-q4_0-unquantized-assistant

https://huggingface.co/google/gemma-4-E4B-it-qat-q4_0-unquantized-assistant


πŸ”“ [HF Models] google - gemma-4-26B-A4B-it-qat-q4_0-unquantized-assistant

https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-unquantized-assistant


πŸ”“ [HF Models] google - gemma-4-31B-it-qat-q4_0-unquantized-assistant

https://huggingface.co/google/gemma-4-31B-it-qat-q4_0-unquantized-assistant


πŸ”“ [HF Models] google - gemma-4-E2B-it-qat-mobile-ct

https://huggingface.co/google/gemma-4-E2B-it-qat-mobile-ct


πŸ”“ [HF Models] google - gemma-4-E4B-it-qat-mobile-ct

https://huggingface.co/google/gemma-4-E4B-it-qat-mobile-ct


πŸ”“ [HF Models] google - gemma-4-E2B-it-qat-mobile-transformers

https://huggingface.co/google/gemma-4-E2B-it-qat-mobile-transformers


πŸ”“ [HF Models] google - gemma-4-E4B-it-qat-mobile-transformers

https://huggingface.co/google/gemma-4-E4B-it-qat-mobile-transformers


πŸ”“ [HF Models] google - gemma-4-12B-it-qat-q4_0-unquantized

https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-unquantized


πŸ”“ [HF Models] google - gemma-4-E4B-it-qat-w4a16-ct

https://huggingface.co/google/gemma-4-E4B-it-qat-w4a16-ct
πŸ“° Google AI Blog - Introducing the Google Colab CLI
Google has announced the Google Colab Command-Line Interface (CLI), a new tool that allows developers and AI agents to connect local terminals to remote Colab runtimes for frictionless execution. The lightweight CLI enables users to easily request high-powered GPUs, run local Python scripts remotely, and seamlessly retrieve artifact logs or models like fine-tuned Gemma 3 adapters. By integrating directly into standard terminal environments, the tool is highly programmable and ready to be used by AI agents such as Antigravity or Claude Code to manage complex machine learning pipelines.

https://developers.googleblog.com/en/introducing-the-google-colab-cli/
❀1
πŸ—“οΈ Weekly GitHub Activity


πŸ¦™ llama.cpp
β”” Release: b9437 β†’ b9544
β”” 107 commits

- Added support for EXAONE 4.5 #21733
- Added support for Granite4 Vision #23545
- Added support for Step3.7-Flash #23845
- Added support for Mellum architecture #23966
- Added support for Granite Multilingual Embeddings R2 #22716
- Added support for StepFun 3.5 MTP #23274
- Added tokenizer support for jina-embeddings-v2-base-zh #18756
- Initial support for Qwen3 SSM recurrent architectures #24031
- Server: Real-time reasoning interruption via new control endpoint #23971
- Server: Added placeholder bitmap for token counting and input_tokens API #23913
- Web UI: Thinking mode toggle, reasoning effort levels, and single-line preview #23434, #23601
- Web UI: Mermaid diagrams support and interactive preview #24032
- Tensor Parallel: Quantized KV cache support #23792
- Multimodal: Added frame merge support for Qwen-VL models #21858
- Vulkan: Optimized Q3_K/Q6_K performance on Intel Xe2/BMG via block loads 1962000
- CUDA: Improved MTP performance via mul_mat_vec_q_moe enrollment into PDL #24087
- Hexagon: Major optimizations for MUL_MAT, FLASH_ATTN, and GDN #23989
- Metal: Templated GLU kernels to support f16/f32 #23882
- Web UI: Custom CSS injection via configuration #23904
- KV-cache: SWA checkpoints store only non-masked cells #23981
- Deprecated llama_set_warmup #24009
- Fix model parameters not being propagated correctly to backend #23893
- Server: Avoid unnecessary checkpoint restore when new tokens are present #24110
- Fix session state corruption in common_prompt_batch_decode #23468

πŸ”— All changes | Latest release


🎨 stable-diffusion.cpp
β”” Release: master-660-d2797b8 β†’ master-679-f3fd359
β”” 19 commits

- Added support for Ideogram 4 models #1609
- Added support for Wan2.2 5B FLF2V #1110
- Implemented PiD support #1585
- Added Adaptive Projected Guidance (APG) and unconditional Skip Layer Guidance (SLG) #593
- Added --stream-layers to stream weights from CPU during generation #1576
- Added img-cfg support for edit models #929
- Optimized performance via pinned host buffer allocation and streaming budget management #1601 #1611
- Fixed Flash Attention KV padding issues #1453

πŸ”— All changes | Latest release


πŸ€— Fresh models trending on HuggingFace:

ideogram-ai/ideogram-4-nf4 β™‘212
bosonai/higgs-audio-v3-tts-4b β™‘153
Hcompany/Holo-3.1-4B β™‘57
LiconStudio/LTX-2.3-Multiple-Subject-Reference β™‘50
VAST-AI/TripoSplat β™‘49
nex-agi/Nex-N2-Pro β™‘48
Hcompany/Holo-3.1-35B-A3B β™‘36
SupraLabs/Supra-50M-Reasoning β™‘30
Aratako/Irodori-TTS-600M-v3-VoiceDesign β™‘29
mudler/parakeet-cpp-gguf β™‘28
nex-agi/Nex-N2-mini β™‘22
Trendyol/Trendyol-TTS β™‘22
litert-community/gemma-4-12B-it-litert-lm β™‘20
latam-gpt/Llama-3.1-70B-LatamGPT-SFT-1.0 β™‘20
Soul-AILab/SoulX-Transcriber β™‘17
Hcompany/Holo-3.1-9B β™‘17
Hcompany/Holo-3.1-0.8B β™‘13
ideogram-ai/ideogram-4-nf4-diffusers β™‘13
πŸ“° HuggingFace - Her Β· ΰ€Ήΰ₯‡ΰ€° β€” a detective for your Claude Code sessions


https://huggingface.co/blog/build-small-hackathon/her-blog
πŸ“° HuggingFace - Building Pakistan Notice Helper: A Small AI Tool for a Very Local Safety Problem


https://huggingface.co/blog/build-small-hackathon/building-pakistan-notice-helper