π° PyTorch - PyTorch Conference North America Announces 2026 Keynotes
PyTorch Conference North America will be held in San Jose, California, on October 20β21, 2026. The 2026 keynote speakers are: Mark Collier, Executive Director, PyTorch Foundation Mazin Gilbert, Executive Director,...
https://pytorch.org/blog/pytorch-conference-north-america-announces-2026-keynotes/
PyTorch Conference North America will be held in San Jose, California, on October 20β21, 2026. The 2026 keynote speakers are: Mark Collier, Executive Director, PyTorch Foundation Mazin Gilbert, Executive Director,...
https://pytorch.org/blog/pytorch-conference-north-america-announces-2026-keynotes/
π [GitHub Releases] turboderp-org/exllamav3 - 1.4.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.4.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.4.0
GitHub
Release 1.4.0 Β· turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs - Release 1.4.0 Β· turboderp-org/exllamav3
π [HF Models] ATH-MaaS - Lumen
https://huggingface.co/ATH-MaaS/Lumen
π [HF Models] ATH-MaaS - Lumen_wmt26
https://huggingface.co/ATH-MaaS/Lumen_wmt26
https://huggingface.co/ATH-MaaS/Lumen
π [HF Models] ATH-MaaS - Lumen_wmt26
https://huggingface.co/ATH-MaaS/Lumen_wmt26
huggingface.co
ATH-MaaS/Lumen Β· Hugging Face
Weβre on a journey to advance and democratize artificial intelligence through open source and open science.
π [HF Models] Wan-AI - Wan2.2-Animate-2-14B-Distilled-Diffusers
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B-Distilled-Diffusers
π [HF Models] Wan-AI - Wan2.2-Animate-2-14B-Diffusers
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B-Diffusers
π [HF Models] Wan-AI - Wan2.2-Animate-2-14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B-Distilled-Diffusers
π [HF Models] Wan-AI - Wan2.2-Animate-2-14B-Diffusers
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B-Diffusers
π [HF Models] Wan-AI - Wan2.2-Animate-2-14B
https://huggingface.co/Wan-AI/Wan2.2-Animate-2-14B
huggingface.co
Wan-AI/Wan2.2-Animate-2-14B-Distilled-Diffusers Β· Hugging Face
Weβre on a journey to advance and democratize artificial intelligence through open source and open science.
π° Google DeepMind - See what 5 builders are making with Gemini Omni
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-builders/
https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-builders/
Google
See what 5 builders are making with Gemini Omni
Gemini Omni makes creating videos as easy as having a conversation. Hereβs how five people use it to edit videos and visualize ideas.
π° LMSys - HPC-Ops Γ SGLang: High-Performance Attention, Router GEMM, and MoE Kernels from Tencent Hunyuan
https://lmsys.org/blog/2026-08-07-hpc-ops-sglang
https://lmsys.org/blog/2026-08-07-hpc-ops-sglang
www.lmsys.org
HPC-Ops Γ SGLang: High-Performance Attention, Router GEMM, and MoE Kernels from Tencent Hunyuan
HPC-Ops is an open-source operator library for LLM inference, deployed in Tencent's large-scale production serving. Its core operators, including Dynamic Attention and Fused MoE, play a critical role ...
π° HuggingFace - TutorMoments: Do AI tutors know when to help and when to hold back?
https://huggingface.co/blog/allenai/tutormoments
https://huggingface.co/blog/allenai/tutormoments
huggingface.co
TutorMoments: Do AI tutors know when to help and when to hold back?
A Blog post by Ai2 on Hugging Face
π° OpenAI - Improving GPTβ5.6 Sol in ChatGPTβand expanding access to GPT-5.6 Luna for free users
ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.6 Luna.
https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt
π° OpenAI - Working with the American Psychological Association on youth mental health and AI
OpenAI and the American Psychological Association advance evidence-based guidance, resources, and safeguards for responsible AI use and youth mental health.
https://openai.com/index/openai-and-apa-partner-to-advance-responsible-ai
π° OpenAI - From asking to doing: How the world is putting ChatGPT to work
New OpenAI Signals data shows how people use ChatGPT worldwide, with country-level insights on adoption, usage trends, and evolving behavior.
https://openai.com/index/how-the-world-is-putting-chatgpt-to-work
ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.6 Luna.
https://openai.com/index/improving-gpt-5-6-sol-in-chatgpt
π° OpenAI - Working with the American Psychological Association on youth mental health and AI
OpenAI and the American Psychological Association advance evidence-based guidance, resources, and safeguards for responsible AI use and youth mental health.
https://openai.com/index/openai-and-apa-partner-to-advance-responsible-ai
π° OpenAI - From asking to doing: How the world is putting ChatGPT to work
New OpenAI Signals data shows how people use ChatGPT worldwide, with country-level insights on adoption, usage trends, and evolving behavior.
https://openai.com/index/how-the-world-is-putting-chatgpt-to-work
OpenAI
Improving GPTβ5.6 Sol in ChatGPTβand expanding access to GPT-5.6 Luna for free users
ChatGPT introduces improved GPT-5.6 Sol with better accuracy and consistency, plus expanded access for free users and unlimited everyday chats with GPT-5.6 Luna.
π [GitHub Releases] sgl-project/sglang - v0.5.17
https://github.com/sgl-project/sglang/releases/tag/v0.5.17
https://github.com/sgl-project/sglang/releases/tag/v0.5.17
GitHub
Release v0.5.17 Β· sgl-project/sglang
Highlights
582 PRs from 194 contributors.
Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linea...
582 PRs from 194 contributors.
Kimi K3 day-0 support: A 2.8T-parameter multimodal LatentMoE (896 experts, top-16, routed in a 3584-dim latent space) with a 1M-token context, 69 KDA linea...
π [GitHub Releases] turboderp-org/exllamav3 - 1.4.1
https://github.com/turboderp-org/exllamav3/releases/tag/v1.4.1
https://github.com/turboderp-org/exllamav3/releases/tag/v1.4.1
GitHub
Release 1.4.1 Β· turboderp-org/exllamav3
Add Mistral-4 (under Mistral3ForConditionalGeneration)
Support logit bias
Adopt LLGuidance instead of Formatron (still supported as optional dependency)
Full Changelog: v1.4.0...v1.4.1
Support logit bias
Adopt LLGuidance instead of Formatron (still supported as optional dependency)
Full Changelog: v1.4.0...v1.4.1
π° Claude Blog - Auto mode is now the default in Claude Code for Pro, Max, and Team plans
https://claude.com/blog/auto-mode-default-in-claude-code
π° Claude Blog - Running auto mode in production
https://claude.com/blog/auto-mode-in-production
π° Claude Blog - Millennium and Anthropic are building a digital risk analyst with Claude
https://claude.com/blog/millennium-and-anthropic-are-building-a-digital-risk-analyst-with-claude
π° Claude Blog - Run Claude Code sessions on your own compute
https://claude.com/blog/run-claude-code-sessions-on-your-own-compute
π° Claude Blog - Inference hooks: inline data loss prevention for Claude Enterprise
https://claude.com/blog/claude-enterprise-inference-hooks
π° Claude Blog - A guide to cost visibility and control in Claude
https://claude.com/blog/a-guide-to-cost-visibility-and-control-in-claude
https://claude.com/blog/auto-mode-default-in-claude-code
π° Claude Blog - Running auto mode in production
https://claude.com/blog/auto-mode-in-production
π° Claude Blog - Millennium and Anthropic are building a digital risk analyst with Claude
https://claude.com/blog/millennium-and-anthropic-are-building-a-digital-risk-analyst-with-claude
π° Claude Blog - Run Claude Code sessions on your own compute
https://claude.com/blog/run-claude-code-sessions-on-your-own-compute
π° Claude Blog - Inference hooks: inline data loss prevention for Claude Enterprise
https://claude.com/blog/claude-enterprise-inference-hooks
π° Claude Blog - A guide to cost visibility and control in Claude
https://claude.com/blog/a-guide-to-cost-visibility-and-control-in-claude
Claude
Auto mode is now the default in Claude Code for Pro, Max, and Team plans | Claude by Anthropic
Claude Code will soon run auto mode by default for Pro, Max, and Team plans, enabling longer-running autonomous work, and catching more dangerous commands.
ποΈ Weekly GitHub Activity
π¦ llama.cpp
β Release: b10233 β b10331
β 98 commits
- Added Multi-Token Prediction (MTP) support for DeepSeek V3.2 #26457 and Qwen3-Next #25589
- Added support for Qwen3-TTS #26254 and DeepSeek OCR multi-row batching #26154
- Introduced Docker-based tool isolation in the server #26507
- Added an LRU scheduler to the server router to prevent evicting busy models #26572
- Upgraded the WebUI with filesystem @mentions #26715, slash commands #26716, a working directory picker #26518, and rich contenteditable input #26717
- Implemented the DeepSeek V4 Lightning Indexer on Metal #25893 and SYCL #26568 to boost prompt processing performance
- Fused CUDA kernels for rms_norm + mul + rope to improve inference latency #26767
- Enabled dynamic allocation for split graph inputs to prevent crashes when running wide MoE models #22789
- Extended SYCL oneDNN SDPA to non-FP16 KV caches (Q4_0-Q8_0, FP32) #25874
- Added a --delete-splits option to gguf-split to free up disk space during file merges #26538
- Added an upcoming deprecation notice for changing the default server port from 8080 to 9931 #26508
- Fixed CUDA data-races when reusing shared memory in block_reduce #26385
- Fixed Metal NORM/RMS_NORM calculation errors on row lengths leaving a partial simdgroup #26708
- Fixed gguf-py reader to guard against OOM and integer overflow on crafted files #25401
π All changes | Latest release
π¨ stable-diffusion.cpp
β Release: master-810-db99efd β master-813-bfbef5b
β 3 commits
- Added support for the minimax-h3 model #1854
- Added a trained Minimax VAE Latent2rgb projection #1856
π All changes | Latest release
π΅ audio.cpp
β Release: release-0.5 β release-0.5.1
β 12 commits
- Added a --list-devices CLI flag and list_backend_devices() API to enumerate available backend devices with their hardware names #171
- Added support for Irodori-TTS v4 and aligned its tokenizer #184
- Migrated multiple models to spec v1, including HTDemucs, Mel-Band RoFormer, HeartMuLa, Hviske ASR, SeedVC, Sortformer, and Higgs Audio STT #177 #179
- Enabled Metal ConvTranspose fast path with MioCodec layout-safe multiplies cdd5196
π All changes | Latest release
π¦ llama.cpp
β Release: b10233 β b10331
β 98 commits
- Added Multi-Token Prediction (MTP) support for DeepSeek V3.2 #26457 and Qwen3-Next #25589
- Added support for Qwen3-TTS #26254 and DeepSeek OCR multi-row batching #26154
- Introduced Docker-based tool isolation in the server #26507
- Added an LRU scheduler to the server router to prevent evicting busy models #26572
- Upgraded the WebUI with filesystem @mentions #26715, slash commands #26716, a working directory picker #26518, and rich contenteditable input #26717
- Implemented the DeepSeek V4 Lightning Indexer on Metal #25893 and SYCL #26568 to boost prompt processing performance
- Fused CUDA kernels for rms_norm + mul + rope to improve inference latency #26767
- Enabled dynamic allocation for split graph inputs to prevent crashes when running wide MoE models #22789
- Extended SYCL oneDNN SDPA to non-FP16 KV caches (Q4_0-Q8_0, FP32) #25874
- Added a --delete-splits option to gguf-split to free up disk space during file merges #26538
- Added an upcoming deprecation notice for changing the default server port from 8080 to 9931 #26508
- Fixed CUDA data-races when reusing shared memory in block_reduce #26385
- Fixed Metal NORM/RMS_NORM calculation errors on row lengths leaving a partial simdgroup #26708
- Fixed gguf-py reader to guard against OOM and integer overflow on crafted files #25401
π All changes | Latest release
π¨ stable-diffusion.cpp
β Release: master-810-db99efd β master-813-bfbef5b
β 3 commits
- Added support for the minimax-h3 model #1854
- Added a trained Minimax VAE Latent2rgb projection #1856
π All changes | Latest release
π΅ audio.cpp
β Release: release-0.5 β release-0.5.1
β 12 commits
- Added a --list-devices CLI flag and list_backend_devices() API to enumerate available backend devices with their hardware names #171
- Added support for Irodori-TTS v4 and aligned its tokenizer #184
- Migrated multiple models to spec v1, including HTDemucs, Mel-Band RoFormer, HeartMuLa, Hviske ASR, SeedVC, Sortformer, and Higgs Audio STT #177 #179
- Enabled Metal ConvTranspose fast path with MioCodec layout-safe multiplies cdd5196
π All changes | Latest release
π€ Fresh models trending on HuggingFace:
β‘255 deepgrove/maple-preview
β‘197 lightx2v/Minimax-h3-Turbo
β‘168 PinkCherry_MiniMax-H3 | PinkFluffyBunny-MiniMax-H3
β‘139 Kijai/MiniMax-H3-experimental
β‘97 SyzygyResearch/Mach-1-Additive-35B
β‘94 Kijai/MiniMax-H3-TAE
β‘93 endless-frontier/BigBang-v1
β‘81 Akahsizrr/fuse-1-Lite
β‘75 badtheorylabs/BTL-4 | BTL-4-Compact
β‘49 harrrshall/BarunLM-35M
β‘44 TenStrip/10Eros-Max
β‘42 sand-ai/MAGI-2-preview
β‘40 thedeoxen/Krea-2-pose-controlnet
β‘39 ReadyArt/gemma-4-31B-it-scotoma-2
β‘39 DeepBeepMeep/MiniMax-H3
β‘35 zerofata/G4-MeroMero-v2-31B
β‘32 SupraLabs/Supra2-100M-Instruct | Supra2-100M-Base
β‘29 KBlueLeaf/TIPOv2-1B-A200M
β‘27 ASLP-lab/MeanVC2
β‘26 quimmedes/Deepwen-3.6
β‘19 BananaMind/BananaMind-2-Pro-Preview
β‘19 0xSero/deepseek-v4-flash-0731-spark
β‘255 deepgrove/maple-preview
20B-A1B ternary-weight Mixture-of-Experts LLM optimized for high-speed on-device reasoning and math.
β‘197 lightx2v/Minimax-h3-Turbo
Image-to-video and text-to-video model based on MiniMax-H3 for fast video generation.
β‘168 PinkCherry_MiniMax-H3 | PinkFluffyBunny-MiniMax-H3
Text-to-video model based on MiniMax-H3 specialized in organic motion and detail generation.
β‘139 Kijai/MiniMax-H3-experimental
Experimental 4-bit quantized version of the MiniMax-H3 video model with int8-convrot activation support for ComfyUI.
β‘97 SyzygyResearch/Mach-1-Additive-35B
35B ternary 1.7-bit LLM based on Qwen designed for highly efficient, high-accuracy reasoning.
β‘94 Kijai/MiniMax-H3-TAE
Lightweight VAE model designed for quick video preview generation in MiniMax-H3 workflows.
β‘93 endless-frontier/BigBang-v1
Qwen3.6-35B-A3B fine-tune for synthetic frontier tasks for advanced scientific research, coding, and tool-use, significantly outperforming base model on benchmarks. [paper]
β‘81 Akahsizrr/fuse-1-Lite
5.72B MoE LLM fusing LiquidAI's host model with Qwen3.6 coding experts for lightweight code generation.
β‘75 badtheorylabs/BTL-4 | BTL-4-Compact
35B agentic reasoning LLM fine-tuned for tool use, programming, and long-horizon software engineering tasks.
β‘49 harrrshall/BarunLM-35M
35M parameter decoder-only base SLM featuring hybrid local-global attention for ultra-efficient text generation.
β‘44 TenStrip/10Eros-Max
Video diffusion model based on MiniMax-H3 that blends weight patterns from LTX, Wan, and Krea models to improve aesthetic detail.
β‘42 sand-ai/MAGI-2-preview
114B MoE audio-video generator activating 6B parameters per token to produce synchronized 10-second videos with sound.
β‘40 thedeoxen/Krea-2-pose-controlnet
OpenPose ControlNet LoRA for Krea2 Turbo that controls character posing based on input skeleton maps.
β‘39 ReadyArt/gemma-4-31B-it-scotoma-2
31B LLM fine-tuned from Gemma 4 to reduce stylistic ai tics and cautious behavioral constraints during generation.
β‘39 DeepBeepMeep/MiniMax-H3
Single-file video and image generation models compatible with the WanGP interface for low-VRAM deployment.
β‘35 zerofata/G4-MeroMero-v2-31B
31B Creative narrative & RP tune of Gemma 4.
β‘32 SupraLabs/Supra2-100M-Instruct | Supra2-100M-Base
100M parameter instruct-tuned SLM based on the Qwen3 architecture with a 2K token context window.
β‘29 KBlueLeaf/TIPOv2-1B-A200M
1B MoE sparse prompt expansion model designed to optimize short user prompts for T2I generators.
β‘27 ASLP-lab/MeanVC2
18M parameter zero-shot voice conversion model designed for ultra-low-latency streaming audio.
β‘26 quimmedes/Deepwen-3.6
35B MoE LLM fine-tuned with DeepSeek-style reasoning and specialized skills for 2D and 3D game development and Blender pipelines.
β‘19 BananaMind/BananaMind-2-Pro-Preview
139M parameter base SLM preview checkpoint featuring a digit-aware BPE tokenizer and a 3K token context.
β‘19 0xSero/deepseek-v4-flash-0731-spark
REAP-pruned, quantized MoE LLM optimized for ultra-long context inference on a single DGX Spark.
huggingface.co
deepgrove/maple-preview Β· Hugging Face
Weβre on a journey to advance and democratize artificial intelligence through open source and open science.
π [HF Models] dots-studio - dots.tts-mf-2steps-interleave-tn
https://huggingface.co/dots-studio/dots.tts-mf-2steps-interleave-tn
π [HF Models] dots-studio - dots.tts-mf-2steps
https://huggingface.co/dots-studio/dots.tts-mf-2steps
π [HF Models] dots-studio - dots.tts-mf-1step
https://huggingface.co/dots-studio/dots.tts-mf-1step
π [HF Models] dots-studio - dots.tts-mf-2steps-stts
https://huggingface.co/dots-studio/dots.tts-mf-2steps-stts
https://huggingface.co/dots-studio/dots.tts-mf-2steps-interleave-tn
π [HF Models] dots-studio - dots.tts-mf-2steps
https://huggingface.co/dots-studio/dots.tts-mf-2steps
π [HF Models] dots-studio - dots.tts-mf-1step
https://huggingface.co/dots-studio/dots.tts-mf-1step
π [HF Models] dots-studio - dots.tts-mf-2steps-stts
https://huggingface.co/dots-studio/dots.tts-mf-2steps-stts
huggingface.co
dots-studio/dots.tts-mf-2steps-stts Β· Hugging Face
Weβre on a journey to advance and democratize artificial intelligence through open source and open science.
π° HuggingFace - Making Knowledge Distillation Cheap Enough to Run at Scale
https://huggingface.co/blog/MultiverseComputingCAI/efficient-knowledge-distillation
https://huggingface.co/blog/MultiverseComputingCAI/efficient-knowledge-distillation
huggingface.co
Making Knowledge Distillation Cheap Enough to Run at Scale
A Blog post by Multiverse Computing on Hugging Face
π [HF Models] Motif-Technologies - Motif-3-Base
https://huggingface.co/Motif-Technologies/Motif-3-Base
π [HF Models] Motif-Technologies - Motif-3
https://huggingface.co/Motif-Technologies/Motif-3
https://huggingface.co/Motif-Technologies/Motif-3-Base
π [HF Models] Motif-Technologies - Motif-3
https://huggingface.co/Motif-Technologies/Motif-3