🔄 [GitHub Releases] turboderp-org/exllamav3 - 1.1.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.1.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.1.0
GitHub
Release 1.1.0 · turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs - Release 1.1.0 · turboderp-org/exllamav3
🆕 [HF Models] openbmb - MiniCPM-RobotTrack
https://huggingface.co/openbmb/MiniCPM-RobotTrack
🆕 [HF Models] openbmb - MiniCPM-RobotManip
https://huggingface.co/openbmb/MiniCPM-RobotManip
https://huggingface.co/openbmb/MiniCPM-RobotTrack
🆕 [HF Models] openbmb - MiniCPM-RobotManip
https://huggingface.co/openbmb/MiniCPM-RobotManip
huggingface.co
openbmb/MiniCPM-RobotTrack · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🔄 [GitHub Releases] PygmalionAI/aphrodite-engine - v0.22.0
https://github.com/dphnAI/sonar/releases/tag/v0.22.0
https://github.com/dphnAI/sonar/releases/tag/v0.22.0
GitHub
Release v0.22.0 · dphnAI/sonar
What's Changed
feat: add native Metal support by @AlpinDale in #1668
chore: optimize metal backend performance by @AlpinDale in #1669
perf: optimize GDN performance on Metal by @AlpinDale in #...
feat: add native Metal support by @AlpinDale in #1668
chore: optimize metal backend performance by @AlpinDale in #1669
perf: optimize GDN performance on Metal by @AlpinDale in #...
🗓️ Weekly GitHub Activity
🦙 llama.cpp
└ Release: b9966 → b10068
└ 102 commits
- Added support for Hunyuan 3 (hy_v3) with MTP speculative decoding #25395 #25641
- Added support for Minimax2 Eagle3 speculative decoding 259ae1d
- Added support for BitNetForCausalLM GGUF conversion #25769
- Implemented GGML_OP_LIGHTNING_INDEXER for DeepSeek V3.2/V4 on CPU and CUDA #24231 #25545
- Added fused hyper-connection ops for DeepSeek V4 to reduce graph splits #25585 #25702
- Added CUDA Virtual Devices support and enabled CUDA graphs on Volta and Turing architectures #25228 #25749
- Added Flash Attention via oneDNN graph API for SYCL on Intel Battlemage #25222
- Optimized CUDA MoE gate/up activation quantization, improving prefill times on RTX 5090 and Blackwell #25441
- Added auto-download of DeepSeek-Flash and Eagle3 speculative decoding sidecars from Hugging Face #25811
- Server now supports CORS configuration options and accepts null sampling parameters to request defaults #25655 #25538
- Added KleidiAI SME2 f32 kernel and improved hardware-specific kernel dispatch #24414 #25478
- Fixed CUDA crash when querying memory on devices with no available memory #25157
- Fixed Tensor Parallel execution for Phi3, Bert, Plamo2/3, and ChatGLM #25536
- Fixed quantization crash on DeepSeek-V4 i32 routing tables #25787
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-775-b5d8120 → master-782-b290693
└ 7 commits
- Support for AnimateDiff SD 1.5 motion modules v2 and v3 #1784 with img2video capabilities via the --init-img parameter #1789
- Support for ADetailer #1785
- Support for PiD 1.5 #1790
- Configurable reference image processing for edit models #1780
- Fixed cross attention and output projection token protection for Anima LoRAs #1786
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
thinkingmachines/Inkling ♡1060
OpenMOSS-Team/MOSS-VL-Realtime ♡76
ai-sage/GigaAM-Multilingual ♡56
nineninesix/diamond-1.0 ♡43
ai-sage/GigaChat3.1-Audio-10B-A1.8B ♡38
rzgar/Bernini-R-S2V ♡37
acvlab/ABot-World-0-5B-LF ♡29
fal/ideogram-v4-instant ♡27
InternScience/Agents-A1-4B ♡26
sensenova/SenseNova-U1-8B-MoT-Infographic-V3 ♡26
OpenMOSS-Team/MOSS-VL-Instruct-0708 ♡23
fal/ideogram-v4-fast ♡23
t-tech/T-Search ♡23
GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking ♡23
mente-ai/uyu-2-28B ♡20
OpenMOSS-Team/MOSS-VL-Base-0708 ♡17
yijunwang2/krea2-outpaint ♡17
yijunwang2/krea2-reid ♡15
🦙 llama.cpp
└ Release: b9966 → b10068
└ 102 commits
- Added support for Hunyuan 3 (hy_v3) with MTP speculative decoding #25395 #25641
- Added support for Minimax2 Eagle3 speculative decoding 259ae1d
- Added support for BitNetForCausalLM GGUF conversion #25769
- Implemented GGML_OP_LIGHTNING_INDEXER for DeepSeek V3.2/V4 on CPU and CUDA #24231 #25545
- Added fused hyper-connection ops for DeepSeek V4 to reduce graph splits #25585 #25702
- Added CUDA Virtual Devices support and enabled CUDA graphs on Volta and Turing architectures #25228 #25749
- Added Flash Attention via oneDNN graph API for SYCL on Intel Battlemage #25222
- Optimized CUDA MoE gate/up activation quantization, improving prefill times on RTX 5090 and Blackwell #25441
- Added auto-download of DeepSeek-Flash and Eagle3 speculative decoding sidecars from Hugging Face #25811
- Server now supports CORS configuration options and accepts null sampling parameters to request defaults #25655 #25538
- Added KleidiAI SME2 f32 kernel and improved hardware-specific kernel dispatch #24414 #25478
- Fixed CUDA crash when querying memory on devices with no available memory #25157
- Fixed Tensor Parallel execution for Phi3, Bert, Plamo2/3, and ChatGLM #25536
- Fixed quantization crash on DeepSeek-V4 i32 routing tables #25787
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-775-b5d8120 → master-782-b290693
└ 7 commits
- Support for AnimateDiff SD 1.5 motion modules v2 and v3 #1784 with img2video capabilities via the --init-img parameter #1789
- Support for ADetailer #1785
- Support for PiD 1.5 #1790
- Configurable reference image processing for edit models #1780
- Fixed cross attention and output projection token protection for Anima LoRAs #1786
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
thinkingmachines/Inkling ♡1060
OpenMOSS-Team/MOSS-VL-Realtime ♡76
ai-sage/GigaAM-Multilingual ♡56
nineninesix/diamond-1.0 ♡43
ai-sage/GigaChat3.1-Audio-10B-A1.8B ♡38
rzgar/Bernini-R-S2V ♡37
acvlab/ABot-World-0-5B-LF ♡29
fal/ideogram-v4-instant ♡27
InternScience/Agents-A1-4B ♡26
sensenova/SenseNova-U1-8B-MoT-Infographic-V3 ♡26
OpenMOSS-Team/MOSS-VL-Instruct-0708 ♡23
fal/ideogram-v4-fast ♡23
t-tech/T-Search ♡23
GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking ♡23
mente-ai/uyu-2-28B ♡20
OpenMOSS-Team/MOSS-VL-Base-0708 ♡17
yijunwang2/krea2-outpaint ♡17
yijunwang2/krea2-reid ♡15
GitHub
model: add Hy3 (hy_v3) support with MTP speculative decoding by satindergrewal · Pull Request #25395 · ggml-org/llama.cpp
Overview
Adds support for Tencent's Hy3 (hy_v3 / HYV3ForCausalLM, 299B MoE, 80 layers + 1 MTP layer), including its multi-token-prediction head as a draft-mtp speculative target. Addresses ...
Adds support for Tencent's Hy3 (hy_v3 / HYV3ForCausalLM, 299B MoE, 80 layers + 1 MTP layer), including its multi-token-prediction head as a draft-mtp speculative target. Addresses ...
🔓 [HF Models] nvidia - Cosmos3-Edge
https://huggingface.co/nvidia/Cosmos3-Edge
🔓 [HF Models] nvidia - Cosmos3-Super-Text2Image-4Step
https://huggingface.co/nvidia/Cosmos3-Super-Text2Image-4Step
🔓 [HF Models] nvidia - Cosmos3-Super-Image2Video-4Step
https://huggingface.co/nvidia/Cosmos3-Super-Image2Video-4Step
🔓 [HF Models] nvidia - Cosmos3-Edge-Policy-DROID
https://huggingface.co/nvidia/Cosmos3-Edge-Policy-DROID
https://huggingface.co/nvidia/Cosmos3-Edge
🔓 [HF Models] nvidia - Cosmos3-Super-Text2Image-4Step
https://huggingface.co/nvidia/Cosmos3-Super-Text2Image-4Step
🔓 [HF Models] nvidia - Cosmos3-Super-Image2Video-4Step
https://huggingface.co/nvidia/Cosmos3-Super-Image2Video-4Step
🔓 [HF Models] nvidia - Cosmos3-Edge-Policy-DROID
https://huggingface.co/nvidia/Cosmos3-Edge-Policy-DROID
huggingface.co
nvidia/Cosmos3-Edge · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Google AI Blog - Run Ray on TPU, Part 1: The foundations
Ray 2.55 introduces official, first-class support for Google Cloud TPUs, enabling developers to run distributed Python workloads on Google's accelerators using the familiar Ray task-and-actor APIs. To handle the strict networking requirement of keeping multi-host TPU "slices" together over their Inter-Chip Interconnect (ICI), the KubeRay Operator on GKE automatically provisions and labels the underlying hardware layout. Ray Core utilizes these labels via its slice_placement_group() primitive to atomically reserve complete slices, allowing developers to deploy jobs through KubeRay, Ray Train, or Ray Serve simply by declaring a hardware topology (like "4x4") without writing custom placement code.
https://developers.googleblog.com/en/run-ray-on-tpu-part-1-the-foundations/
Ray 2.55 introduces official, first-class support for Google Cloud TPUs, enabling developers to run distributed Python workloads on Google's accelerators using the familiar Ray task-and-actor APIs. To handle the strict networking requirement of keeping multi-host TPU "slices" together over their Inter-Chip Interconnect (ICI), the KubeRay Operator on GKE automatically provisions and labels the underlying hardware layout. Ray Core utilizes these labels via its slice_placement_group() primitive to atomically reserve complete slices, allowing developers to deploy jobs through KubeRay, Ray Train, or Ray Serve simply by declaring a hardware topology (like "4x4") without writing custom placement code.
https://developers.googleblog.com/en/run-ray-on-tpu-part-1-the-foundations/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Ray 2.55 introduces official, first-class support for Google Cloud TPUs, enabling developers to run distributed Python workloads on Google's accelerators using the familiar Ray task-and-actor APIs. To handle the strict networking requirement of keeping multi…
📰 OpenAI - Safety and alignment in an era of long-horizon models
https://openai.com/index/safety-alignment-long-horizon-models
https://openai.com/index/safety-alignment-long-horizon-models
OpenAI
Safety and alignment in an era of long-horizon models
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
📰 HuggingFace - Grabette: an open system to record robot-manipulation data
https://huggingface.co/blog/grabette
https://huggingface.co/blog/grabette
huggingface.co
Grabette: an open system to record robot-manipulation data
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Google Model Cards - Gemini 3.6 Flash
https://deepmind.google/models/model-cards/gemini-3-6-flash/
📰 Google Model Cards - Gemini 3.5 Flash-Lite
https://deepmind.google/models/model-cards/gemini-3-5-flash-lite/
https://deepmind.google/models/model-cards/gemini-3-6-flash/
📰 Google Model Cards - Gemini 3.5 Flash-Lite
https://deepmind.google/models/model-cards/gemini-3-5-flash-lite/
Google DeepMind
Gemini 3.6 Flash - Model Card
🆕 [HF Models] nvidia - NV-JEPA-DNA-HyenaDNA
https://huggingface.co/nvidia/NV-JEPA-DNA-HyenaDNA
🆕 [HF Models] nvidia - NV-JEPA-DNA-NTv3
https://huggingface.co/nvidia/NV-JEPA-DNA-NTv3
🆕 [HF Models] nvidia - NV-JEPA-DNA-DNABERT2
https://huggingface.co/nvidia/NV-JEPA-DNA-DNABERT2
https://huggingface.co/nvidia/NV-JEPA-DNA-HyenaDNA
🆕 [HF Models] nvidia - NV-JEPA-DNA-NTv3
https://huggingface.co/nvidia/NV-JEPA-DNA-NTv3
🆕 [HF Models] nvidia - NV-JEPA-DNA-DNABERT2
https://huggingface.co/nvidia/NV-JEPA-DNA-DNABERT2
huggingface.co
nvidia/NV-JEPA-DNA-HyenaDNA · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Google Gemma Blog - Scaling Agentic RL: High-Throughput Agentic Training with Tunix
Tunix is Google’s new JAX-native post-training library designed to eliminate TPU idling bottlenecks when training multi-turn, tool-using LLM reasoning agents. It maximizes hardware throughput by combining highly concurrent, asynchronous rollouts with a decoupled producer-consumer pipeline, ensuring the trainer is constantly fed even while agents wait on network I/O or environment steps. Additionally, Tunix provides plug-and-play abstractions and continuous macro-level profiling, allowing developers to easily integrate custom open-source environments and optimize complex distributed workflows without massive code rewrites.
https://developers.googleblog.com/en/scaling-agentic-rl-high-throughput-agentic-training-with-tunix/
Tunix is Google’s new JAX-native post-training library designed to eliminate TPU idling bottlenecks when training multi-turn, tool-using LLM reasoning agents. It maximizes hardware throughput by combining highly concurrent, asynchronous rollouts with a decoupled producer-consumer pipeline, ensuring the trainer is constantly fed even while agents wait on network I/O or environment steps. Additionally, Tunix provides plug-and-play abstractions and continuous macro-level profiling, allowing developers to easily integrate custom open-source environments and optimize complex distributed workflows without massive code rewrites.
https://developers.googleblog.com/en/scaling-agentic-rl-high-throughput-agentic-training-with-tunix/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Optimize agentic RL training with Tunix, Google’s JAX-native library. Eliminate TPU idling with async rollouts and easily plug in custom OSS environments.
📰 Claude Blog - How Anthropic secures its AI-native software development lifecycle
https://claude.com/blog/how-anthropic-secures-its-ai-native-software-development-lifecycle
📰 Claude Blog - How Datadog built a “universal machine tool” for Claude Code
https://claude.com/blog/how-datadog-built-a-universal-machine-tool-for-claude-code
📰 Claude Blog - Working at the frontier: How Rakuten builds agents overnight with Claude Fable 5
https://claude.com/blog/working-at-the-frontier-rakuten
https://claude.com/blog/how-anthropic-secures-its-ai-native-software-development-lifecycle
📰 Claude Blog - How Datadog built a “universal machine tool” for Claude Code
https://claude.com/blog/how-datadog-built-a-universal-machine-tool-for-claude-code
📰 Claude Blog - Working at the frontier: How Rakuten builds agents overnight with Claude Fable 5
https://claude.com/blog/working-at-the-frontier-rakuten
Claude
How Anthropic secures its AI-native software development lifecycle | Claude by Anthropic
Anthropic Deputy CISO Jason Clinton details how the Security Engineering team secures an AI-native SDLC where AI authors 80% of merged code.
📰 HuggingFace - The State of Simulation for Physical AI: An Overview
https://huggingface.co/blog/nvidia/state-of-simulation-for-physical-ai
https://huggingface.co/blog/nvidia/state-of-simulation-for-physical-ai
huggingface.co
The State of Simulation for Physical AI: An Overview
A Blog post by NVIDIA on Hugging Face
❤1
📰 NVIDIA - Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI
What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.
https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/
📰 NVIDIA - NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI
Agentic AI shifts more of the critical execution path onto the CPU. Agents operate in sandboxes to execute code, invoke tools, retrieve context…
https://developer.nvidia.com/blog/inside-nvidia-vera-cpu-olympus-cores-built-for-maximum-single-threaded-performance-in-agentic-ai/
📰 NVIDIA - Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token…
https://developer.nvidia.com/blog/setting-a-world-record-for-moe-pre-training-on-nvidia-gb300-nvl72/
📰 NVIDIA - NVIDIA NVLink: The Scale-Up Network for AI Factories
The demand for AI continues to accelerate. Workloads are getting larger, models are becoming more complex, and there is mounting pressure to deploy AI compute…
https://developer.nvidia.com/blog/nvidia-nvlink-the-scale-up-network-for-ai-factories/
What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale.
https://developer.nvidia.com/blog/inside-nvidia-rubin-gpu-architecture-powering-the-era-of-agentic-ai/
📰 NVIDIA - NVIDIA Vera CPU: Olympus Cores Built for Maximum Single-Thread Performance in Agentic AI
Agentic AI shifts more of the critical execution path onto the CPU. Agents operate in sandboxes to execute code, invoke tools, retrieve context…
https://developer.nvidia.com/blog/inside-nvidia-vera-cpu-olympus-cores-built-for-maximum-single-threaded-performance-in-agentic-ai/
📰 NVIDIA - Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
Frontier model pre-training has converged on mixture of experts (MoE), which is fundamentally changing what limits large-scale AI training. As compute per token…
https://developer.nvidia.com/blog/setting-a-world-record-for-moe-pre-training-on-nvidia-gb300-nvl72/
📰 NVIDIA - NVIDIA NVLink: The Scale-Up Network for AI Factories
The demand for AI continues to accelerate. Workloads are getting larger, models are becoming more complex, and there is mounting pressure to deploy AI compute…
https://developer.nvidia.com/blog/nvidia-nvlink-the-scale-up-network-for-ai-factories/
NVIDIA Technical Blog
Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI
What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale. These factories are now tasked with…
📰 Qwen Research - Qwen-Image-3.0: Rich Content, Authentic Details, Deep Knowledge
https://qwen.ai/blog?id=qwen-image-3.0
https://qwen.ai/blog?id=qwen-image-3.0
qwen.ai
Qwen Studio
Qwen Studio offers comprehensive functionality spanning chatbot, image and video understanding, image generation, document processing, web search integration, tool utilization, and artifacts.
🔄 [GitHub Releases] invoke-ai/InvokeAI - v6.13.7
https://github.com/invoke-ai/InvokeAI/releases/tag/v6.13.7
https://github.com/invoke-ai/InvokeAI/releases/tag/v6.13.7
GitHub
Release v6.13.7 · invoke-ai/InvokeAI
InvokeAI Version 6.13.7
⚠️ This is a security patch release
All recent versions of InvokeAI through 6.13.6 contain a security hole that potentially allows an attacker to recover the contents of the...
⚠️ This is a security patch release
All recent versions of InvokeAI through 6.13.6 contain a security hole that potentially allows an attacker to recover the contents of the...
❤1