🔄 [GitHub Releases] pytorch/pytorch - PyTorch 2.12.1 Release, bug fix release
https://github.com/pytorch/pytorch/releases/tag/v2.12.1
https://github.com/pytorch/pytorch/releases/tag/v2.12.1
GitHub
Release PyTorch 2.12.1 Release, bug fix release · pytorch/pytorch
This release is meant to fix the following regressions and silent correctness issues:
Regression fixes
Fix nondeterministic outputs in test_batch_invariance with FLASH_ATTN on NVIDIA B200 GPUs (#1...
Regression fixes
Fix nondeterministic outputs in test_batch_invariance with FLASH_ATTN on NVIDIA B200 GPUs (#1...
📰 PyTorch - Schedule Now Available for KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China
AI is transforming how we build, deploy, and operate technology. Open source is making it possible. KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China will take place September...
https://pytorch.org/blog/schedule-now-available-for-kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/
AI is transforming how we build, deploy, and operate technology. Open source is making it possible. KubeCon + CloudNativeCon + OpenInfra Summit + PyTorch Conference China will take place September...
https://pytorch.org/blog/schedule-now-available-for-kubecon-cloudnativecon-openinfra-summit-pytorch-conference-china/
📰 LMSys - Optimizing Ling-2.6-1T on TPU with SGLang-JAX: Hiding MoE Data Movement Behind Compute with One Pallas Kernel
https://lmsys.org/blog/2026-06-17-ling-2-6-tpu
https://lmsys.org/blog/2026-06-17-ling-2-6-tpu
www.lmsys.org
Optimizing Ling-2.6-1T on TPU with SGLang-JAX: Hiding MoE Data Movement Behind Compute with One Pallas Kernel
SGLang-JAX now supports efficient serving of inclusionAI's Ling-2.6-1T on TPU v7x. With a working baseline in place, profiling pointed to the Mixture-of-Experts (MoE) path as the main bottleneck: each...
📰 HuggingFace - Is it agentic enough? Benchmarking open models on your own tooling
https://huggingface.co/blog/is-it-agentic-enough
https://huggingface.co/blog/is-it-agentic-enough
huggingface.co
Is it agentic enough? Benchmarking open models on your own tooling
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🆕 [HF Models] LiquidAI - LFM2.5-ColBERT-350M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-ColBERT-350M-GGUF
🆕 [HF Models] LiquidAI - LFM2.5-Embedding-350M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-Embedding-350M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-ColBERT-350M-GGUF
🆕 [HF Models] LiquidAI - LFM2.5-Embedding-350M-GGUF
https://huggingface.co/LiquidAI/LFM2.5-Embedding-350M-GGUF
huggingface.co
LiquidAI/LFM2.5-ColBERT-350M-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 PyTorch - JUST LAUNCHED! PyTorch Certified Associate (PTCA)
Prove You Can Work with the Technology Behind Modern AI Linux Foundation Education and PyTorch Foundation have launched the PyTorch Certified Associate (PTCA), a new certification designed for early-stage practitioners...
https://pytorch.org/blog/just-launched-pytorch-certified-associate-ptca/
Prove You Can Work with the Technology Behind Modern AI Linux Foundation Education and PyTorch Foundation have launched the PyTorch Certified Associate (PTCA), a new certification designed for early-stage practitioners...
https://pytorch.org/blog/just-launched-pytorch-certified-associate-ptca/
📰 HuggingFace - Beyond LoRA: Can you beat the most popular fine-tuning technique?
https://huggingface.co/blog/peft-beyond-lora
https://huggingface.co/blog/peft-beyond-lora
huggingface.co
Beyond LoRA: Can you beat the most popular fine-tuning technique?
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 PyTorch - From Minutes to Seconds: LLM-Guided Autotuning for Helion Kernels
TL;DR Helion, PyTorch’s domain-specific language (DSL) for performance portable machine learning kernels, heavily relies on autotuning for performance. Currently Helion searches utilize the Likelihood-Free Bayesian Optimization (LFBO) to find the...
https://pytorch.org/blog/from-minutes-to-seconds-llm-guided-autotuning-for-helion-kernels/
TL;DR Helion, PyTorch’s domain-specific language (DSL) for performance portable machine learning kernels, heavily relies on autotuning for performance. Currently Helion searches utilize the Likelihood-Free Bayesian Optimization (LFBO) to find the...
https://pytorch.org/blog/from-minutes-to-seconds-llm-guided-autotuning-for-helion-kernels/
📰 HuggingFace - MosaicLeaks: Can your research agent keep a secret?
https://huggingface.co/blog/ServiceNow/mosaicleaks
https://huggingface.co/blog/ServiceNow/mosaicleaks
huggingface.co
MosaicLeaks: Can your research agent keep a secret?
A Blog post by ServiceNow on Hugging Face
📰 Anthropic Research - Project Fetch: Phase two
https://www.anthropic.com/research/project-fetch-phase-two
https://www.anthropic.com/research/project-fetch-phase-two
Anthropic
Project Fetch: Phase two
Results from our latest test of whether Claude can help Anthropic employees perform sophisticated robotics tasks. We found that Claude Opus 4.7, operating without human assistance, was about 20 times faster than the fastest human team at all tasks completed…
🆕 [HF Models] jdopensource - JoyAI-VL-Interaction-Preview
https://huggingface.co/jdopensource/JoyAI-VL-Interaction-Preview
https://huggingface.co/jdopensource/JoyAI-VL-Interaction-Preview
huggingface.co
jdopensource/JoyAI-VL-Interaction-Preview · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Google AI Blog - How A2A is Building a World of Collaborative Agents
Celebrating the first anniversary of the Agent-to-Agent (A2A) protocol, this blog post highlights how the framework enables autonomous AI agents to securely collaborate and hand off tasks without the rigidity of traditional APIs. By delegating complex workflows to specialized peer agents, A2A prevents context pollution, ensures data privacy, and simplifies application design through modularity. To demonstrate this ecosystem in action, the post spotlights FoldRun—an agentic interface for life sciences that orchestrates complex protein structure predictions—alongside diverse A2A use cases spanning commerce, data streaming, DevOps, and telecommunications.
https://developers.googleblog.com/en/how-a2a-is-building-a-world-of-collaborative-agents/
Celebrating the first anniversary of the Agent-to-Agent (A2A) protocol, this blog post highlights how the framework enables autonomous AI agents to securely collaborate and hand off tasks without the rigidity of traditional APIs. By delegating complex workflows to specialized peer agents, A2A prevents context pollution, ensures data privacy, and simplifies application design through modularity. To demonstrate this ecosystem in action, the post spotlights FoldRun—an agentic interface for life sciences that orchestrates complex protein structure predictions—alongside diverse A2A use cases spanning commerce, data streaming, DevOps, and telecommunications.
https://developers.googleblog.com/en/how-a2a-is-building-a-world-of-collaborative-agents/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Discover how the Agent-to-Agent (A2A) protocol is shifting AI from isolated tools to a collaborative ecosystem, enabling secure, autonomous agent handoffs and scalable workflows like FoldRun.
📰 Claude Blog - Steering Claude Code: CLAUDE.md files, skills, hooks, rules, subagents and more
https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more
📰 Claude Blog - Centrally manage authorization for MCP connectors
https://claude.com/blog/enterprise-managed-auth
📰 Claude Blog - Claude Code now supports artifacts
https://claude.com/blog/artifacts-in-claude-code
📰 Claude Blog - Meet the winners of our Claude Opus 4.8 Build Day hackathon
https://claude.com/blog/meet-the-winners-of-our-claude-opus-4-8-build-day-hackathon
📰 Claude Blog - Claude Design now stays on brand for daily work
https://claude.com/blog/claude-design-stays-on-brand-for-daily-work
📰 Claude Blog - Secure access to the Claude Platform with Workload Identity Federation
https://claude.com/blog/workload-identity-federation
https://claude.com/blog/steering-claude-code-skills-hooks-rules-subagents-and-more
📰 Claude Blog - Centrally manage authorization for MCP connectors
https://claude.com/blog/enterprise-managed-auth
📰 Claude Blog - Claude Code now supports artifacts
https://claude.com/blog/artifacts-in-claude-code
📰 Claude Blog - Meet the winners of our Claude Opus 4.8 Build Day hackathon
https://claude.com/blog/meet-the-winners-of-our-claude-opus-4-8-build-day-hackathon
📰 Claude Blog - Claude Design now stays on brand for daily work
https://claude.com/blog/claude-design-stays-on-brand-for-daily-work
📰 Claude Blog - Secure access to the Claude Platform with Workload Identity Federation
https://claude.com/blog/workload-identity-federation
Claude
Steering Claude Code: when to use CLAUDE.md, skills, hooks, and subagents | Claude by Anthropic
Seven ways to steer Claude Code—CLAUDE.md files, rules, skills, subagents, hooks, and more—and when to use each, based on context cost and authority.
🆕 [HF Models] FunAudioLLM - Fun-ASR-Nano-GGUF
https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-GGUF
🆕 [HF Models] FunAudioLLM - Paraformer-GGUF
https://huggingface.co/FunAudioLLM/Paraformer-GGUF
🆕 [HF Models] FunAudioLLM - SenseVoiceSmall-GGUF
https://huggingface.co/FunAudioLLM/SenseVoiceSmall-GGUF
🆕 [HF Models] FunAudioLLM - fsmn-vad-GGUF
https://huggingface.co/FunAudioLLM/fsmn-vad-GGUF
https://huggingface.co/FunAudioLLM/Fun-ASR-Nano-GGUF
🆕 [HF Models] FunAudioLLM - Paraformer-GGUF
https://huggingface.co/FunAudioLLM/Paraformer-GGUF
🆕 [HF Models] FunAudioLLM - SenseVoiceSmall-GGUF
https://huggingface.co/FunAudioLLM/SenseVoiceSmall-GGUF
🆕 [HF Models] FunAudioLLM - fsmn-vad-GGUF
https://huggingface.co/FunAudioLLM/fsmn-vad-GGUF
🗓️ Weekly GitHub Activity
🦙 llama.cpp
└ Release: b9627 → b9743
└ 116 commits
- Added support for Cohere2MoE (North Code / Tiny Aya) and GLM-5.2 models. #24615 #24770
- Integrated Eagle3 speculative decoding support for Qwen 3.5 and 3.6. #24593
- Updated OpenVINO backend to 2026.2 with context-shift, Q5_1 weights, and Gemma 4 support. #24503
- Introduced a model management API to the server router for remote model downloads and deletion. #23976
- UI improvements: added HEIC/HEIF image support, SVG/Mermaid rendering with source toggles, and markdown rendering for thinking blocks. #24137 #24080 #24611
- Optimized AMX performance on CPU and improved i-quants prefill speeds for WebGPU. #24806 #24530
- Enhanced Metal backend with concat support for F16/BF16 and rope_back operator. #24724 #24725
- Fixed significant whitespace issues in chat grammar generation and double-escaping in tool-call parsing. #24624 #24667
- Server now includes real-time generation speed metrics and JSONL conversation exports. #24291 #24688
- SYCL backend updates: added Conv2D/Conv3D support, dev-to-dev memcpy, and set F16 as default. #24600 #24476 #23996
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-694-276025e → master-709-92a3b73
└ 15 commits
- Added RPC support for remote compute execution #1629
- Implemented PuLID-Flux identity-injection support for Flux models #1595
- Added support for cancelling ongoing generations with partial image batch returns #1124
- Introduced disk parameters backend support #1651
- Added backend-specific max-VRAM budgets bb90bfa
- Fixed handling of oversized Vulkan parameter tensors #1662
- Synchronized core library with latest GGML #1656
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
WeiboAI/VibeThinker-3B ♡511
prefeitura-rio/Rio-3.5-Open-397B ♡327
owensong/Inflect-Nano-v1 | gguf ♡140
Zyphra/ZONOS2 ♡118
datalab-to/lift ♡86
poolside/Laguna-M.1 ♡74
Boogu/Boogu-Image-0.1-Edit ♡67
AlexWortega/SIQ-1-35B ♡59
Boogu/Boogu-Image-0.1-Turbo ♡37
SupraLabs/Supra-1.5-50M-Instruct-exp ♡37
Boogu/Boogu-Image-0.1-Base ♡32
HKUSTAudio/AudioX-Turbo ♡29
FINAL-Bench/Darwin-398B-JGOS ♡28
Multilingual-Multimodal-NLP/LoopCoder-V2 ♡26
YTan2000/Qwen3.6-27B-MTP-TQ3_4S ♡16
Danrisi/UltraReal_FineTune_Anima_base1_v3 ♡14
catnip-ai-tech/MaineCoon ♡14
🦙 llama.cpp
└ Release: b9627 → b9743
└ 116 commits
- Added support for Cohere2MoE (North Code / Tiny Aya) and GLM-5.2 models. #24615 #24770
- Integrated Eagle3 speculative decoding support for Qwen 3.5 and 3.6. #24593
- Updated OpenVINO backend to 2026.2 with context-shift, Q5_1 weights, and Gemma 4 support. #24503
- Introduced a model management API to the server router for remote model downloads and deletion. #23976
- UI improvements: added HEIC/HEIF image support, SVG/Mermaid rendering with source toggles, and markdown rendering for thinking blocks. #24137 #24080 #24611
- Optimized AMX performance on CPU and improved i-quants prefill speeds for WebGPU. #24806 #24530
- Enhanced Metal backend with concat support for F16/BF16 and rope_back operator. #24724 #24725
- Fixed significant whitespace issues in chat grammar generation and double-escaping in tool-call parsing. #24624 #24667
- Server now includes real-time generation speed metrics and JSONL conversation exports. #24291 #24688
- SYCL backend updates: added Conv2D/Conv3D support, dev-to-dev memcpy, and set F16 as default. #24600 #24476 #23996
🔗 All changes | Latest release
🎨 stable-diffusion.cpp
└ Release: master-694-276025e → master-709-92a3b73
└ 15 commits
- Added RPC support for remote compute execution #1629
- Implemented PuLID-Flux identity-injection support for Flux models #1595
- Added support for cancelling ongoing generations with partial image batch returns #1124
- Introduced disk parameters backend support #1651
- Added backend-specific max-VRAM budgets bb90bfa
- Fixed handling of oversized Vulkan parameter tensors #1662
- Synchronized core library with latest GGML #1656
🔗 All changes | Latest release
🤗 Fresh models trending on HuggingFace:
WeiboAI/VibeThinker-3B ♡511
prefeitura-rio/Rio-3.5-Open-397B ♡327
owensong/Inflect-Nano-v1 | gguf ♡140
Zyphra/ZONOS2 ♡118
datalab-to/lift ♡86
poolside/Laguna-M.1 ♡74
Boogu/Boogu-Image-0.1-Edit ♡67
AlexWortega/SIQ-1-35B ♡59
Boogu/Boogu-Image-0.1-Turbo ♡37
SupraLabs/Supra-1.5-50M-Instruct-exp ♡37
Boogu/Boogu-Image-0.1-Base ♡32
HKUSTAudio/AudioX-Turbo ♡29
FINAL-Bench/Darwin-398B-JGOS ♡28
Multilingual-Multimodal-NLP/LoopCoder-V2 ♡26
YTan2000/Qwen3.6-27B-MTP-TQ3_4S ♡16
Danrisi/UltraReal_FineTune_Anima_base1_v3 ♡14
catnip-ai-tech/MaineCoon ♡14
GitHub
chat: add dedicated Cohere2MoE (North Code) parser by pwilkin · Pull Request #24615 · ggml-org/llama.cpp
Overview
The Cohere2 MoE template is pretty special, so using the autoparser even with workarounds didn't really work. Needed a dedicated parser.
Additional information
Please use the templ...
The Cohere2 MoE template is pretty special, so using the autoparser even with workarounds didn't really work. Needed a dedicated parser.
Additional information
Please use the templ...
🔓 [HF Models] inclusionAI - Sing-Guard-2b
https://huggingface.co/inclusionAI/Sing-Guard-2b
🔓 [HF Models] inclusionAI - Sing-Guard-8b
https://huggingface.co/inclusionAI/Sing-Guard-8b
🔓 [HF Models] inclusionAI - Sing-Guard-4b
https://huggingface.co/inclusionAI/Sing-Guard-4b
https://huggingface.co/inclusionAI/Sing-Guard-2b
🔓 [HF Models] inclusionAI - Sing-Guard-8b
https://huggingface.co/inclusionAI/Sing-Guard-8b
🔓 [HF Models] inclusionAI - Sing-Guard-4b
https://huggingface.co/inclusionAI/Sing-Guard-4b
huggingface.co
inclusionAI/SingGuard-2b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🆕 [HF Models] jdopensource - JoyAI-Image-Edit-Plus-Diffusers
https://huggingface.co/jdopensource/JoyAI-Image-Edit-Plus-Diffusers
https://huggingface.co/jdopensource/JoyAI-Image-Edit-Plus-Diffusers
huggingface.co
jdopensource/JoyAI-Image-Edit-Plus-Diffusers · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 HuggingFace - PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters
https://huggingface.co/blog/PaddlePaddle/pp-ocrv6
https://huggingface.co/blog/PaddlePaddle/pp-ocrv6
huggingface.co
PP-OCRv6 on Hugging Face: 50-Language OCR from 1.5M to 34.5M Parameters
A Blog post by PaddlePaddle on Hugging Face
🆕 [HF Models] inclusionAI - Sing-Guard-2b-GGUF
https://huggingface.co/inclusionAI/Sing-Guard-2b-GGUF
🆕 [HF Models] inclusionAI - Sing-Guard-4b-GGUF
https://huggingface.co/inclusionAI/Sing-Guard-4b-GGUF
🆕 [HF Models] inclusionAI - Sing-Guard-8b-GGUF
https://huggingface.co/inclusionAI/Sing-Guard-8b-GGUF
https://huggingface.co/inclusionAI/Sing-Guard-2b-GGUF
🆕 [HF Models] inclusionAI - Sing-Guard-4b-GGUF
https://huggingface.co/inclusionAI/Sing-Guard-4b-GGUF
🆕 [HF Models] inclusionAI - Sing-Guard-8b-GGUF
https://huggingface.co/inclusionAI/Sing-Guard-8b-GGUF
huggingface.co
inclusionAI/SingGuard-2b-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.