π [GitHub Releases] turboderp-org/exllamav3 - 0.0.42
https://github.com/turboderp-org/exllamav3/releases/tag/v0.0.42
https://github.com/turboderp-org/exllamav3/releases/tag/v0.0.42
GitHub
Release 0.0.42 Β· turboderp-org/exllamav3
Fix MTP drafting when MTP model is not on the target model's output device
Full Changelog: v0.0.41...v0.0.42
Full Changelog: v0.0.41...v0.0.42
π [GitHub Releases] vllm-project/vllm - v0.23.0
https://github.com/vllm-project/vllm/releases/tag/v0.23.0
https://github.com/vllm-project/vllm/releases/tag/v0.23.0
GitHub
Release v0.23.0 Β· vllm-project/vllm
vLLM v0.23.0 Release Notes
Please note that Minimax M3 is not yet supported in this version. Please follow vLLM recipe for usage guides for M3.
Highlights
This release features 408 commits from 200...
Please note that Minimax M3 is not yet supported in this version. Please follow vLLM recipe for usage guides for M3.
Highlights
This release features 408 commits from 200...
π [GitHub Releases] sgl-project/sglang - v0.5.13
https://github.com/sgl-project/sglang/releases/tag/v0.5.13
https://github.com/sgl-project/sglang/releases/tag/v0.5.13
GitHub
Release v0.5.13 Β· sgl-project/sglang
Highlights
New Model Support:
Autoregressive: Nemotron 3 Ultra (Day-0, blog), Step-3.7-Flash, Command A+
Diffusion: Cosmos3, LingBot-World, SANA-WM, Ernie-Image, FLUX.2-Klein 4B/9B, Ideogram 4
Sp...
New Model Support:
Autoregressive: Nemotron 3 Ultra (Day-0, blog), Step-3.7-Flash, Command A+
Diffusion: Cosmos3, LingBot-World, SANA-WM, Ernie-Image, FLUX.2-Klein 4B/9B, Ideogram 4
Sp...
π° Anthropic - Statement on the US government directive to suspend access to Fable 5 and Mythos 5
https://www.anthropic.com/news/fable-mythos-access
https://www.anthropic.com/news/fable-mythos-access
Anthropic
Statement on the US government directive to suspend access to Fable 5 and Mythos 5
The US government has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States.
π€·ββ2π€―1
π [HF Models] inclusionAI - VISTA-9B
https://huggingface.co/inclusionAI/VISTA-9B
π [HF Models] inclusionAI - VISTA-4B
https://huggingface.co/inclusionAI/VISTA-4B
https://huggingface.co/inclusionAI/VISTA-9B
π [HF Models] inclusionAI - VISTA-4B
https://huggingface.co/inclusionAI/VISTA-4B
huggingface.co
inclusionAI/VISTA-9B Β· Hugging Face
Weβre on a journey to advance and democratize artificial intelligence through open source and open science.
ποΈ Weekly GitHub Activity
π¦ llama.cpp
β Release: b9544 β b9627
β 83 commits
- Support for EAGLE3 speculative decoding #18039
- Added Gemma 4 Multi-Token Prediction (MTP) and assistant draft-model support #23398 #24282
- New architecture support for Cohere2-MoE #24260
- Multi-Token Multi-Domain (MTMD) models now support video input and a batching API #24269 #24384
- WebUI implemented as a Progressive Web App (PWA) with offline caching #23871
- WebUI added an opt-in sandboxed JavaScript execution tool #24244
- GGML core bumped to version 0.15.1 e08c226
- WebGPU performance improvements for prefill and k-quants #24225
- Vulkan added fast paths for contiguous transfers and dot2 product extension support #23973 #24123
- Fixed CUDA ssm_scan_f32 data-races and CPU rms_norm_back in-place aliasing #24360 #24305
- Server added prompt logging to local directories #22031
π All changes | Latest release
π¨ stable-diffusion.cpp
β Release: master-679-f3fd359 β master-694-276025e
β 15 commits
- Added circular RoPE support for ideogram4 #1627
- Introduced free_sd_images function to manage memory for C API #1633
- Optimized performance by capping planner budget when models exceed streaming limits #1612
- Normalized APG diff_norm calculations by tensor size #1620
- Fixed SD3 conditioning crash when clip_l text encoder is missing #1638
- Corrected mask shape for masked flash attention #1625
- Resolved LoKR application issue by correctly marking w2_a tensors #1650
π All changes | Latest release
π€ Fresh models trending on HuggingFace:
bosonai/higgs-audio-v3-tts-4b β‘414
nex-agi/Nex-N2-mini β‘193
prefeitura-rio/Rio-3.5-Open-397B β‘108
RazzzHF/Realism_Engine_Ideogram_4 β‘90
silx-ai/Quasar-Preview β‘63
mindlab-research/Macaron-V1-Preview-749B β‘57
Zyphra/ZONOS2 β‘56
BennyDaBall/Z-Image-Engineer-V6 β‘46
PaddlePaddle/pp-ocrv6 β‘43
MooreThreads/MusaCoder-27B β‘35
apodex/Apodex-1.0-mini β‘31
Muhammadreza/alduin-4b-it-base β‘28
zjunlp/LabVLA β‘26
fancyfeast/bigasp-3 β‘20
Photoroom/prxpixel-t2i β‘20
LatentForce-ai/Cassini-1.0 β‘20
Gryphe/Pantheon-Reasoning-26B-A4B-1.1 β‘19
libertywing/FlashMemory-Deepseek-V4 β‘19
dx8152/Flux2-Klein-9B-Migration β‘19
apodex/Apodex-1.0-4B-SFT β‘18
Gryphe/Gemma-4-31B-StyleTune β‘17
VAGOsolutions/SauerkrautLM-LFM2.5-GLiNER β‘16
tsolful/zjourney-Ideogram-4-Fantasy-Realism-Refiner β‘14
π¦ llama.cpp
β Release: b9544 β b9627
β 83 commits
- Support for EAGLE3 speculative decoding #18039
- Added Gemma 4 Multi-Token Prediction (MTP) and assistant draft-model support #23398 #24282
- New architecture support for Cohere2-MoE #24260
- Multi-Token Multi-Domain (MTMD) models now support video input and a batching API #24269 #24384
- WebUI implemented as a Progressive Web App (PWA) with offline caching #23871
- WebUI added an opt-in sandboxed JavaScript execution tool #24244
- GGML core bumped to version 0.15.1 e08c226
- WebGPU performance improvements for prefill and k-quants #24225
- Vulkan added fast paths for contiguous transfers and dot2 product extension support #23973 #24123
- Fixed CUDA ssm_scan_f32 data-races and CPU rms_norm_back in-place aliasing #24360 #24305
- Server added prompt logging to local directories #22031
π All changes | Latest release
π¨ stable-diffusion.cpp
β Release: master-679-f3fd359 β master-694-276025e
β 15 commits
- Added circular RoPE support for ideogram4 #1627
- Introduced free_sd_images function to manage memory for C API #1633
- Optimized performance by capping planner budget when models exceed streaming limits #1612
- Normalized APG diff_norm calculations by tensor size #1620
- Fixed SD3 conditioning crash when clip_l text encoder is missing #1638
- Corrected mask shape for masked flash attention #1625
- Resolved LoKR application issue by correctly marking w2_a tensors #1650
π All changes | Latest release
π€ Fresh models trending on HuggingFace:
bosonai/higgs-audio-v3-tts-4b β‘414
nex-agi/Nex-N2-mini β‘193
prefeitura-rio/Rio-3.5-Open-397B β‘108
RazzzHF/Realism_Engine_Ideogram_4 β‘90
silx-ai/Quasar-Preview β‘63
mindlab-research/Macaron-V1-Preview-749B β‘57
Zyphra/ZONOS2 β‘56
BennyDaBall/Z-Image-Engineer-V6 β‘46
PaddlePaddle/pp-ocrv6 β‘43
MooreThreads/MusaCoder-27B β‘35
apodex/Apodex-1.0-mini β‘31
Muhammadreza/alduin-4b-it-base β‘28
zjunlp/LabVLA β‘26
fancyfeast/bigasp-3 β‘20
Photoroom/prxpixel-t2i β‘20
LatentForce-ai/Cassini-1.0 β‘20
Gryphe/Pantheon-Reasoning-26B-A4B-1.1 β‘19
libertywing/FlashMemory-Deepseek-V4 β‘19
dx8152/Flux2-Klein-9B-Migration β‘19
apodex/Apodex-1.0-4B-SFT β‘18
Gryphe/Gemma-4-31B-StyleTune β‘17
VAGOsolutions/SauerkrautLM-LFM2.5-GLiNER β‘16
tsolful/zjourney-Ideogram-4-Fantasy-Realism-Refiner β‘14
GitHub
[Speculative decoding] feat: add EAGLE3 speculative decoding support by ruixiang63 Β· Pull Request #18039 Β· ggml-org/llama.cpp
ImportantThe old PR has been backed up in this branch: https://github.com/ruixiang63/llama.cpp/tree/eagle3-v1-backup
The new commits in this PR have been rebased onto the latest master branch, refa...
The new commits in this PR have been rebased onto the latest master branch, refa...
π [GitHub Releases] turboderp-org/exllamav3 - 0.0.43
https://github.com/turboderp-org/exllamav3/releases/tag/v0.0.43
https://github.com/turboderp-org/exllamav3/releases/tag/v0.0.43
GitHub
Release 0.0.43 Β· turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs - Release 0.0.43 Β· turboderp-org/exllamav3
π° OpenAI - New OpenAI Academy courses for the next era of work
https://openai.com/index/academy-courses-applying-ai-at-work
π OpenAI - How an astrophysicist uses Codex to help simulate black holes
https://openai.com/index/using-codex-to-simulate-black-holes
π OpenAI - BBVA puts AI at the core of banking with OpenAI
https://openai.com/index/bbva
π OpenAI - Creating new simulations of black holes with Codex
https://openai.com/index/creating-new-simulations-black-holes
π OpenAI - Ad Tools Terms
https://openai.com/policies/ad-tools-terms
π OpenAI - OpenAI Academy
https://openai.com/academy
π OpenAI - How Preply combines AI and human tutors to personalize learning
https://openai.com/index/preply
https://openai.com/index/academy-courses-applying-ai-at-work
π OpenAI - How an astrophysicist uses Codex to help simulate black holes
https://openai.com/index/using-codex-to-simulate-black-holes
π OpenAI - BBVA puts AI at the core of banking with OpenAI
https://openai.com/index/bbva
π OpenAI - Creating new simulations of black holes with Codex
https://openai.com/index/creating-new-simulations-black-holes
π OpenAI - Ad Tools Terms
https://openai.com/policies/ad-tools-terms
π OpenAI - OpenAI Academy
https://openai.com/academy
π OpenAI - How Preply combines AI and human tutors to personalize learning
https://openai.com/index/preply
OpenAI
New OpenAI Academy courses for the next era of work
OpenAI introduces three Academy courses that help people build practical AI skills, create repeatable workflows, and apply agents in everyday work.
π [HF Models] tencent - Hy-Embodied-0.5-VLA-RoboTwin
https://huggingface.co/tencent/Hy-Embodied-0.5-VLA-RoboTwin
π [HF Models] tencent - Hy-Embodied-0.5-VLA-UMI
https://huggingface.co/tencent/Hy-Embodied-0.5-VLA-UMI
https://huggingface.co/tencent/Hy-Embodied-0.5-VLA-RoboTwin
π [HF Models] tencent - Hy-Embodied-0.5-VLA-UMI
https://huggingface.co/tencent/Hy-Embodied-0.5-VLA-UMI
huggingface.co
tencent/Hy-Embodied-0.5-VLA-RoboTwin Β· Hugging Face
Weβre on a journey to advance and democratize artificial intelligence through open source and open science.
π [HF Models] microsoft - FastContext-1.0-4B-SFT
https://huggingface.co/microsoft/FastContext-1.0-4B-SFT
https://huggingface.co/microsoft/FastContext-1.0-4B-SFT
π [HF Models] swiss-ai - Apertus-v1.1-1.5B-Instruct
https://huggingface.co/swiss-ai/Apertus-v1.1-1.5B-Instruct
π [HF Models] swiss-ai - Apertus-v1.1-4B-Instruct
https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct
π [HF Models] swiss-ai - Apertus-v1.1-0.5B-Instruct
https://huggingface.co/swiss-ai/Apertus-v1.1-0.5B-Instruct
π [HF Models] swiss-ai - Apertus-v1.1-4B
https://huggingface.co/swiss-ai/Apertus-v1.1-4B
π [HF Models] swiss-ai - Apertus-v1.1-1.5B
https://huggingface.co/swiss-ai/Apertus-v1.1-1.5B
https://huggingface.co/swiss-ai/Apertus-v1.1-1.5B-Instruct
π [HF Models] swiss-ai - Apertus-v1.1-4B-Instruct
https://huggingface.co/swiss-ai/Apertus-v1.1-4B-Instruct
π [HF Models] swiss-ai - Apertus-v1.1-0.5B-Instruct
https://huggingface.co/swiss-ai/Apertus-v1.1-0.5B-Instruct
π [HF Models] swiss-ai - Apertus-v1.1-4B
https://huggingface.co/swiss-ai/Apertus-v1.1-4B
π [HF Models] swiss-ai - Apertus-v1.1-1.5B
https://huggingface.co/swiss-ai/Apertus-v1.1-1.5B
huggingface.co
swiss-ai/Apertus-v1.1-1.5B-Instruct Β· Hugging Face
Weβre on a journey to advance and democratize artificial intelligence through open source and open science.
π° LMSys - The next generation of speculative decoding: DFlash and Spec V2
https://lmsys.org/blog/2026-06-15-next-generation-speculative-decoding-dflash-v2
https://lmsys.org/blog/2026-06-15-next-generation-speculative-decoding-dflash-v2
www.lmsys.org
The next generation of speculative decoding: DFlash and Spec V2
Using Modal and Z Lab's DFlash speculative decoding models with SGLangβs newly default Spec V2 engine, you can achieve state-of-the-art latencies for LLM inference serving. Our new, jointly-released D...
π° OpenAI - OpenAI Merchant Feed Terms of Service
https://openai.com/policies/merchant-feed-terms-of-service
π° OpenAI - Introducing the OpenAI Partner Network
https://openai.com/index/introducing-openai-partner-network
π OpenAI - Training to cycle across Antarctica with ChatGPT
https://openai.com/index/cycling-across-antarctica
https://openai.com/policies/merchant-feed-terms-of-service
π° OpenAI - Introducing the OpenAI Partner Network
https://openai.com/index/introducing-openai-partner-network
π OpenAI - Training to cycle across Antarctica with ChatGPT
https://openai.com/index/cycling-across-antarctica
OpenAI
OpenAI Merchant Feed Terms of Service
Review OpenAI Merchant Feed Terms of Service, including content requirements, data use, compliance policies, and rights for sharing product data.
π° NVIDIA - Boosting MoE Training Throughput with Advanced Fusion Kernels
Mixture-of-experts (MoE) models have quickly become a foundational component of modern, large-scale AI systems. They are widely adopted because they enableβ¦
https://developer.nvidia.com/blog/boosting-moe-training-throughput-with-advanced-fusion-kernels/
π° NVIDIA - Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models
Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts itβ¦
https://developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action-models/
Mixture-of-experts (MoE) models have quickly become a foundational component of modern, large-scale AI systems. They are widely adopted because they enableβ¦
https://developer.nvidia.com/blog/boosting-moe-training-throughput-with-advanced-fusion-kernels/
π° NVIDIA - Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models
Quick glossary for readers new to VLA/WAM terminology VLA Vision-Language-Action model: a robot policy that starts from a pretrained VLM backbone and adapts itβ¦
https://developer.nvidia.com/blog/pretrained-to-imagine-fine-tuned-to-act-the-rise-of-world-action-models/
NVIDIA Technical Blog
Boosting MoE Training Throughput with Advanced Fusion Kernels
Mixture-of-experts (MoE) models have quickly become a foundational component of modern, large-scale AI systems. They are widely adopted because they enable substantially larger model capacity whileβ¦
π [HF Models] inclusionAI - Ling-2.6-flash-base
https://huggingface.co/inclusionAI/Ling-2.6-flash-base
π [HF Models] inclusionAI - Ling-2.6-1T-base
https://huggingface.co/inclusionAI/Ling-2.6-1T-base
https://huggingface.co/inclusionAI/Ling-2.6-flash-base
π [HF Models] inclusionAI - Ling-2.6-1T-base
https://huggingface.co/inclusionAI/Ling-2.6-1T-base
huggingface.co
inclusionAI/Ling-2.6-flash-base Β· Hugging Face
Weβre on a journey to advance and democratize artificial intelligence through open source and open science.
π° Google AI Blog - Unlocking the Power of the TPU Stack: Introducing our new Developer Hub
Google has officially launched the TPU Developer Hub, a centralized educational resource designed to help model builders and developers maximize the performance of Google Cloud TPUs. The hub offers code-first resources, open-source recipes, and deep-dive documentation covering hardware architecture, software optimization, debugging, parallelism, and networking. These materials are tailored for both human developers and AI-assisted tools to streamline everything from large-scale training to low-latency inference workloads.
https://developers.googleblog.com/en/unlocking-the-power-of-the-tpu-stack-introducing-our-new-developer-hub/
Google has officially launched the TPU Developer Hub, a centralized educational resource designed to help model builders and developers maximize the performance of Google Cloud TPUs. The hub offers code-first resources, open-source recipes, and deep-dive documentation covering hardware architecture, software optimization, debugging, parallelism, and networking. These materials are tailored for both human developers and AI-assisted tools to streamline everything from large-scale training to low-latency inference workloads.
https://developers.googleblog.com/en/unlocking-the-power-of-the-tpu-stack-introducing-our-new-developer-hub/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Announcing the TPU Developer Hub: a code-first, agent-friendly resource for ML developers. Get actionable tutorials on hardware architecture, XProf debugging, and Pallas kernels.
π° Claude Blog - Meet the winners of the Built with Opus 4.7 Claude Code hackathon
https://claude.com/blog/meet-the-winners-of-built-with-opus-4-7-claude-code-hackathon
https://claude.com/blog/meet-the-winners-of-built-with-opus-4-7-claude-code-hackathon
Claude
Meet the winners of the Built with Opus 4.7 Claude Code hackathon | Claude by Anthropic
We chatted with the winners of our Built with Opus 4.7 hackathon about their projects, tackling medical training, electronics repair, computer science education, interactive play, home repair, and factory maintenance.
π° OpenAI - Predicting model behavior before release by simulating deployment
https://openai.com/index/deployment-simulation
https://openai.com/index/deployment-simulation
OpenAI
Predicting model behavior before release by simulating deployment
OpenAI introduces Deployment Simulation, a method to predict AI model behavior before deployment using real conversation data to improve safety and evaluation accuracy.