🆕 [HF Models] google - diffusiongemma-26B-A4B-it
https://huggingface.co/google/diffusiongemma-26B-A4B-it
https://huggingface.co/google/diffusiongemma-26B-A4B-it
📰 PyTorch - Portable vLLM Model Inference Kernels in Helion
TL;DR Helion kernels were integrated into vLLM for FP8 inference using Qwen3 models and evaluated across NVIDIA H100 and B200 GPUs. The experiments show that Helion provides a productive PyTorch-native...
https://pytorch.org/blog/portable-vllm-model-inference-kernels-in-helion/
TL;DR Helion kernels were integrated into vLLM for FP8 inference using Qwen3 models and evaluated across NVIDIA H100 and B200 GPUs. The experiments show that Helion provides a productive PyTorch-native...
https://pytorch.org/blog/portable-vllm-model-inference-kernels-in-helion/
📰 Google AI Blog - DiffusionGemma: The Developer Guide
DiffusionGemma is an experimental text-generation model built on the Gemma 4 architecture that uses diffusion-based parallel generation instead of token-by-token autoregression, enabling much faster inference, bidirectional context awareness, and real-time self-correction while remaining deployable on consumer GPUs. Its architecture generates and refines 256-token blocks in parallel through iterative denoising, allowing it to handle complex constraint-based tasks such as Sudoku more effectively than traditional language models and demonstrating strong gains from fine-tuning. The model integrates with vLLM and other popular inference frameworks, giving developers access to a new non-autoregressive approach that combines high performance, efficient long-context scaling, and straightforward customization and deployment.
https://developers.googleblog.com/en/diffusiongemma-the-developer-guide/
DiffusionGemma is an experimental text-generation model built on the Gemma 4 architecture that uses diffusion-based parallel generation instead of token-by-token autoregression, enabling much faster inference, bidirectional context awareness, and real-time self-correction while remaining deployable on consumer GPUs. Its architecture generates and refines 256-token blocks in parallel through iterative denoising, allowing it to handle complex constraint-based tasks such as Sudoku more effectively than traditional language models and demonstrating strong gains from fine-tuning. The model integrates with vLLM and other popular inference frameworks, giving developers access to a new non-autoregressive approach that combines high performance, efficient long-context scaling, and straightforward customization and deployment.
https://developers.googleblog.com/en/diffusiongemma-the-developer-guide/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Learn how DiffusionGemma reimagines text generation with diffusion-based parallel decoding, bidirectional context, and self-correction, delivering faster inference and new possibilities for customization and deployment.
📰 NVIDIA - Run DiffusionGemma on NVIDIA for Developer-Ready, High-Throughput Text Generation
Developers building real-time AI—such as chat assistants, copilots, and agentic workflows—are often constrained by token-by-token generation speed.
https://developer.nvidia.com/blog/run-diffusiongemma-on-nvidia-for-developer-ready-high-throughput-text-generation/
Developers building real-time AI—such as chat assistants, copilots, and agentic workflows—are often constrained by token-by-token generation speed.
https://developer.nvidia.com/blog/run-diffusiongemma-on-nvidia-for-developer-ready-high-throughput-text-generation/
NVIDIA Technical Blog
Run DiffusionGemma on NVIDIA for Developer-Ready, High-Throughput Text Generation
Developers building real-time AI—such as chat assistants, copilots, and agentic workflows—are often constrained by token-by-token generation speed. This limits responsiveness, increases serving costs…
📰 HuggingFace - Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP
https://huggingface.co/blog/torch-mlp-fusion
https://huggingface.co/blog/torch-mlp-fusion
huggingface.co
Profiling in PyTorch (Part 2): From nn.Linear to a Fused MLP
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 Anthropic - DXC will integrate Claude into the systems banks, airlines, and other regulated industries rely on
https://www.anthropic.com/news/dxc-anthropic-alliance
📰 Anthropic - Introducing Claude Corps
https://www.anthropic.com/news/claude-corps
https://www.anthropic.com/news/dxc-anthropic-alliance
📰 Anthropic - Introducing Claude Corps
https://www.anthropic.com/news/claude-corps
Anthropic
DXC will integrate Claude into the systems banks, airlines, and other regulated industries rely on
Anthropic is announcing a multi-year global alliance with DXC Technology, one of the world’s largest IT services companies.
📰 NVIDIA - One-Click Multi-Tenant Security with NVIDIA Quantum InfiniBand
NVIDIA Quantum InfiniBand now offers intent-based security profiles in Unified Fabric Manager (UFM) that enable multi-tenant fabric security in a single click.
https://developer.nvidia.com/blog/one-click-multi-tenant-security-with-nvidia-quantum-infiniband/
NVIDIA Quantum InfiniBand now offers intent-based security profiles in Unified Fabric Manager (UFM) that enable multi-tenant fabric security in a single click.
https://developer.nvidia.com/blog/one-click-multi-tenant-security-with-nvidia-quantum-infiniband/
NVIDIA Technical Blog
One-Click Multi-Tenant Security with NVIDIA Quantum InfiniBand
NVIDIA Quantum InfiniBand now offers intent-based security profiles in Unified Fabric Manager (UFM) that enable multi-tenant fabric security in a single click. NVIDIA Quantum InfiniBand supports three…
📰 PyTorch - PyTorch Meetup Singapore: A milestone in APAC
TL;DR Eighty engineers, researchers, and community builders gathered for the inaugural PyTorch Meetup Singapore. Hosted at the Red Hat Asia Pacific office and organised by Sudhir Dharanendraiah, Ayush Satyam, Sumantro...
https://pytorch.org/blog/pytorch-meetup-singapore-a-milestone-in-apac/
TL;DR Eighty engineers, researchers, and community builders gathered for the inaugural PyTorch Meetup Singapore. Hosted at the Red Hat Asia Pacific office and organised by Sudhir Dharanendraiah, Ayush Satyam, Sumantro...
https://pytorch.org/blog/pytorch-meetup-singapore-a-milestone-in-apac/
📰 HuggingFace - olmo-eval: An evaluation workbench for the model development loop
https://huggingface.co/blog/allenai/olmo-eval
https://huggingface.co/blog/allenai/olmo-eval
huggingface.co
olmo-eval: An evaluation workbench for the model development loop
A Blog post by Ai2 on Hugging Face
🔄 [GitHub Releases] turboderp-org/exllamav3 - 0.0.41
https://github.com/turboderp-org/exllamav3/releases/tag/v0.0.41
https://github.com/turboderp-org/exllamav3/releases/tag/v0.0.41
GitHub
Release 0.0.41 · turboderp-org/exllamav3
An optimized quantization and inference library for running LLMs locally on modern consumer-class GPUs - Release 0.0.41 · turboderp-org/exllamav3
📰 Claude Blog - The evolution of agentic surfaces: building with Claude Managed Agents
https://claude.com/blog/building-with-claude-managed-agents
📰 Claude Blog - New in Claude Managed Agents: run agents on a schedule and store environment variables in vaults
https://claude.com/blog/whats-new-in-claude-managed-agents
📰 Claude Blog - Building intelligent apps for Apple platforms with Claude in the Foundation Models framework
https://claude.com/blog/claude-for-foundation-models
📰 Claude Blog - Observability for developers building connectors
https://claude.com/blog/observability-for-developers-building-connectors
https://claude.com/blog/building-with-claude-managed-agents
📰 Claude Blog - New in Claude Managed Agents: run agents on a schedule and store environment variables in vaults
https://claude.com/blog/whats-new-in-claude-managed-agents
📰 Claude Blog - Building intelligent apps for Apple platforms with Claude in the Foundation Models framework
https://claude.com/blog/claude-for-foundation-models
📰 Claude Blog - Observability for developers building connectors
https://claude.com/blog/observability-for-developers-building-connectors
Claude
The evolution of agentic surfaces: building with Claude Managed Agents | Claude by Anthropic
Claude Managed Agents allows teams to build and deploy agents in production environments reliably at scale. Here’s why and how teams are using it.
📰 Anthropic - Results from the first Anthropic Public Record
https://www.anthropic.com/news/anthropic-public-record
📰 Anthropic - TCS and Anthropic partner to bring Claude to regulated industries
https://www.anthropic.com/news/tcs-anthropic-partnership
https://www.anthropic.com/news/anthropic-public-record
📰 Anthropic - TCS and Anthropic partner to bring Claude to regulated industries
https://www.anthropic.com/news/tcs-anthropic-partnership
Anthropic
Results from the first Anthropic Public Record
Anthropic Public Record is a national survey of attitudes and opinions towards AI.
📰 NVIDIA - NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark
AI agents have fundamentally changed the complexity of inference workloads. Until now, the industry has struggled to define a standard for measuring how…
https://developer.nvidia.com/blog/nvidia-achieves-leading-agentic-coding-performance-on-first-agentic-ai-benchmark/
📰 NVIDIA - Deploy Long-Context Reasoning and Agentic Workflows with MiniMax M3 on NVIDIA Accelerated Infrastructure
As enterprise AI adoption scales, developers are increasingly forced to stitch together fragmented pipelines—separate models for text, vision…
https://developer.nvidia.com/blog/deploy-long-context-reasoning-and-agentic-workflows-with-minimax-m3-on-nvidia-accelerated-infrastructure/
AI agents have fundamentally changed the complexity of inference workloads. Until now, the industry has struggled to define a standard for measuring how…
https://developer.nvidia.com/blog/nvidia-achieves-leading-agentic-coding-performance-on-first-agentic-ai-benchmark/
📰 NVIDIA - Deploy Long-Context Reasoning and Agentic Workflows with MiniMax M3 on NVIDIA Accelerated Infrastructure
As enterprise AI adoption scales, developers are increasingly forced to stitch together fragmented pipelines—separate models for text, vision…
https://developer.nvidia.com/blog/deploy-long-context-reasoning-and-agentic-workflows-with-minimax-m3-on-nvidia-accelerated-infrastructure/
NVIDIA Technical Blog
NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark
AI agents have fundamentally changed the complexity of inference workloads. Until now, the industry has struggled to define a standard for measuring how inference systems perform under these…