GenAI monitor
551 subscribers
4.37K links
AI frontier model updates & open source LLM releases
Download Telegram
📰 PyTorch - Serving DeepSeek-V4 on GB300 with SGLang: 5x Higher Throughput at the Same Interactivity Since Day-0
TL;DR: DeepSeek-V4 support was live in SGLang on Day-0, but the Day-0 stack was only the starting point. Since launch, we have coordinated a set of kernel, runtime, and hardening...

https://pytorch.org/blog/serving-deepseek-v4-on-gb300-with-sglang-5x-higher-throughput-at-the-same-interactivity-since-day-0/
📰 NVIDIA - Boost Inference Performance up to 15x on NVIDIA Blackwell Using DFlash Speculative Decoding
As AI systems move from single-turn interactions to coordinated multiagent workflows, low-latency inference becomes increasingly important.

https://developer.nvidia.com/blog/boost-inference-performance-up-to-15x-on-nvidia-blackwell-using-dflash-speculative-decoding/


📰 NVIDIA - How Telcos Build Autonomous Networks with Agentic AI
Telecom operators are adopting AI across network operations, customer care, and back-office workflows, but most are still early in the journey to autonomy.

https://developer.nvidia.com/blog/how-telcos-build-autonomous-networks-with-agentic-ai/
🆕 [HF Models] Qwen - Qwen-AgentWorld-35B-A3B


https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B
📰 PyTorch - TokenSpeed-Kernel: Portable APIs and High-Performance Kernels for Multi-Silicon LLM Inference
TL;DR The TokenSpeed-kernel is a standalone, open-source subsystem designed to solve backend complexity in LLM inference. It introduces a clean, layered API and registry system that decouples the high-level runtime...

https://pytorch.org/blog/lightseek-tokenspeed-kernel/
📰 HuggingFace - Which tokens does a hybrid model predict better?


https://huggingface.co/blog/allenai/hybrid-token-prediction


📰 HuggingFace - Run a vLLM Server on HF Jobs in One Command


https://huggingface.co/blog/vllm-jobs