📰 Claude Blog - Working at the frontier: How Hebbia builds AI for financial diligence that can't miss a detail
https://claude.com/blog/working-at-the-frontier-how-hebbia-builds-ai-for-financial-diligence-that-cant-miss-a-detail
https://claude.com/blog/working-at-the-frontier-how-hebbia-builds-ai-for-financial-diligence-that-cant-miss-a-detail
Claude
Working at the frontier: How Hebbia builds AI for financial diligence that can't miss a detail | Claude by Anthropic
How Anthropic's Claude Fable 5 beat Hebbia's finance-specific model evaluations, achieving their biggest accuracy gain yet.
📰 Anthropic Research - Claude’s values across models and languages
https://www.anthropic.com/research/claude-values-models-languages
📰 Anthropic Research - Claude plays robotics
https://www.anthropic.com/research/claude-plays-robotics
https://www.anthropic.com/research/claude-values-models-languages
📰 Anthropic Research - Claude plays robotics
https://www.anthropic.com/research/claude-plays-robotics
Anthropic
Claude’s values across models and languages
We analyzed 300,000 real conversations to measure the values Claude expresses across models and languages, compressed into four interpretable axes.
📰 NVIDIA - NVIDIA Ising Decoding Cuts Color Code Logical Error Rates by Over 300X
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes…
https://developer.nvidia.com/blog/nvidia-ising-decoding-cuts-color-code-logical-error-rates-by-over-300x/
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes…
https://developer.nvidia.com/blog/nvidia-ising-decoding-cuts-color-code-logical-error-rates-by-over-300x/
NVIDIA Technical Blog
NVIDIA Ising Decoding Cuts Color Code Logical Error Rates by Over 300X
Useful quantum computers will require fault tolerant logical operations. Researchers are actively exploring many different quantum error correction (QEC) codes to enable this…
🆕 [HF Models] inclusionAI - SingGuard-NSFA-9B-GGUF
https://huggingface.co/inclusionAI/SingGuard-NSFA-9B-GGUF
🆕 [HF Models] inclusionAI - SingGuard-NSFA-4B-GGUF
https://huggingface.co/inclusionAI/SingGuard-NSFA-4B-GGUF
🆕 [HF Models] inclusionAI - SingGuard-NSFA-2B-GGUF
https://huggingface.co/inclusionAI/SingGuard-NSFA-2B-GGUF
🆕 [HF Models] inclusionAI - SingGuard-NSFA-0.8B-GGUF
https://huggingface.co/inclusionAI/SingGuard-NSFA-0.8B-GGUF
https://huggingface.co/inclusionAI/SingGuard-NSFA-9B-GGUF
🆕 [HF Models] inclusionAI - SingGuard-NSFA-4B-GGUF
https://huggingface.co/inclusionAI/SingGuard-NSFA-4B-GGUF
🆕 [HF Models] inclusionAI - SingGuard-NSFA-2B-GGUF
https://huggingface.co/inclusionAI/SingGuard-NSFA-2B-GGUF
🆕 [HF Models] inclusionAI - SingGuard-NSFA-0.8B-GGUF
https://huggingface.co/inclusionAI/SingGuard-NSFA-0.8B-GGUF
huggingface.co
inclusionAI/SingGuard-NSFA-9B-GGUF · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🔄 [GitHub Releases] sgl-project/sglang - v0.5.15.post1
https://github.com/sgl-project/sglang/releases/tag/v0.5.15.post1
https://github.com/sgl-project/sglang/releases/tag/v0.5.15.post1
GitHub
Release v0.5.15.post1 · sgl-project/sglang
v0.5.15.post1 includes a few patches, mostly for GLM 5.2
#30454 #30627: Fix DSA model launching on non Cuda/HIP devices
#30858: Fix flashinfer dependency on Cuda 12 images
#31001: Fix NaN outputs ...
#30454 #30627: Fix DSA model launching on non Cuda/HIP devices
#30858: Fix flashinfer dependency on Cuda 12 images
#31001: Fix NaN outputs ...
🔄 [GitHub Releases] vllm-project/vllm - v0.25.1
https://github.com/vllm-project/vllm/releases/tag/v0.25.1
https://github.com/vllm-project/vllm/releases/tag/v0.25.1
GitHub
Release v0.25.1 · vllm-project/vllm
vLLM v0.25.1
Highlights
This release features 2 commits from 2 contributors (1 new)!
v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0.
Bug Fixes
Avoid blocking model ...
Highlights
This release features 2 commits from 2 contributors (1 new)!
v0.25.1 is a patch release containing two targeted bug fixes on top of v0.25.0.
Bug Fixes
Avoid blocking model ...
🆕 [HF Models] nvidia - Nemotron-3-Embed-8B-BF16
https://huggingface.co/nvidia/Nemotron-3-Embed-8B-BF16
🆕 [HF Models] nvidia - Nemotron-3-Embed-1B-BF16
https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16
https://huggingface.co/nvidia/Nemotron-3-Embed-8B-BF16
🆕 [HF Models] nvidia - Nemotron-3-Embed-1B-BF16
https://huggingface.co/nvidia/Nemotron-3-Embed-1B-BF16
huggingface.co
nvidia/Nemotron-3-Embed-8B-BF16 · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
🔄 [GitHub Releases] turboderp-org/exllamav3 - 1.0.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.0.0
https://github.com/turboderp-org/exllamav3/releases/tag/v1.0.0
GitHub
Release 1.0.0 · turboderp-org/exllamav3
-> Small writeup with charts.
Remove flash-attention-2 and xformers dependencies
New attention kernel with online cache quantization, dual input for SWA layers and attention sinks
New conv1d ke...
Remove flash-attention-2 and xformers dependencies
New attention kernel with online cache quantization, dual input for SWA layers and attention sinks
New conv1d ke...
📰 Google DeepMind - Reconstructing Pelé’s “lost” goal
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/reconstructing-peles-lost-goal/
https://blog.google/innovation-and-ai/models-and-research/google-deepmind/reconstructing-peles-lost-goal/
Google
Reconstructing Pelé’s “lost” goal
See how Google DeepMind AI technology reconstructed Pelé’s legendary 1959 lost goal at Rua Javari in our new mini-documentary.
📰 Google AI Blog - Unlocking the Next Era of On-Device AI with Google Tensor and Pixel
At Google I/O Connect India, Google showcased the future of 100% private, on-device AI powered by the custom Tensor SoC and TPU for the new Pixel 10 family. The event debuted the lightweight Gemma 4 E2B model, which runs natively on the device to enable completely offline multimodal features like AI chat, real-time image recognition, and personal agent tasks. Developers can start building these secure, edge-based applications today by accessing the newly announced Tensor SDK beta and its accompanying open-source resources.
https://developers.googleblog.com/en/unlocking-the-next-era-of-on-device-ai-with-google-tensor-and-pixel/
At Google I/O Connect India, Google showcased the future of 100% private, on-device AI powered by the custom Tensor SoC and TPU for the new Pixel 10 family. The event debuted the lightweight Gemma 4 E2B model, which runs natively on the device to enable completely offline multimodal features like AI chat, real-time image recognition, and personal agent tasks. Developers can start building these secure, edge-based applications today by accessing the newly announced Tensor SDK beta and its accompanying open-source resources.
https://developers.googleblog.com/en/unlocking-the-next-era-of-on-device-ai-with-google-tensor-and-pixel/
Googleblog
Google for Developers Blog - News about Web, Mobile, AI and Cloud
Build the next generation of offline AI apps. Read our Google I/O Connect recap to learn about Gemma 4 for Pixel 10 and access the new Tensor SDK beta.
📰 LMSys - Serving GLM5.2 NVFP4 Agentic Workload with SGLang: Reaching 500 TPS in 2 Weeks
https://lmsys.org/blog/2026-07-13-glm52-optimization
https://lmsys.org/blog/2026-07-13-glm52-optimization
www.lmsys.org
Serving GLM5.2 NVFP4 Agentic Workload with SGLang: Reaching 500 TPS in 2 Weeks
- More than 500 TPS on 8xB300 (bs=1)
- Sync free speculative decoding for GLM 5.2 MTP
- Built-in IndexShare MTP with Spec V2
- 2.33x faster TopK-V2 for ISL 80k
- Indexer prologue fusion
- Gemm kernels...
- Sync free speculative decoding for GLM 5.2 MTP
- Built-in IndexShare MTP with Spec V2
- 2.33x faster TopK-V2 for ISL 80k
- Indexer prologue fusion
- Gemm kernels...
📰 OpenAI - How to manage AI investments in the agentic era
https://openai.com/index/managing-ai-investments-in-agentic-era
📰 OpenAI - How data science teams use ChatGPT Work
https://openai.com/academy/codex-for-work/how-data-science-teams-use-codex
📰 OpenAI - How sales teams use ChatGPT Work
https://openai.com/academy/codex-for-work/how-sales-teams-use-codex
https://openai.com/index/managing-ai-investments-in-agentic-era
📰 OpenAI - How data science teams use ChatGPT Work
https://openai.com/academy/codex-for-work/how-data-science-teams-use-codex
📰 OpenAI - How sales teams use ChatGPT Work
https://openai.com/academy/codex-for-work/how-sales-teams-use-codex
OpenAI
How to manage AI investments in the agentic era
Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.
📰 NVIDIA - Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when…
https://developer.nvidia.com/blog/lessons-from-the-leaderboard-what-5000-kagglers-taught-us-about-improving-ai-reasoning/
📰 NVIDIA - How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes…
https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/
📰 NVIDIA - Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning…
https://developer.nvidia.com/blog/post-train-nvidia-cosmos-3-in-one-day-using-agent-skills/
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when…
https://developer.nvidia.com/blog/lessons-from-the-leaderboard-what-5000-kagglers-taught-us-about-improving-ai-reasoning/
📰 NVIDIA - How to Run an Autoresearch Workflow with RL Agent Skills and NVIDIA NeMo
Coding AI agents are becoming practical operators for long-running machine learning (ML) workflows. They can inspect repositories, set up runtimes…
https://developer.nvidia.com/blog/how-to-run-an-autoresearch-workflow-with-rl-agent-skills-and-nvidia-nemo/
📰 NVIDIA - Post-Train NVIDIA Cosmos 3 in One Day Using Agent Skills
What if autonomous coding AI agents could push your vision reasoning models above 90% accuracy with almost no manual effort? When adapting vision reasoning…
https://developer.nvidia.com/blog/post-train-nvidia-cosmos-3-in-one-day-using-agent-skills/
NVIDIA Technical Blog
Lessons From the Leaderboard: What 5,000+ Kagglers Taught Us About Improving AI Reasoning
The NVIDIA Nemotron Model Reasoning Challenge invited the Kaggle community to explore a focused question: What techniques can improve reasoning accuracy when everyone starts from the same open model…
🆕 [HF Models] microsoft - bitnet-embedding-0.6b
https://huggingface.co/microsoft/bitnet-embedding-0.6b
🆕 [HF Models] microsoft - bitnet-embedding-270m
https://huggingface.co/microsoft/bitnet-embedding-270m
https://huggingface.co/microsoft/bitnet-embedding-0.6b
🆕 [HF Models] microsoft - bitnet-embedding-270m
https://huggingface.co/microsoft/bitnet-embedding-270m
huggingface.co
microsoft/bitnet-embedding-0.6b · Hugging Face
We’re on a journey to advance and democratize artificial intelligence through open source and open science.
📰 HuggingFace - Introducing Real World VoiceEQ: Measuring the human quality of voice AI
https://huggingface.co/blog/real-world-voiceeq
https://huggingface.co/blog/real-world-voiceeq