DeepSeek
1.38K subscribers
52 photos
41 links
Unravel the mystery of AGI with curiousity. Answer the essential questions with long-termism. https://www.deepseek.com
Download Telegram
🧩 DeepSeek Harness v0.1 is now available in Developer Preview!

🔹 We’re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
🔹 Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.

Try it now!
https://github.com/deepseek-ai/deepseek-harness
🔥12🐳10💩2👾21
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! 🚀

🔹 This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge.
🔹 On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.

Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.
🐳92
Multimodality unlocks more agent use cases. 👀

V4-Flash-Vision-Exp works smoothly across agent frameworks, combining visual understanding with a wide range of tools to unlock more practical workflows.
🐳7
Multimodal API support 🔌

🔹 Set model='deepseek-v4-flash-vision-exp'
🔹 Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing
🔹 Supports Chat Completions, Messages & Responses
🔹 Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API.

Docs: https://api-docs.deepseek.com/guides/vision
🐳8
Files API is now live. 📁

🔹 Free to use
🔹 Upload an image once, then reference it by file_id to save request bandwidth
🔹 Reuse the same image across requests—no need to upload it again

Learn more: https://api-docs.deepseek.com/guides/files_api
🐳13
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.

🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
🐳9🤡3
🧠 Asymmetric architecture. More intelligence, less cost.

🔹 552B-parameter MoE.
🔹 New Causal Encoder–Decoder architecture: just 8B active parameters for input, 16B for output.
🔹 New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
🐳9🤡2
💾 Smaller KV cache. Bigger savings.

Compared with the previous generation, V4.1-Flash’s KV cache needs just:
🔹 1/4 the HBM
🔹 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.
🐳11🤡4
V4.1-Flash is now live on the DeepSeek API with native multimodal support.

Set your model to deepseek-flash.

🔹 V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
🔹 Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. We’re phasing out V4-Pro.
🔹 Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.

🤝 Official partners WorkBuddy (including Codebuddy) & OpenCode
now fully support V4.1-Flash. Try it today!
🐳81
💰 More efficient architecture. Lower API prices.

V4.1-Flash lets us serve more users at a lower cost. We’re passing the savings on to you.

🔹 Peak/off-peak pricing continues to balance demand.
🔹 Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save.
🔹 New pricing takes effect at 04:00 UTC on Sept 10, 2026.
🐳7
🌐 Supporting open source. Expanding deployment options.

We’ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.
Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let’s talk.

🔹 Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
🔹 Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
🐳84👍1