DeepSeek
1.39K subscribers
52 photos
41 links
Unravel the mystery of AGI with curiousity. Answer the essential questions with long-termism. https://www.deepseek.com
Download Telegram
API pricing update ๐Ÿ’ฐ

With the V4 lineup release, weโ€™re updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling. ๐Ÿ“‰

New pricing takes effect at 16:00 UTC, Aug 16, 2026 ๐Ÿ•’
๐Ÿ˜ฑ18๐Ÿ‘Ž5๐Ÿ˜4๐Ÿณ4๐Ÿ‘Œ1๐Ÿ˜ญ1
๐Ÿงฉ DeepSeek Harness v0.1 is now available in Developer Preview!

๐Ÿ”น Weโ€™re opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
๐Ÿ”น Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.

Try it now!
https://github.com/deepseek-ai/deepseek-harness
๐Ÿ”ฅ12๐Ÿณ10๐Ÿ’ฉ2๐Ÿ‘พ2โค1
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! ๐Ÿš€

๐Ÿ”น This experimental multimodal model matches DeepSeek-V4-Flash on text capabilitiesโ€”including agents, reasoning, and world knowledge.
๐Ÿ”น On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.

Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.
๐Ÿณ9โค2
Multimodality unlocks more agent use cases. ๐Ÿ‘€

V4-Flash-Vision-Exp works smoothly across agent frameworks, combining visual understanding with a wide range of tools to unlock more practical workflows.
๐Ÿณ7
Multimodal API support ๐Ÿ”Œ

๐Ÿ”น Set model='deepseek-v4-flash-vision-exp'
๐Ÿ”น Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing
๐Ÿ”น Supports Chat Completions, Messages & Responses
๐Ÿ”น Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API.

Docs: https://api-docs.deepseek.com/guides/vision
๐Ÿณ8
Files API is now live. ๐Ÿ“

๐Ÿ”น Free to use
๐Ÿ”น Upload an image once, then reference it by file_id to save request bandwidth
๐Ÿ”น Reuse the same image across requestsโ€”no need to upload it again

Learn more: https://api-docs.deepseek.com/guides/files_api
๐Ÿณ13
๐Ÿš€ Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.

๐Ÿ”น Introducing the smallest model in our new architecture family, with native visual understanding.
๐Ÿ”น Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
๐Ÿณ9๐Ÿคก3
๐Ÿง  Asymmetric architecture. More intelligence, less cost.

๐Ÿ”น 552B-parameter MoE.
๐Ÿ”น New Causal Encoderโ€“Decoder architecture: just 8B active parameters for input, 16B for output.
๐Ÿ”น New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
๐Ÿณ9๐Ÿคก2
๐Ÿ’พ Smaller KV cache. Bigger savings.

Compared with the previous generation, V4.1-Flashโ€™s KV cache needs just:
๐Ÿ”น 1/4 the HBM
๐Ÿ”น 1/8 the SSD storage
Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly.
๐Ÿณ11๐Ÿคก4
โšก V4.1-Flash is now live on the DeepSeek API with native multimodal support.

Set your model to deepseek-flash.

๐Ÿ”น V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
๐Ÿ”น Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. Weโ€™re phasing out V4-Pro.
๐Ÿ”น Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.

๐Ÿค Official partners WorkBuddy (including Codebuddy) & OpenCode
now fully support V4.1-Flash. Try it today!
๐Ÿณ8โค1
๐Ÿ’ฐ More efficient architecture. Lower API prices.

V4.1-Flash lets us serve more users at a lower cost. Weโ€™re passing the savings on to you.

๐Ÿ”น Peak/off-peak pricing continues to balance demand.
๐Ÿ”น Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save.
๐Ÿ”น New pricing takes effect at 04:00 UTC on Sept 10, 2026.
๐Ÿณ7
๐ŸŒ Supporting open source. Expanding deployment options.

Weโ€™ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.
Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Letโ€™s talk.

๐Ÿ”น Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
๐Ÿ”น Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
๐Ÿณ8โค4๐Ÿ‘1