API pricing update ๐ฐ
With the V4 lineup release, weโre updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling. ๐
New pricing takes effect at 16:00 UTC, Aug 16, 2026 ๐
With the V4 lineup release, weโre updating our API pricing and introducing peak and off-peak rates. Off-peak rates are 50% lower than peak, enabling more flexible workload scheduling. ๐
New pricing takes effect at 16:00 UTC, Aug 16, 2026 ๐
๐ฑ18๐5๐4๐ณ4๐1๐ญ1
๐งฉ DeepSeek Harness v0.1 is now available in Developer Preview!
๐น Weโre opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
๐น Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.
Try it now!
https://github.com/deepseek-ai/deepseek-harness
๐น Weโre opening it up to developers building agent harnesses worldwide and open-sourcing the codebase in MIT license.
๐น Powered by the Cordis meta-framework, DeepSeek Harness is an agent harness built around one core idea: Everything is a plugin. Models, tools, skills, sessions, sandboxes, filesystems, loops, orchestration, and UI are ALL implemented as plugins, and can be mixed, matched, replaced, and extended.
Try it now!
https://github.com/deepseek-ai/deepseek-harness
๐ฅ12๐ณ10๐ฉ2๐พ2โค1
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! ๐
๐น This experimental multimodal model matches DeepSeek-V4-Flash on text capabilitiesโincluding agents, reasoning, and world knowledge.
๐น On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.
Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.
๐น This experimental multimodal model matches DeepSeek-V4-Flash on text capabilitiesโincluding agents, reasoning, and world knowledge.
๐น On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8.
Try it with model='deepseek-v4-flash-vision-exp'. DeepSeek Harness 0.1.1 was released today with out-of-the-box support for the new model.
๐ณ9โค2
Multimodality unlocks more agent use cases. ๐
V4-Flash-Vision-Exp works smoothly across agent frameworks, combining visual understanding with a wide range of tools to unlock more practical workflows.
V4-Flash-Vision-Exp works smoothly across agent frameworks, combining visual understanding with a wide range of tools to unlock more practical workflows.
๐ณ7
Multimodal API support ๐
๐น Set model='deepseek-v4-flash-vision-exp'
๐น Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing
๐น Supports Chat Completions, Messages & Responses
๐น Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API.
Docs: https://api-docs.deepseek.com/guides/vision
๐น Set model='deepseek-v4-flash-vision-exp'
๐น Images are tokenized for billing: up to 384 tokens each, at V4-Flash pricing
๐น Supports Chat Completions, Messages & Responses
๐น Supports mixed text + image input; images can be provided via base64, external URLs, or the Files API.
Docs: https://api-docs.deepseek.com/guides/vision
๐ณ8
Files API is now live. ๐
๐น Free to use
๐น Upload an image once, then reference it by file_id to save request bandwidth
๐น Reuse the same image across requestsโno need to upload it again
Learn more: https://api-docs.deepseek.com/guides/files_api
๐น Free to use
๐น Upload an image once, then reference it by file_id to save request bandwidth
๐น Reuse the same image across requestsโno need to upload it again
Learn more: https://api-docs.deepseek.com/guides/files_api
๐ณ13
๐ง Asymmetric architecture. More intelligence, less cost.
๐น 552B-parameter MoE.
๐น New Causal EncoderโDecoder architecture: just 8B active parameters for input, 16B for output.
๐น New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
๐น 552B-parameter MoE.
๐น New Causal EncoderโDecoder architecture: just 8B active parameters for input, 16B for output.
๐น New pre-training methods + larger-scale RL post-training deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro.
๐ณ9๐คก2
โก V4.1-Flash is now live on the DeepSeek API with native multimodal support.
Set your model to deepseek-flash.
๐น V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
๐น Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. Weโre phasing out V4-Pro.
๐น Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.
๐ค Official partners WorkBuddy (including Codebuddy) & OpenCode
now fully support V4.1-Flash. Try it today!
Set your model to deepseek-flash.
๐น V4-Flash & V4-Flash-Vision-Exp are retired. For compatibility, deepseek-v4-flash and deepseek-v4-flash-vision-exp temporarily route to V4.1-Flash.
๐น Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime. Weโre phasing out V4-Pro.
๐น Starting at 04:00 UTC on Sept 14, 2026, all deepseek-v4-pro requests will route to V4.1-Flash at V4.1-Flash rates. This will continue until V4.1-Pro launches.
๐ค Official partners WorkBuddy (including Codebuddy) & OpenCode
now fully support V4.1-Flash. Try it today!
๐ณ8โค1
๐ฐ More efficient architecture. Lower API prices.
V4.1-Flash lets us serve more users at a lower cost. Weโre passing the savings on to you.
๐น Peak/off-peak pricing continues to balance demand.
๐น Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save.
๐น New pricing takes effect at 04:00 UTC on Sept 10, 2026.
V4.1-Flash lets us serve more users at a lower cost. Weโre passing the savings on to you.
๐น Peak/off-peak pricing continues to balance demand.
๐น Off-peak rates are 50% of peak rates. Schedule flexible workloads off-peak to save.
๐น New pricing takes effect at 04:00 UTC on Sept 10, 2026.
๐ณ7
๐ Supporting open source. Expanding deployment options.
Weโll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.
Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Letโs talk.
๐น Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
๐น Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
Weโll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options.
Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Letโs talk.
๐น Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
๐น Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
๐ณ8โค4๐1