DeepSeek V4 API prices rise up to 1,100% today as peak-hour billing replaces flat rates
DeepSeek's V4-Flash and V4-Pro APIs more than quadrupled at peak hours starting at 16:00 UTC today, ending the flat-rate pricing that forced a sector-wide cost collapse at the start of 2026. Increases range from 50% to more than 1,100% depending on model, token type, and time of use, per Quartz.
What changed
The flat rates that made DeepSeek the cheapest major inference provider are gone. Two billing tiers now apply: peak hours run 01:00 to 04:00 UTC and again 06:00 to 10:00 UTC. Off-peak rates are half the peak price for each model and token type, according to DeepSeek's pricing documentation.
For V4-Flash output tokens, the change is most stark. The flat $0.28 per million tokens becomes $1.32 per million at peak and $0.66 per million off-peak, per Quartz. Cache-miss input tokens for V4-Flash move from $0.14 to $0.44 at peak.
The V4-Pro changes are proportionally similar. Output tokens go from $0.87 per million to $3.96 at peak and $1.98 off-peak. Cache-miss input tokens for V4-Pro climb from $0.435 to $1.32 at peak, per Quartz. Cache-hit rates, which are already priced far below cache-miss, take the steepest percentage hits at peak hours, which is where the 1,100% headline figure originates.
DeepSeek said the purpose is to allocate resources "more reasonably," and the company described the tiered structure as a tool to steer developer workloads toward less congested periods, per Fortune.
Why it matters
This is not a typical vendor price adjustment. DeepSeek's original V4 pricing did something unusual: it set a rate so low that Chinese rivals felt immediate competitive pressure to match it. After DeepSeek made deep discounts on V4-Pro permanent in May, ByteDance and Tencent cut their own API rates in kind. AI Weekly described DeepSeek as the vendor that "set the floor" for Chinese AI inference per AI Weekly. The floor just rose significantly.
For developers, the practical shock is in agent-loop workloads scheduled during peak hours. V4-Flash output at $1.32 per million during peak costs 4.7 times what it cost yesterday. A team spending $1,000 per month on V4-Flash output at flat rates could face $4,700 in peak-hour costs if usage patterns do not shift.
The off-peak discount offers a real lever: move batch jobs to off-peak windows and V4-Flash output drops to $0.66 per million, which is 2.4 times the old flat rate rather than 4.7 times. Whether that trade-off works depends on whether the application has latency tolerance. Agent pipelines, nightly report generation, and dataset processing jobs can likely shift. Real-time customer-facing calls cannot.
Competitive positioning still favors DeepSeek. Even at peak rates, V4-Pro output at $3.96 per million remains below what Anthropic's frontier model charges, per Fortune. DeepSeek built its developer base on a large cost advantage. That advantage narrows at peak hours but does not disappear.
Context
DeepSeek announced on August 6 that a "significant" price increase was coming but gave no amounts or date. Bloomberg first reported that notice, and the AI Weekly writeup from that day captured the dynamic well: a company whose V4-Flash was running at roughly 3 cents per benchmark test when it launched in late July was signaling it wanted to raise the floor it set. The August 16 date and the specific rate table came on August 13, via DeepSeek's pricing page update, per Quartz.
The July 31 public beta of DeepSeek-V4-Flash-0731 preceded the pricing notice by six days. The changelog for that release noted that "the official release of DeepSeek-V4-Pro will follow soon," per DeepSeek's own change log. A V4-Pro commercial rate structure now exists, which may signal that a full launch is close.
The deeper context is financial. DeepSeek closed a funding round exceeding $7 billion and has begun IPO preparations, per Quartz. Investors scrutinizing unit economics at an AI lab that charges a few cents per million tokens have a different set of questions than ones backing a lab that charges dollars. The pricing shift tracks with that transition. Founder Liang Wenfeng now has to balance the low-cost developer positioning that drove DeepSeek's rapid adoption against the commercial expectations of external capital.
The peak/off-peak model itself is not novel. AWS, Google Cloud, and Azure all apply time-differentiated pricing to compute-intensive services. What is notable is that DeepSeek, which positioned itself as the answer to those providers' costs, is now using the same billing instrument.
What to watch next
The first signal to track is whether ByteDance, Tencent, and Moonshot hold their discounted rates to capture developers who no longer find a bargain at DeepSeek's peak pricing. If those rivals hold, the "DeepSeek death zone" that analysts described for midrange models may narrow. If they follow DeepSeek up, the whole Chinese AI inference market becomes more expensive in parallel.
The second is whether OpenAI or Anthropic cite the shift as context in their own pricing communications. Before today, any move toward higher US lab pricing ran directly against the reference point that DeepSeek represented. That reference point has moved.
Sources
- DeepSeek is raising AI developer access prices by up to 1,100% starting Sunday - Quartz, August 13, 2026
- DeepSeek Models and Pricing - DeepSeek API Docs (vendor)
- DeepSeek increases prices for AI services by multiple times - Fortune/Bloomberg, August 13, 2026
- DeepSeek Warns Developers of 'Significant' API Price Hike - AI Weekly, August 6, 2026
