DeepSeek V4 just raised its peak input price to ¥9 per million tokens. Zhiyu GLM-5.3 answered with ¥8. This isn't just an AI pricing adjustment—it's a liquidity war, and I've seen this exact pattern play out in DeFi summer 2020 when protocols started competing for TVL with token incentives. The difference? Here, the asset is developer mindshare, and the yield is code generation efficiency.
I've spent the last 24 years in crypto, auditing smart contracts in Mumbai during the ICO mania, and later farming yields on Compound. What I'm seeing now between DeepSeek and Zhiyu is a textbook case of protocol competition disguised as a pricing update. The surface story is simple: DeepSeek hikes peak prices, Zhiyu matches with a slightly better benchmark score. But the real game is being played in the infrastructure layer—caching, off-peak scheduling, and cost structure optimization.
Let me break down the signals.
Context: The Coders Are the New LPs
DeepSeek and Zhiyu are the two dominant Chinese AI model providers. Their API pricing is the new gas fee—except instead of paying for block space, developers pay for token generation. The highest volume use case right now is Coding Agent tasks: autonomous code generation, debugging, and deployment. A single agent session can burn millions of tokens. That's like a DeFi arbitrage bot running on a high-frequency trading loop—every basis point of cost matters.
DeepSeek's V4 model has been the go-to for cost-sensitive developers. Its price increase to ¥9 input / ¥27 output at peak hours is a shock to the system. Zhiyu immediately responded with GLM-5.3 at ¥8 / ¥28, and published a benchmark comparison claiming GLM-5.3 wins 7 out of 9 Agent-related tests. That's a classic yield farming PR move—release a competing product with a higher APY (or lower cost) right when the market leader stumbles.
But here's where the DeFi analogy gets interesting. The real value isn't in the base fee—it's in the caching and off-peak pricing. DeepSeek offers a cache hit rate of ¥0.15 per million tokens during peak hours (¥0.30 for cache), compared to Zhiyu's ¥2. That's a 13x to 1/60 difference. In DeFi terms, this is like having a liquidity pool that charges 0.01% fees for stablecoin pairs while the competitor charges 0.3%. The volume will flow to the cheaper infrastructure, even if the base model is slightly weaker.
Core: Reading the Infrastructure Telegraphed by Pricing
From my experience auditing smart contract codebases, I've learned that the most revealing signals are often hidden in the margins. DeepSeek's cache-to-regular price ratio of 1/60 (peak cache ¥0.3 vs peak input ¥9) is a massive efficiency indicator. It tells me that DeepSeek's KV-cache system is highly optimized—likely achieving reuse rates that allow them to serve cached tokens at near-zero marginal cost. This is akin to a Layer 2 solution that can settle thousands of transactions for the cost of one on-chain.
Zhiyu's cache price of ¥2 vs regular ¥8 is a ratio of 1/4—still decent, but not revolutionary. This suggests their infrastructure is either less optimized or they haven't prioritized cache as a competitive weapon. In the world of crypto, we know that infrastructure wins—the base layer that can handle the most load at the lowest cost becomes the settlement layer of choice. DeepSeek is building that layer for AI.
Furthermore, DeepSeek's off-peak pricing (half price during low-demand hours) is a direct analog to dynamic fee models used by decentralized exchanges. They're using price signals to flatten demand curves, spreading token consumption across time zones. This requires sophisticated demand forecasting and scheduling—a playbook I've seen in DeFi across projects like Uniswap v3's concentrated liquidity or Curve's gauge voting.
Now, let's talk about the benchmark war. Zhiyu's GLM-5.3 claims to lead on 7 of 9 Agent benchmarks, with margins of 2-4 points. I've seen this before in DeFi TVL competitions—one protocol claims higher TVL, but upon closer inspection, they've counted double-counted liquidity or incentivized wash trading. In AI, benchmark contamination is real. The 9 benchmarks are all Agent-focused, deliberately avoiding general knowledge, math, or multilingual tasks. It's like comparing two DeFi protocols only on total value locked without looking at daily active users or revenue.
DeepSeek's performance on Terminal Bench 2.1 (88.2 vs 87.9) and NL2Repo (leading) shows that the gap is marginal. In crypto, a 2% difference in APY doesn't justify the effort of migrating positions—especially when switching costs are high. Developers have to adapt toolchains, retest pipelines, and reconfigure integrations. The cost of migrating from DeepSeek to Zhiyu is likely higher than the ¥1 per million tokens saved.
Contrarian: Price Is Not the Decisive Factor—Switching Costs Are
The common narrative is that DeepSeek's price increase will drive users to Zhiyu's cheaper, stronger model. But I've seen the opposite in DeFi. When a protocol raises fees, users don't immediately flee—they look at the entire ecosystem. DeepSeek has an open-source model, a large developer community, and a deep optimization for the code generation use case. Zhiyu is closed-source and targets enterprise government clients. The lock-in effect is real.
My contrarian take: The real battle is not in the base price, but in the caching infrastructure and the developer tooling. DeepSeek's cache pricing is so aggressive that it will dominate high-reuse scenarios like code autocomplete, template generation, and iterative debugging. Zhiyu's model may win on one-off complex tasks, but those are lower volume. The total token consumption will skew toward cache-friendly applications, giving DeepSeek an edge in total wallet share.
Also, consider the timing. DeepSeek's price increase could be a signal that they are approaching capacity—their GPU clusters are hitting utilization limits. That's a classic supply constraint scenario, similar to Ethereum's gas spikes during NFT mania. Zhiyu's launch right after is opportunistic, but it also reveals that they are not the market leader—they are following the leader's moves. In crypto, the first mover with the best infrastructure often maintains dominance, even if a competitor releases a slightly better token.
Takeaway: Infrastructure Is Permanent, Yields Are Transient
I don't predict trends; I ride the volatility. But the pattern here is clear: the AI API market is entering a phase of margin compression and infrastructure differentiation. The winners will be those who, like the best DeFi protocols, optimize for the long tail of use cases through caching, off-peak scheduling, and elastic scaling.
DeepSeek's caching strategy is its moat. Zhiyu's benchmark lead is a short-term narrative. The ultimate test will be which model delivers the most reliable, low-cost code generation at scale. As I've learned from building DeFi yield strategies: Speed is a feature, not a bug, until it breaks. When the market recovers, the infrastructure that survived the bear will dominate.
For developers and investors watching this space, ignore the benchmark scores. Watch the cache pricing, the off-peak discounts, and the uptime SLAs. That's where the real value lies.
Curation is the new consensus mechanism. Choose your model provider like you choose your L1—not just on price, but on the community, the tooling, and the resilience of the infrastructure underneath.