Updated Aug 31, 2026 · Aug 25–31, 2026 · 240 models · 176 ZDR+NPT

Best models on AI Gateway for Aug 25–31, 2026

This week (Aug 25–31, 2026), the best ZDR + no-training pick on AI Gateway isDeepSeek V4 Flashat$0.113 / 1Mondeepinfra31% off. The frontier pick isGPT 5.6 Sol.

Independent ranking from live catalog, 7-day adoption, discounts, and DeepsecBench. Snapshot of the weekly ranking. See the current week.

Best ZDR + no-training models this week

2

Zero data retention and no training on prompts

BangWorkhorseCheap

DeepSeek V4 Flash

16.5 Deepsec3.06 bang22.1% tokens31% off
Frontier

GPT 5.6 Sol

35.6 Deepsec1.40 bang

Best bang-for-buck models (including models that train)

2

Includes models that train or skip ZDR

BangWorkhorseCheap

DeepSeek V4 Flash

ZDR + no training16.5 Deepsec3.06 bang22.1% tokens31% off
Frontier

GPT 5.6 Sol

ZDR + no training35.6 Deepsec1.40 bang

Ranked AI Gateway models

ZDR + no training, ranked

ModelBlendTokens
DeepSeek V4 Flash· 31% off$0.11322.1%
Qwen 3.8 Max$3.000.0%
GPT 5.6 Terra$4.500.0%
GPT 5.6 Luna$0.4506.0%
Kimi K3· 7% off$5.600.0%
GLM 5.3$2.150.0%
GPT 5.6 Sol$11.250.0%
Claude Opus 5$10.005.2%
GPT 5.5$11.250.0%
Claude Sonnet 5$4.004.1%

Which labs developers actually use

10
Token share vs spend
deepseek33.1% tok3.9% $
anthropic17.8% tok60.4% $
openai12.2% tok16.9% $
stepfun10.2% tok2.0% $
zai8.8% tok4.6% $
minimax6.7% tok0.3% $
google5.2% tok7.5% $
moonshotai2.2% tok3.4% $
meta1.5% tok0.2% $
xiaomi0.9% tok0.0% $

Frequently asked questions

What is the best AI Gateway model this week?
This week (Aug 25–31, 2026), the best ZDR + no-training pick on AI Gateway is DeepSeek V4 Flash at $0.113 / 1M blended on deepinfra (31% off). The frontier pick is GPT 5.6 Sol.
What does ZDR + no training mean?
ZDR is zero data retention: the provider does not keep prompts. No training means prompts are not used to train models. A privacy pick requires the AI Gateway catalog to mark both `zdr` and `no_training` as all or some, matching the official ?zdr=true and ?npt=true filters.
How is bang-for-buck calculated?
Bang-for-buck is DeepsecBench score divided by that run's cost. Value score is 7-day mean token share divided by blended $/1M, using a 3× input + 1× output mix. Discounted picks use a cheaper ZDR endpoint than list price.
Which models count as capable?
Capable models support tool use and have at least 128,000 tokens of context. Vision is not required. Cheap routers must also blend at or under $0.50 / 1M; workhorses at or under $3.00 / 1M.
Is this an official Vercel product?
No. bestmodels.dev is an independent weekly ranking. Catalog, adoption, and DeepsecBench numbers come from Vercel AI Gateway data licensed CC BY 4.0. We are not affiliated with Vercel.