Worth reading
DeepSeek raised prices and made the clock part of the bill
DeepSeek raised V4 API prices on August 16, 2026 and split every rate into peak and off-peak halves. V4-Flash output went from $0.28 to $1.32 during peak hours.
Verified records, your tests, and real cost — one clear decision.
One clear model decision
Three steps. No universal winner.
Worth reading
DeepSeek raised V4 API prices on August 16, 2026 and split every rate into peak and off-peak halves. V4-Flash output went from $0.28 to $1.32 during peak hours.
New and changed
DeepSeek raised V4 API prices on August 16, 2026 and split every rate into peak and off-peak halves. V4-Flash output went from $0.28 to $1.32 during peak hours.
Gemini 3.7 Flash went GA on August 13, 2026 at $0.75/$3.75 per 1M tokens, the same rate 3.6 Flash was cut to. Both return to $1.50/$7.50 on January 1, 2027.
Z.AI's API release adds post-training gains, always-on reasoning, and three compatible protocols at $1.40/$4.40.
Google's stable multimodal model starts at $0.75/$3.75 through 2026, with a 1M-token window and a dated price change.
Grok 4.6 targets long coding agents, starts at $2/$6, doubles at 200K prompt tokens, and has no numeric text-output cap.
Grok 4.5 review: xAI's coding specialist, built with Cursor, priced at $2/$6, with a smaller 500K context window than Grok 4.3 despite what some blogs claim.
GLM-5.2 review: Zhipu's MIT-licensed, open-weight 753B model at $1.40/$4.40 per million tokens, with coding benchmarks Z.AI says beat GPT-5.5 and Opus 4.7.
OpenAI Presence bundles voice and chat agents with policies, evaluations, approved actions, and escalation. What buyers know—and what remains undisclosed.
Anthropic's current Opus at the same $5/$25 rate — what changes over 4.8, and who should move.
A candidate for visual and structured-output work based on OpenAI's published positioning; validate it on your own prompts.
A retired multimodal model that Google deprecated in March 2026 in favor of Gemini 3.1 Pro.
Cursor, GitHub Copilot, Windsurf, and Cody on the same Markdown-exporter task.
ElevenLabs, OpenAI Whisper, and Cartesia Sonic on latency, accuracy, and naturalness.
Advertised context capacity versus reproducible retrieval checks you can run on your own documents.
The cheapest model for chat, coding, RAG, agents, classification, and summarization.
Where Llama 4, Mistral Large 3, DeepSeek-V4, and Qwen 3.6 stand against the closed labs.
MMLU is saturated above 90%. The benchmarks worth tracking now.
Claude Opus 4.7, GPT-5, and Gemini 3.1 Pro Preview: sourced specifications, editorial tradeoffs, and a local evaluation plan.
Llama 4, Mistral, DeepSeek, Qwen, the small-model tier, and what it takes to self-host.
What AI costs by model and workload, and where teams overspend.
An interactive comparison covering pricing, benchmarks, context windows, and capability ratings for all 33 models in the shared index, across frontier, mid, and open-weight tiers.
All 33 models ranked by benchr Rating →
All 113 articles in the archive →
Provider guides: OpenAI, Anthropic, Google, and open weights →
benchr is an evidence-led reference, not a generic listicle. Figures are identified as official provider data, third-party benchmark results, or benchr editorial estimates so readers can judge the evidence behind each claim.