One detail changed the comparison: the three big consumer assistants now cost almost exactly the same. ChatGPT Plus is $20 a month. Claude Pro is $20 a month, or about $17 if you pay for the year. Google AI Pro is $19.99. So the old tiebreaker, price, is gone. You're not choosing a cheaper plan. You're choosing a default model and the app it lives in.
And the defaults are where this gets interesting, because the three companies point you at very different models the moment you open the box.
| Assistant | Paid plan | Default model | What the free tier gives you |
|---|---|---|---|
| ChatGPT | $20/mo (Plus) | GPT-5.5 Instant | GPT-5.5 with a roughly 10-message-per-5-hours cap, then a smaller fallback model |
| Claude | $20/mo, $17/mo annual | Claude Sonnet 4.6 | Sonnet 4.6 with web search, memory, and voice; usage capped per session |
| Gemini | $19.99/mo (Google AI Pro) | Gemini 3.5 Flash | Flash chat plus a daily allotment of Gemini 3.1 Pro, image generation, and a few deep-research reports |
Defaults, quotas, and included tools change faster than annual comparison pages. Verify each live plan page before paying. The four sections below are reproducible evaluation designs, not a record of private benchr runs or a hidden scoreboard.
Use the same input, a pinned date, and an explicit rubric. Hide the assistant name from reviewers where possible, record factual errors and revision time, and repeat enough examples to avoid choosing from one lucky answer.
Task one: a quick answer you can check
Ask each assistant the same current-information question and require links. Score whether every citation opens, supports the adjacent claim, has an appropriate date, and avoids omitted caveats. Measure verification time rather than link count alone. The AI search guide provides a fuller citation rubric.
Decision rule: choose the assistant with the highest citation-validity rate on your topics, not the one with the most confident prose.
Task two: drafting an email or message
Give all three the same awkward reply, tone guide, prohibited promises, and length limit. Blind-score instruction compliance, factual additions, policy risk, tone, and edit distance from the version you would send. Anthropic's positioning makes Claude a reasonable candidate, but positioning is not an independent result.
Decision rule: choose the lowest reviewed revision cost across a representative set of your messages.
Task three: explain something so it sticks
Prepare several questions with expert-verified answer keys, specify the learner's level, and ask each assistant to separate fact, analogy, and uncertainty. Score coverage, factual accuracy, calibration, and whether the explanation creates a misleading simplification.
Decision rule: let subject-matter review decide; never infer medical, legal, or financial reliability from an engaging explanation.
Task four: images and a casual creative spin
First check which image tools and quotas each live plan actually includes. Then reuse one brief and score text fidelity, composition, accessibility, editing controls, export quality, and the number of retries. Video is a separate workflow covered in the AI video guide.
Decision rule: compare the accepted image cost and edit time, not a one-off attractive sample or an old free-tier quota.
Current answers
Verify Citation validity and freshnessDrafting and messages
Blind review Compliance and edit distanceExplaining to learn
Expert key Accuracy and calibrationImages and creative
Live-plan test Accepted cost and edit timeThese four rubrics do not add up to a universal score. Weight them by your actual workload and keep the raw examples so another reviewer can reproduce the decision.
Its documented tools and interface match your workflow. Verify the current default model and plan limits, then score it on the same held-out tasks as the alternatives.
Writing, editing, and instruction-bound replies dominate your workload. Blind-review tone, factual restraint, and revision effort instead of relying on model reputation.
Google Workspace integration or the plan's current search and media tools matter to you. Confirm availability and quotas on Google's live plan page before assigning value.
Do not assume one or two subscriptions is automatically right. Start with bounded trials, include review time in total cost, and keep a second paid plan only when it produces distinct measured value. Re-run the sample when a default model or plan changes. For underlying model evidence, see Opus 4.8 vs GPT-5.5 and the Gemini lifecycle review.