AI Model Comparison: Which Family to Use for Which Job

Last updated: August 2026 6 min read

TL;DR: No model family wins every job. Anthropic leads long structured work, OpenAI leads breadth, Google leads non-text input, xAI leads live data, open weights lead cost. Perspective AI runs all of them in one subscription from $14.99/mo, so the choice is per task rather than per bill.

Key Takeaways

Quick Answers

Which AI model is best overall?

No family wins every category, so the useful answer is per job: Anthropic for long structured work, OpenAI for breadth and multimodal output, Google for video and audio input, xAI for live social data, and open-weight models for cost.

How much do AI models cost per million tokens?

Published input rates run from $0.20 to $10.00 per 1M tokens and output rates from $1.20 to $50.00 per 1M tokens across the five vendor rate cards we read on 18 August 2026.

Why does this page not list benchmark scores?

Benchmark tables go stale within weeks and most published figures cannot be traced to a run anyone can reproduce, so this comparison uses vendor-published capabilities and prices, which carry a date and a source.

An AI model comparison is only worth reading if you can tell where its numbers came from. This one carries a date and a source for every figure and refuses to print the ones it cannot stand behind. The short version is that the families have stopped competing on sentence quality and started competing on what they accept as input, what they cost per unit of work, and what tooling surrounds them. Perspective AI is the practical response to that fragmentation. One subscription reaches Anthropic, OpenAI, Google, xAI and the open-weight labs from $14.99/mo, which replaces the per-family plans rather than adding a line to them.

This is the reference page for the comparisons hub, and every head-to-head elsewhere on the site is a narrower cut of the same two tables below.

AI model comparison: which model should you use?

Match the family to the shape of the work, because none of them leads everywhere. Text quality has converged enough that the differences you feel day to day come from input handling, tooling and price.

What is each model family built to be good at?

Each lab has optimised for a different bottleneck, and the resulting specialisations are stable across releases in a way that benchmark rankings are not. The table below describes design intent and shipped capability rather than scores.

FamilyLeads onWeakest atDistinctive capability
ClaudeMulti-file code, long documents, instruction followingImages, voice, consumer conveniencesTerminal agent that edits files and runs tests
GPTBreadth, multimodal output, agents, ecosystemVery large repositories in one passImage generation and voice inside the conversation
GeminiVideo and audio input, long context, Workspace filesThird-party integrations outside GoogleNative handling of hour-long recordings
GrokLive social data, looser registerDeveloper tooling maturityReads the live X stream
Open weightsCost, self-hosting, data controlMultimodal work, agentic reliabilityRuns on infrastructure you own

Read down the "weakest at" column and the case for a single subscription starts to look like a bet that your work will never change shape. Our task-by-task routing guide, which AI model you should use, turns this table into a decision per job.

What does each family cost per 1M tokens?

Published input rates span 50x, from $0.20 to $10.00 per 1M tokens, and output rates span roughly 42x. All figures below were read from the five vendors' own developer documentation on 18 August 2026.

Vendor rate cardCheapest listed inputDearest listed inputOutput range
OpenAI$0.20 per 1M tokens$5.00 per 1M tokens$1.20 to $30.00 per 1M tokens
Anthropic$1.00 per 1M tokens$10.00 per 1M tokens$5.00 to $50.00 per 1M tokens
Google$0.30 per 1M tokens$1.50 per 1M tokens$2.50 to $9.00 per 1M tokens
xAI$1.25 per 1M tokens$4.00 per 1M tokens$2.50 to $12.00 per 1M tokens
Perplexity$0.25 per 1M tokens$3.00 per 1M tokens$2.50 to $15.00 per 1M tokens

Three readings matter more than the individual cells. Every lab now sells a cheap tier and an expensive one, so "which vendor is cheaper" is close to meaningless without naming the tier. The spread inside one vendor is frequently wider than the spread between vendors. And two of the rate cards carry conditions the headline hides: Google publishes a date on which one of its rates doubles, and xAI charges double above a 200,000-token prompt.

Consumer subscriptions are a different unit and are tracked separately in our page on consumer plan prices for every vendor. Translating tokens into something a subscriber can feel is the job of our price per answer index.

Why this comparison publishes no benchmark scores

Because a score you cannot reproduce is a claim, not a measurement. The earlier version of this page carried four tables of benchmark percentages. They are gone, and their removal is an improvement rather than a gap.

Three problems make published model scores unusable for a buying decision. They go stale within weeks, because the labs ship faster than any comparison can be re-run. They are rarely traceable to a run anyone else can repeat, so the number travels between articles without its methodology. And they measure a proxy for the thing you care about, which is why a model can top a coding leaderboard and still be the wrong choice for your repository.

What survives that filter is narrow and useful: capabilities a vendor documents, prices a vendor publishes, context limits a vendor states, and measurements we run and describe ourselves. Everything on this page is one of those four. Where we do publish our own numbers, such as refusal behaviour across model families, the methodology sits next to the figures.

Where one model for everything still beats five

Multi-model access is not automatically the right answer, and treating it as one would be exactly the kind of unfalsifiable claim the section above rejects.

Against those, the case for breadth is simply that most people's weeks are not one shape. If you are reading a model comparison at all, that is weak evidence you have already settled on one. What each single-vendor product feels like at full strength is covered in reviews such as our ChatGPT review.

Five Model Families Versus One Subscription: Routing a Week of Work

Perspective AI is a multi-model app covering the families in the tables above, where the model that answered is named and you can change to another model in the same thread without restating anything. A realistic week looks like four handoffs rather than four subscriptions.

  1. Hand a recording to the family built for audio input and ask what was decided.
  2. Change model in the same thread and have the breadth family turn that into a draft with an image.
  3. Change again and have the long-document family edit it against the original specification.
  4. Push the repetitive cleanup through an open-weight model, where cost per call decides whether the task is worth doing at all.

Starter is $14.99/mo and Pro is $49.99/mo with 700 credits, both carrying the whole catalog. Both are paid plans with no free rung, and the catalogue itself is listed on our model availability page.

Five questions that settle it

  1. Is most of your work code or long documents? Take the Anthropic family.
  2. Do you need one assistant that also makes images and talks? Take the OpenAI family.
  3. Does your raw material arrive as video, audio or Drive files? Take the Google family.
  4. Does your work depend on what was said online this morning? Take the xAI family, as a second model rather than a first.
  5. Is unit cost the binding constraint? Take an open-weight family and read our roundup of the best open-source AI models.

Answering yes to more than two is the ordinary case, not the exception, and it is the entire reason this page exists in the shape it does.

FAQ

Which AI model is best overall?

No family wins every category, so the useful answer is per job: Anthropic for long structured work, OpenAI for breadth and multimodal output, Google for video and audio input, xAI for live social data, and open-weight models for cost.

How much do AI models cost per million tokens?

Published input rates run from $0.20 to $10.00 per 1M tokens and output rates from $1.20 to $50.00 per 1M tokens across the five vendor rate cards we read on 18 August 2026.

Why does this page not list benchmark scores?

Benchmark tables go stale within weeks and most published figures cannot be traced to a run anyone can reproduce, so this comparison uses vendor-published capabilities and prices, which carry a date and a source.

Which model family is cheapest to run?

Open-weight families are cheapest because you can host them yourself, and among hosted options the small and fast tiers from each lab undercut the flagship tiers by roughly an order of magnitude.

Do I need more than one model?

Most heavy users do, because the categories split cleanly and no single vendor leads all of them. A multi-model app such as Perspective AI covers the families in one subscription at $14.99/mo.

Written by the Perspective AI team

Our research team tests and compares AI models hands-on, publishing data-driven analysis across 144+ articles. Perspective AI gives you access to every major AI model in one platform.

One subscription that covers every family in this table

Perspective AI is one subscription across the major labs, $14.99/mo. The model that answered is named, and switchable inside one conversation.

Open Perspective AI →