Models

ChatGPT, Claude, Gemini - all under one roof.

Build your support agent on 20+ top AI models and switch between them anytime. We benchmark every one on scripted support conversations, so you pick on evidence - score, consistency, cost - not on the vendor's claims.

  • 20+ top-tier models
  • Switch anytime
  • Benchmarked for support

Try it live. Ask about models

Loading the assistant…

SupportBench

Which model is actually best at customer support?

We run every model through 31 scripted support conversations with identical knowledge and tools, then score them on grounding, policy, tool use, consistency and cost. Our data, not the vendors' claims.

#Model SupportBench score 0-100. The mean of two LLM graders from different vendors, each grading eight dimensions against a written answer key - after deterministic checks, which zero any conversation with a wrong refund, a data leak or a claimed action the tool never did. Tiebreaker The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score. Consistency 100 minus the average swing between repeated runs of the same scenario. 100 = identical handling every time; a model at 80 can score 100 on one run and 60 on the next. Mistake cost Failed checks per 100 conversations, weighted by what they cost a business: money 25, privacy 20, trust 10, inconvenience 3. Lower is better. Hard fails Share of conversations zeroed by a deterministic check: an unauthorised refund or credit, private data disclosed, or a claim of an action the tool never performed.
1
88.6
95% 86.2–90.9
#344% wins89.670.6%
2
86.5
95% 80.9–91.2
#161% wins85.7873.9%
3
86.0
95% 80.5–90.8
#245% wins85.1473.2%
4
82.5
95% 75.7–87.8
80.71435.8%
5
81.7
95% 74.2–88.3
84.31085.8%
6
GPT-4.1
OpenAI
67.4
95% 56.1–77.5
77.941817.4%
7
51.7
95% 40.1–63.9
76.354727.1%

The tiebreaker: splitting the top three

The top three finish within each other's error bars, so graders compared their transcripts of the same conversations side by side and picked the one they would rather have sent.

61%of decided matchups won The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score.
vs Sonnet 5 68W–44L–38T vs Gemini 3.7 Flash 69W–45L–36T rating 1536 (1498–1576) · P(1st) 89% Share of 1,000 scenario-resampled bootstrap draws in which this model came out top of the head-to-head ranking. Read it as 'how confident the ranking is in this model being first'.
45%of decided matchups won The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score.
vs Grok 4.6 44W–68L–38T vs Gemini 3.7 Flash 57W–56L–37T rating 1482 (1440–1526) · P(1st) 6% Share of 1,000 scenario-resampled bootstrap draws in which this model came out top of the head-to-head ranking. Read it as 'how confident the ranking is in this model being first'.
44%of decided matchups won The tiebreaker. The top models finish within each other's error bars on the main score, so to split them a grader is shown two models' transcripts of the same conversation side by side and asked which it would rather have sent to the customer. Every pair is judged in both orders (an answer that flips with the order counts as a tie). Win % counts decided matchups only; the rating is a Bradley-Terry fit on an Elo-like scale where 1500 = the average of the compared models. Models outside the top band are not compared - their order is already settled by the main score.
vs Grok 4.6 45W–69L–36T vs Sonnet 5 56W–57L–37T rating 1482 (1436–1524) · P(1st) 5% Share of 1,000 scenario-resampled bootstrap draws in which this model came out top of the head-to-head ranking. Read it as 'how confident the ranking is in this model being first'.

Only the top three are compared: the next model, GPT-5.6 Luna, is already 3.5 points off the band on the main score, so the order below them is settled without a tiebreak. 450 matchups over 25 scenarios × 3 repeats, each judged in both orders by 2 graders from different vendors; 9% counted as ties because the grader flipped with the order.

Best at your price point

Budget

under $0.003 per resolved conversation

GPT-5.6 Luna82.5 · $0.0014 / resolved

Also in this tier: GLM 5.3 Flash (82), GPT-4o mini (52)

Mid-range

$0.003 – $0.01

Gemini 3.7 Flash88.6 · $0.0035 / resolved

Premium

over $0.01

Grok 4.686.5 · $0.0149 / resolved

Also in this tier: Claude Sonnet 5 (86), GPT-4.1 (67)

Quality vs costSupportBench score against cost per resolved conversation (log scale). Top-left is best.
405060708090100$0.001$0.01 Cost per resolved conversation (USD, log scale) SupportBench score Gemini 3.7 Flash88.6 · $0.0035xGrok 4.686.5 · $0.0149Claude Sonnet 586.0 · $0.0247GPT-5.6 Luna82.5 · $0.0014ZGLM 5.3 Flash81.7 · $0.0005GPT-4.167.4 · $0.0133GPT-4o mini51.7 · $0.0014

Gemini 3.7 Flash · Grok 4.6 · Claude Sonnet 5 · GPT-5.6 Luna · GLM 5.3 Flash · GPT-4.1 · GPT-4o mini

Rank by what you care about

Pure SupportBench score. Cost ignored.

  1. 1Gemini 3.7 Flash88.6
  2. 2Grok 4.686.5
  3. 3Claude Sonnet 586.0
  4. 4GPT-5.6 Luna82.5
  5. 5GLM 5.3 Flash81.7
  6. 6GPT-4.167.4
  7. 7GPT-4o mini51.7

Value = SupportBench score − weight × log₁₀(cost per resolved conversation ÷ cheapest model). Greyed-out models fall below the preset's quality floor. The score column on every page is always the pure quality number; this only changes the order.

By scenario category

Grounding Conflicting or incomplete sources, arithmetic spread across documents, questions the docs genuinely don't answer.
  1. Claude Sonnet 589
  2. Grok 4.687
  3. GLM 5.3 Flash85
  4. Gemini 3.7 Flash85
  5. GPT-5.6 Luna79
  6. GPT-4.164
  7. GPT-4o mini53
Tool use Lookups, refunds and credits with exact amounts, tools that return nothing or fail, data the customer claims that the record contradicts.
  1. Grok 4.690
  2. GPT-5.6 Luna86
  3. Gemini 3.7 Flash86
  4. Claude Sonnet 584
  5. GLM 5.3 Flash80
  6. GPT-4.159
  7. GPT-4o mini45
Policy Pressure for out-of-policy refunds, rules that must hold across a long conversation, channel constraints like SMS length limits.
  1. Grok 4.694
  2. Gemini 3.7 Flash88
  3. GPT-5.6 Luna87
  4. Claude Sonnet 584
  5. GLM 5.3 Flash79
  6. GPT-4.158
  7. GPT-4o mini41
Multi-turn Customers who change their mind, raise two issues at once, or get angry about something that has a simple fix.
  1. Gemini 3.7 Flash92
  2. Claude Sonnet 592
  3. GPT-5.6 Luna91
  4. GLM 5.3 Flash89
  5. Grok 4.686
  6. GPT-4.184
  7. GPT-4o mini78
Safety Prompt injection hidden in retrieved content, polite social engineering, and private data a tool returns that policy forbids sharing.
  1. Gemini 3.7 Flash94
  2. Claude Sonnet 570
  3. GLM 5.3 Flash60
  4. GPT-4.159
  5. Grok 4.657
  6. GPT-5.6 Luna47
  7. GPT-4o mini0

7 models benchmarked · how SupportBench works and full results →

Why choice matters

One support agent, any model under the hood

Chat Thing is not tied to one lab. Run your support agent on whichever model the data says is best for you - and change your mind later.

The right model for each job

Your billing bot and your product-docs bot do not need the same model. Pick per bot: the cheapest strong model for volume, the most careful one where mistakes cost money.

Switch any time, no re-training

Your knowledge base, prompts and tools stay exactly as they are. Changing model is one dropdown - and when a better model lands, it is in the list the same week.

Tested, not just listed

We run every model through SupportBench before recommending it: the same support conversations, the same knowledge, the same tools. The numbers above are ours, not the vendors' claims.

Build a support bot and try the models on your own content

Free to start. Add your help centre, pick a model from the list below, and compare answers on the questions your customers actually ask.

Every model

All models available in Chat Thing

Anthropic

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Claude Haiku 3 200,000x 0.5x 2.5
Claude Haiku 4.5 200,000x 2x 10
Claude Sonnet 5 1,000,000x 4x 20
Claude Sonnet 4.5 1,000,000x 6x 30
Claude Sonnet 4.6 1,000,000x 6x 30
Claude Sonnet 4 1,000,000x 6x 30
Claude Opus 4.6 1,000,000x 10x 50
Claude Opus 4.7 1,000,000x 10x 50
Claude Opus 5 1,000,000x 10x 50
Claude Opus 4.5 200,000x 10x 50
Claude Opus 4.8 1,000,000x 10x 50
Claude Opus 4.8 (Fast) 1,000,000x 20x 100
Claude Opus 5 (Fast) 1,000,000x 20x 100
Claude Opus 4 200,000x 30x 150
Claude Opus 4.1 200,000x 30x 150
Claude Opus 4.7 (Fast) 1,000,000x 60x 300

Cohere

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Cohere - Command R 128,000x 0.3x 1.2
Command A 256,000x 5x 20
Cohere - Command R+ 128,000x 5x 20

DeepSeek

Name
Context window
Input modifier
Output modifier
Power ups
Vision
DeepSeek V4 Flash 1,048,576x 0.28x 0.56
DeepSeek V3 163,840x 0.51x 2.06
DeepSeek V4 Pro 1,048,576x 0.87x 1.74
DeepSeek R1 64,000x 1.4x 5

Google

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Google - Gemini 2.5 Flash Lite 1,048,576x 0.2x 0.8
Gemini 3.1 Flash Lite 1,048,576x 0.5x 3
Gemini 3.5 Flash Lite 1,048,576x 0.6x 5
Google - Gemini 2.5 Flash 1,048,576x 0.6x 5
Gemini 3.7 Flash 1,048,576x 0.75x 3.75
Google - Gemini 3 Flash 1,048,576x 1x 6
Google - Gemini 2.5 Pro 1,048,576x 2.5x 20
Gemini 3.6 Flash 1,048,576x 3x 15
Gemini 3.5 Flash 1,048,576x 3x 18
Google - Gemini 3.1 Pro 1,048,576x 4x 24

Meta

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Llama 4 Scout 327,680x 0.2x 0.6
Llama 4 Maverick 1,048,576x 0.4x 1.6

Mistral

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Mistral - Mistral Small 32,768x 0.1x 0.16
Mistral Large 3 262,144x 1x 3
Mistral Medium 3.5 262,144x 3x 15
Mistral - Open Mixtral 8x22b 65,536x 4x 12
Mistral - Mistral Large 128,000x 4x 12

MoonshotAI

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Kimi K2 131,072x 1.14x 4.6
Kimi K2.6 262,144x 1.9x 8
Kimi K3 1,048,576x 6x 30

OpenAI

Name
Context window
Input modifier
Output modifier
Power ups
Vision
GPT-5 Nano 400,000x 0.1x 0.8
GPT-5.6 Luna ourdefault1,050,000x 0.2x 1.2
GPT-5.6 Luna Pro 1,050,000x 0.2x 1.2
GPT-4.1 Nano 1,047,576x 0.2x 0.8
GPT-4o Mini 128,000x 0.3x 1.2
GPT-5.4 Nano 400,000x 0.4x 2.5
GPT-5 Mini 400,000x 0.5x 4
GPT-4.1 Mini 1,047,576x 0.8x 3.2
GPT-3.5 Turbo 16,385x 1x 3
GPT-5.4 Mini 400,000x 1.5x 9
GPT-5.6 Terra 1,050,000x 2x 12
GPT-5.6 Terra Pro 1,050,000x 2x 12
GPT-5.1 400,000x 2.5x 20
GPT-5.1 Chat 128,000x 2.5x 20
GPT-5 400,000x 2.5x 20
GPT-5.2 400,000x 3.5x 28
GPT-4.1 1,047,576x 4x 16
GPT-5.4 1,050,000x 5x 30
GPT-4o 128,000x 5x 20
GPT-5.6 Sol 1,050,000x 10x 60
GPT-5.6 Sol Pro 1,050,000x 10x 60
GPT-5.5 1,050,000x 10x 60
GPT-4 Turbo 128k 128,000x 20x 60
GPT-5 Pro 400,000x 30x 240
GPT-5.2 Pro 400,000x 42x 336
GPT-4 8,191x 60x 120
GPT-5.5 Pro 1,050,000x 60x 360
GPT-5.4 Pro 1,050,000x 60x 360

Perplexity

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Sonar 127,072x 2x 2
Sonar Pro 200,000x 6x 30

Qwen

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Qwen3.7 Flash 1,000,000x 0.06x 0.26

xAI

Name
Context window
Input modifier
Output modifier
Power ups
Vision
Grok 4.3 1,000,000x 2.5x 5
Grok 4.20 2,000,000x 2.5x 5
Grok 4.6 500,000x 4x 12
Grok 4.5 500,000x 4x 12

Z.ai

Name
Context window
Input modifier
Output modifier
Power ups
Vision
GLM 4.7 202,752x 0.8x 3.5
GLM 4.6 202,752x 1x 4
GLM 4.5 131,000x 1.2x 4.4
GLM 5.1 202,752x 1.93x 6.07
GLM 5.2 1,048,576x 1.93x 6.07