Quick answer: which AI model is best for a sales chatbot?
DM Champ Max, our custom-tuned model. In our August 2026 test it scored higher on real sales conversations than any other model, including GPT-5.5 and Claude Opus 5, it never made up a fact, and it costs about 2.5 cents per message, everything included. A typical sales conversation on our platform runs to about 9 AI replies, so that's roughly 22 cents per conversation on Max, against $2 to $5 on the premium models. The models that come close cost 3 to 24 times more per message if you run everything a chatbot does on one of them, and still 3 to 5 times more if you push the background jobs onto the cheapest model that held up, and you build everything around them yourself.
If you just want to know what to pick: Max for any business where a wrong answer costs you a sale. Mini if you handle a very high volume of chats and want the lowest price. Keep reading if you want to see the numbers.
Why leaderboard rankings don't tell you which chatbot model is best
Every AI company says their model is "#1 in benchmarks." The problem is what those benchmarks measure. They ask a model trivia questions, one at a time, and count how many it gets right. That tells you almost nothing about running a sales chatbot.
A real sales chat is a customer pushing back on their fifth message. It's someone asking about your prices and your refund policy, which no public test has ever seen. It's a customer writing in Spanish, then switching to English. A model can win every trivia contest and still invent a discount that doesn't exist, or accidentally paste its own internal notes into the chat.
So instead of trusting leaderboards, we tested models the way a sales chatbot actually gets used.
How we tested 17 AI models on real sales conversations
We took twelve real customer conversations in several languages, and replayed each one, six messages deep, with every model playing the part of the chatbot. Every model, ours included, ran through the same pipeline: the same instructions, the same knowledge-base search, the same conversation memory, the same rules, and the same tools (booking, tagging, handing over to a human) that a real DM Champ chatbot has. The only thing that changed from row to row was the model answering.
Then two independent AI judges, the latest Claude Opus and Claude Fable models, read every conversation and scored two things:
- Did the chatbot actually sell? Did it understand what the customer needed, handle objections, and move them towards booking or buying? Scored out of 50. The judge is strict and most of the replayed conversations are mid-funnel questions with no sale on the table, so no model gets near 50; read the column as a ranking, not a grade.
- Did the chatbot make anything up? Every specific claim the bot made (a price, a policy, a feature) was checked against what the business actually says. Anything the business never said counts as an invented fact.
We also priced every model on what a message really costs, not just the reply. A sales chatbot runs about 19 small AI jobs behind every message: picking the right answer from the knowledge base, checking for spam, tagging the contact, pulling out their email, deciding whether to follow up. If you build a chatbot on a model's API yourself, you pay for all of those calls. So our cost column is the all-in price per message handled, at August 2026 list prices.
The results: AI models ranked for sales chatbots
| Model | Sales conversation score (/50) | Invented facts (in ~70 replies) | ~Cost per message, all-in |
|---|---|---|---|
| ★ DM Champ Max | 27.1 — highest in the test | 0 | $0.025 (0.25 credits, flat) |
| Grok 4.5 | 26.2 | 6 | ~$0.22 |
| Claude Opus 5 | 26.1 | 7 | ~$0.60 |
| ★ DM Champ Mini | 25.6 | ~1 | $0.015 (0.15 credits, flat) |
| Qwen 3.8 27B | 25.3 | 3 | ~$0.08 |
| GPT-5.5 | 24.9 | 5 | ~$0.45 |
| GPT-5.2 | 22.9 | 2 | ~$0.17 |
| Kimi K3 | 22.2 | 7 | ~$0.39 |
| Claude Sonnet 4.6 | 20.8 | 4 | ~$0.27 |
| GPT-5-mini | 20.2 | 2 | ~$0.05 |
| Gemini 3.6 Flash | 19.8 | 6 | ~$0.10 |
| Qwen 3.8 Flash | 19.6 | 6 | ~$0.02 |
| Claude Sonnet 5 | 18.5 | 7 | ~$0.27 |
| GLM-5 | 18.1 | 6 | ~$0.07 |
| DeepSeek V4-Pro | 13.6 | 5 | ~$0.08 |
| Claude Haiku 4.5 | 11.8 | 17 | ~$0.10 |
| Llama 4 Maverick | 11.6 | 8 | ~$0.01 |
What the table tells you
Max leads a tight top group. Grok 4.5, Claude Opus 5, Mini, Qwen 3.8 27B and GPT-5.5 all sell well too. What separates them is the other two columns.
Max and Mini are the cheapest things in that group, and the only ones that come ready to use. On a single-model build Grok is nine times the price of Max per message, Opus 5 is 24 times; even with the background jobs moved to the cheapest model that held up, they're still 3 to 5 times Max (the maths is under the background-jobs table). The one other model that's also cheap, Qwen 3.8 27B, is an open-source model you host and wire up yourself, and it still invented 3 facts.
Every other model made things up. Even the best ones invented five to seven facts in about 70 replies: a price here, a feature there. Max invented none. That's not because it runs on a magic model. It's because it's built to stay inside what your business actually says.
Below GPT-5.2, quality falls off a cliff. DeepSeek, Haiku and Llama scored less than half of the leaders. Claude Haiku 4.5 invented 17 facts, more than double any other model.
Newer is not better. Claude Sonnet 5 scored below the older Claude Sonnet 4.6 on the same conversations. If you're picking a model by release date, you're guessing.
What a real conversation costs
Per-message prices are abstract, so here's the same thing per conversation. Across DM Champ, the average sales conversation runs to about 9 AI replies (half of them are 5 or fewer, and 1 in 10 goes past 19). Using that average:
| Model | ~Cost per average conversation (9 AI replies) |
|---|---|
| ★ DM Champ Max | ~$0.22 |
| ★ DM Champ Mini | ~$0.13 |
| Qwen 3.8 Flash | ~$0.18 |
| Qwen 3.8 27B | ~$0.70 |
| Gemini 3.6 Flash | ~$0.90 |
| GPT-5.2 | ~$1.50 |
| Grok 4.5 | ~$1.95 |
| Claude Sonnet 4.6 | ~$2.40 |
| Kimi K3 | ~$3.45 |
| GPT-5.5 | ~$3.95 |
| Claude Opus 5 | ~$5.30 |
A business handling 1,000 conversations a month pays about $220 on Max. The same volume on GPT-5.5 built yourself is about $3,950, on Opus 5 about $5,300, and that's before you've paid anyone to build and maintain it.
Here's the whole field on one map. Up means more reliable, right means more expensive:
Where the cheap AI models break
The bottom of the cost column looks tempting. Here's what those prices actually buy you, from real transcripts in this test:
| Model | What happened |
|---|---|
| Nearly every model | A customer asked to continue the chat on WhatsApp. The bot promised their conversation history would follow them there. The business never said that. It was the most common invention in the whole test. |
| Claude Haiku 4.5 | 17 invented details across the conversations: prices, conditions, features. |
| Qwen 3.8 (both sizes), background jobs | Produced this as a customer-facing reply: <meta_data>Message Date: Friday, March 13…</meta_data>Perfecto. That's internal code that should never be visible. In the same background-job test Qwen Flash returned nothing at all on roughly 1 in 10 calls; the larger 27B version, about 1 in 20. |
| Some open-source "reasoning" models | The entire reply was the model's private thinking, not an answer. |
| Several models, in French | Switched from casual tu to formal vous halfway through, which reads like a different person took over the chat. |
Two kinds of failure, and both cost you money. Confident inventions: the customer now believes something that isn't true, and you find out when they ask for the discount that doesn't exist. Leaked machinery: the customer sees code or internal notes and stops trusting the bot. Neither shows up on a leaderboard. Both show up in your inbox.
A note on Qwen 3.8. The 27B version is the surprise of this test: top-six on sales conversations, only 3 invented facts, and about $0.08 a message. If you're a developer happy to host an open-source model and build everything around it, it's the best bargain in the field. The catch is the row above: in the background jobs it leaked internal tags into replies and the Flash version went silent on 1 in 10 calls. Those are fixable with engineering work. That engineering work is the product.
The 25 jobs behind every message
A chatbot reply is only the part you see. DM Champ runs 25 different AI jobs around it: 19 that run on every message (finding the right answer, spotting spam, tagging, extracting contact details, writing follow-ups, deciding when a chat is over) and 6 that run when you set up a campaign (reading your website and writing the chatbot's entire playbook). We ran every model through all 25, using the real prompts from the product.
Two scores per job: quality (did it make the right call?) and clean format (did it produce something usable, with no leaked code or "sorry, I didn't understand"?). The failure rate is how often the model returned nothing usable at all.
Everyday chat jobs (19 per message, scored out of 10)
| Model | Quality | Clean format | Failure rate | ~Cost per job (API price) |
|---|---|---|---|---|
| ★ DM Champ Max / Mini | 8.9 | 9.7 | 3.7% | included in the message credit |
| GPT-5.5 | 8.9 | 9.8 | 3.3% | ~$0.018 |
| Gemini 3.6 Flash | 8.9 | 9.8 | 5.7% | ~$0.005 |
| Kimi K3 | 8.9 | 9.7 | 4.2% | ~$0.017 |
| Claude Opus 5 | 8.8 | 9.6 | 2.6% | ~$0.026 |
| Grok 4.5 | 8.8 | 9.6 | 3.8% | ~$0.009 |
| GLM-5 | 8.7 | 9.4 | 5.8% | ~$0.003 |
| Qwen 3.8 Flash | 8.7 | 9.7 | 4.9% | ~$0.001 |
| Claude Sonnet 4.6 | 8.6 | 9.4 | 4.6% | ~$0.011 |
| Qwen 3.8 27B | 8.6 | 9.6 | 6.7% | ~$0.0035 |
| GPT-5.2 | 8.5 | 9.7 | 4.4% | ~$0.007 |
| Claude Sonnet 5 | 8.5 | 9.5 | 6.8% | ~$0.011 |
| DeepSeek V4-Pro | 8.5 | 9.7 | 9.6% | ~$0.003 |
| GPT-5-mini | 8.2 | 9.3 | 7.6% | ~$0.002 |
| Claude Haiku 4.5 | 8.2 | 8.9 | 10.3% | ~$0.004 |
| Llama 4 Maverick | 7.4 | 8.4 | 15.1% | ~$0.0005 |
These scores are closer together than the conversation scores, because the easy jobs (is this spam? what's this person's email?) are easy for everyone. The gap opens on the hard, customer-facing jobs. Writing a reply a human would actually send is the hardest job in the whole test, for every model.
The cost column is what one of these small jobs costs on the model's API. Small individually, but a chatbot runs about 19 of them per message, which is how a "cheap" model ends up at the all-in prices in the main table. On DM Champ they're never billed separately: the flat message credit covers the reply and every job behind it.
The sensible mixed build, a premium model for the reply and the cheapest model that held up here (Qwen 3.8 Flash) for the 19 jobs, still lands at about $0.13 a message on Opus 5 or GPT-5.5 and about $0.07 on Grok 4.5. That's 3 to 5 times Max, with two integrations to build and keep running.
Setting up a campaign: writing the chatbot's playbook
This is the big one. The model reads a company's website and writes the complete brain of the chatbot: who it is, what it's trying to achieve, how the conversation should flow, the rules, the FAQs. Every section has to be there, correct and based on what the website actually says.
| Model | Playbook quality (/10) | Pass rate | Failure rate |
|---|---|---|---|
| ★ DM Champ Max | 8.6 | 81% | 0% |
| Claude Sonnet 4.6 | 8.3 | 83% | 0% |
| Kimi K3 | 8.3 | 76% | 0% |
| Grok 4.5 | 7.9 | 70% | 2.5% |
| Gemini 3.6 Flash | 7.9 | 78% | 5% |
| Claude Sonnet 5 | 7.9 | 69% | 2.9% |
| Claude Opus 5 | 7.8 | 68% | 10% |
| GPT-5.2 | 7.8 | 73% | 7.5% |
| GPT-5.5 | 7.8 | 72% | 5% |
| DeepSeek V4-Pro | 7.7 | 69% | 2.5% |
| GLM-5 | 7.5 | 67% | 7.5% |
| Qwen 3.8 27B | 7.5 | 60% | 5.4% |
| Qwen 3.8 Flash | 7.4 | 59% | 5.8% |
| Claude Haiku 4.5 | 7.3 | 55% | 7.5% |
| GPT-5-mini | 7.0 | 41% | 12.5% |
| Llama 4 Maverick | 5.6 | 23% | 32.5% |
Max had the top quality score, ahead of Claude Sonnet 4.6 and Claude Opus 5, with a 0% failure rate. Sonnet 4.6 came closest: a slightly higher pass rate, a lower quality score. Look at the bottom too: the cheap models fail 12 to 33% of the time here, producing playbooks with missing sections or made-up details. The playbook is the whole chatbot. You can't trust a cheap model to write it.
On price: a model reads your whole website and writes a long document here, so a single playbook run costs many times a normal reply on any of the paid models. On DM Champ a campaign build costs the same flat 0.25 credits (about $0.025) as a one-line reply on Max.
Why DM Champ Max wins
Max and Mini are our own custom-tuned models, built for one job: to sell well while staying inside what your business actually says. That's exactly what the invented-facts column measures, and it's where a general-purpose model, however strong, has no particular reason to be good.
It's also why we don't offer a menu of 100 models. Every model breaks differently. GPT-5.5 sells well but writes mediocre playbooks. Claude Sonnet 4.6 writes the best playbooks of the rest but is mid-pack in conversations. Qwen is brilliant when it answers and silent one time in ten. The work that makes one model reliable doesn't carry over to the next. So we picked one system and keep improving it. That's why Max is at the top of the playbook table.
Max or Mini? Mini is the lighter custom-tuned model of the two, tuned for volume. The trade-off is a slightly higher chance of a small slip reaching a customer. Mini for high-volume inboxes where price matters most. Max for anywhere a wrong price, policy or promise costs you a sale.
Every model, ranked for a sales chatbot
- DM Champ Max — our custom-tuned model: the highest sales conversation score in the test, the best playbook writer, zero invented facts, 0.25 credits (about $0.025) per message, flat, with all 25 background jobs included. No API key, no surprise bills.
- DM Champ Mini — the lighter custom-tuned model, a sliver behind Max on conversations at 0.15 credits (about $0.015), flat. The best cheap option in the test by a wide margin.
- Grok 4.5 (xAI) — the most consistent of the rest. About $0.22 per message if you build on it yourself. Not available on DM Champ.
- Claude Opus 5 (Anthropic) — strong at everything and the most reliable on background jobs, but at about $0.60 per message it's too expensive to run on every reply.
- Qwen 3.8 27B (Alibaba, open source) — top-six on conversations, only 3 invented facts, about $0.08 a message. The best bargain in the test if you can host it and are ready to fix the tag leaks and weak playbooks yourself.
- GPT-5.5 (OpenAI) — top six on conversations, middle of the pack on playbooks, about $0.45 per message. Not available on DM Champ.
- GPT-5.2 (OpenAI) — the most honest of the rest (only 2 invented facts) at about $0.17. A good choice if accuracy matters more to you than selling.
- Kimi K3 (Moonshot) — the surprise of the test: excellent on background jobs, a top-three playbook writer. Just not cheap at about $0.39.
- Claude Sonnet 4.6 (Anthropic) — the best playbook writer after Max, mid-pack in conversations, and still better than its own successor. Available on DM Champ's Pro tier (1 credit per message) or with your own Anthropic key.
- Gemini 3.6 Flash (Google) — good at background jobs, weak in conversations. Cheap per word but about $0.10 per message once it's doing everything a chatbot does.
- Qwen 3.8 Flash (Alibaba, open source) — mid-table on conversations with 6 invented facts, but about $0.02 a message. Went silent on 1 in 10 background jobs and leaked tags. Fine for an internal tool, not for customers as-is.
- Claude Sonnet 5 (Anthropic) — worse than Sonnet 4.6 on everything we measured. Don't pay for the version number.
- GLM-5, GPT-5-mini, DeepSeek V4-Pro — middling to poor across the board. DeepSeek fails nearly 1 in 10 background jobs; GPT-5-mini only passes 41% of playbooks.
- Claude Haiku 4.5 (Anthropic) — near the bottom in conversations and the worst fact-inventor in the test, with 17.
- Llama 4 Maverick (Meta) — last on every measure that matters.
How to choose the right AI model for your chatbot
- You handle real customer sales or support chats, and a wrong answer costs you money → Max.
- You handle a very high volume of chats and want the lowest price per message → Mini.
- You specifically want Claude, or want AI usage billed to your own Anthropic account → the Pro tier, or bring your own Anthropic key. Our BYOK vs Max cost guide does the maths.
- You're a developer building on a model's API directly and can afford premium prices → Grok 4.5 for value, GPT-5.5 or Opus 5 for peak quality, GPT-5.2 if you care most about accuracy, Qwen 3.8 27B if you can self-host and want the cheapest strong model.
- You're building an internal tool where a wrong answer costs nothing → any cheap model. Qwen 3.8 Flash is remarkable value when an occasional dropped reply doesn't matter.
Frequently asked questions
Q: What is the best AI model for a sales chatbot in 2026? A: In our test, DM Champ Max. It had the highest sales conversation score of 17 models, made up zero facts, and costs about $0.025 per message with everything included. Among the models you'd build on yourself, Grok 4.5, Claude Opus 5, GPT-5.5 and the open-source Qwen 3.8 27B are the strongest, at 3 to 24 times the price per message on a single-model build, before you've built anything around them.
Q: Is ChatGPT (GPT-5.5) good for a sales chatbot? A: GPT-5.5 sells well: it was in the top five on conversations. But it invented 5 facts in about 70 replies, it was middle of the pack at writing chatbot playbooks, and it costs about $0.45 per message once you count the background work a chatbot does. It's a good model. It isn't a finished chatbot.
Q: Why is Max first and not GPT-5.5 or Claude Opus 5? A: Because choosing a chatbot model comes down to three things: how well it sells, what it costs, and whether you can actually send what it writes. Max has the top score and costs a fraction of the others. With any of the others you still have to build everything around it before it can send a single reply.
Q: Why did a newer Claude model score worse than an older one? A: We don't know why, and it doesn't matter: Claude Sonnet 5 scored below Claude Sonnet 4.6 on the same conversations and the same background jobs. New models are tuned for many things, and your sales inbox may not be one of them. It's the best argument for testing on real conversations instead of trusting release notes.
Q: Which AI model makes up the fewest facts? A: DM Champ Max, with zero in our test. Among the others, GPT-5.2 and GPT-5-mini did best with 2 each. Claude Haiku 4.5 was worst with 17.
Q: Can I use a cheaper model on DM Champ with my own API key? A: DM Champ supports your own key for Anthropic Claude only. There's no need to bring a budget model: Max and Mini already cover the low-cost case, and they come tuned and tested. The pricing page has the credit costs.
How we tested (the details)
- Same conditions for every model. Every model, Max and Mini included, ran with DM Champ's actual instructions, knowledge-base search, conversation memory, rules and tools, on twelve real customer conversations in several languages, replayed six messages deep. Nothing was added or removed for our own models. All writing to real accounts was switched off.
- Two independent judges. The latest Claude Opus and Claude Fable models, each reading every reply against the business's actual instructions and knowledge base.
- 25 background jobs, using the real prompts from the product, graded by Claude Opus on quality and clean format separately. When a model's API returned nothing, that's counted in the failure rate, not as a zero in the quality score.
- Cost is the August 2026 list price of one full sales reply with knowledge-base search, plus 19 background calls at each model's per-call price. Conversation length (about 9 AI replies on average, median 5) is measured across all DM Champ conversations with at least one AI reply in the 90 days to late August 2026.
We re-run this whenever a major model ships.
