Step-by-step tutorial Optimization

Best AI model for a chatbot: how to test and compare

Choose a chatbot model with your own questions, a simple scorecard and real side-by-side results instead of relying on model rankings.

Intermediate26 min readJuly 16, 2026
Best AI model for a chatbot: how to test and compare

The best AI model for a chatbot is the one that performs well on your own fixed questions, safety rules, latency target, quota and data-location needs.

There is no universally best chatbot model. A public support bot may need fast, consistent answers at high volume, while an internal assistant may benefit from slower reasoning on complicated questions. The right choice is the least expensive model that reliably meets your quality, speed, tool-use and processing-location requirements.

Do not begin with a public leaderboard. Begin with questions your visitors actually ask and an approved expected answer for each one. General benchmarks help create a shortlist, but only a task-specific comparison shows how a model behaves with your sources, prompt and tools.

This tutorial uses five paraphrases of one known fact from the fictional Northstar Services PDF. The exact fact is NORDSTERN-42. A real WebChatAgent run compared Vertex Gemini 2.5 Flash with Claude Haiku 4.5 while keeping the chatbot, knowledge, prompt and questions unchanged.

Privacy-protected two-click player

Best AI Model for a Chatbot? Real Side-by-Side Test

Find the best AI model for a chatbot by comparing answer quality, safe behavior, latency, quota cost and processing location on identical questions.

YouTube · 3:33 · English

The YouTube player stays blocked until you choose Play. Loading it connects your browser to YouTube and may transfer technical data to Google.

Open directly on YouTube

What you will have at the end

  • A documented shortlist based on current model options
  • A weighted scorecard with hard requirements and pass thresholds
  • One fixed 10–20-question regression set with approved expected answers
  • Measured answer quality and median latency from Model Arena
  • A written keep-or-switch decision including quota and re-indexing impact

Before you start

  • A WebChatAgent chatbot with representative indexed sources
  • 10–20 real questions covering facts, paraphrases, ambiguity and safe non-answers
  • An approved expected answer or scoring rule for every question
  • A maximum acceptable response time and monthly quota budget
  • Any mandatory provider or EU-processing requirement

Best means fit for a measured workload

A fast model is often enough for direct questions whose answers are clearly present in the indexed sources. A smart model can add value when the assistant must combine several passages, follow detailed instructions, call tools or handle nuanced wording. “Smart” is not automatically better for a simple FAQ.

Use hard gates before weighted scoring. If EU processing is mandatory, a non-EU candidate fails even with excellent answers. Among the eligible models, score source fidelity, completeness, instruction following, safe abstention, latency and quota use. Provider changes can require re-indexing because embeddings are provider-specific, so include that operational work in the decision.

Set pass criteriaCompare identical workloadsDocument and monitor

01–05

Set it up step by step

1

Read the current model picker and define the scorecard

Turn business requirements into pass criteria before comparing answers.

Open the intended chatbot and select Configuration. In the AI Model section, open the current model picker. Read the complete label for every realistic candidate: provider, model family, EU marker, speed class and quota multiplier. The captured English Light Mode picker shows fast and smart choices with multipliers such as 4× and 6×.

Write the decision rules before running a test. A practical starter scorecard gives source fidelity 40%, instruction following and safe abstention 25%, response time 20% and quota efficiency 15%. Treat legal processing-location requirements, required tools and a maximum latency as hard pass/fail gates rather than bonus points.

  • Record the exact model labels and test date because the picker can change.
  • Compare only models that satisfy every hard requirement.
  • Do not select a model merely because its name is newer or marked smart.
Turn business requirements into pass criteria before comparing answers.
2

Create a known-answer baseline with the current model

A stronger candidate must beat a complete, repeatable baseline.

Open Knowledge Optimizer → Knowledge Test and run a fixed question set against the current model. Include direct facts, natural paraphrases, one ambiguous request and at least one question the sources cannot answer. Approve the expected result for each question before the run so fluent wording cannot hide a factual error.

The verified example used five natural questions about the code inside the Northstar Services demo PDF. Every answer returned NORDSTERN-42, producing 5 answered, 0 partial, 0 not answered and a 100% score. This is a strong fast-model baseline. A model with a higher quota multiplier must now show a real improvement elsewhere.

  • Keep the sources, source versions, prompt and temperature unchanged.
  • Save the exact questions for regression testing after future model releases.
  • If all candidates miss the same fact, repair the source or retrieval first.
A stronger candidate must beat a complete, repeatable baseline.
3

Run a controlled side-by-side comparison in Model Arena

Change the model candidates and nothing else.

Open the Model Arena tab. Select the current model and one eligible challenger. In the verified run, the candidates were Vertex Gemini 2.5 Flash and Claude Haiku 4.5. Choose “Reuse questions from the last test run” so both models receive the same five NORDSTERN-42 questions.

Check the usage estimate before starting; the interface shows that this comparison consumes about 20 test-chat messages. Run enough representative questions to reduce the influence of one unusually fast or lucky answer. For a production decision, use 10–20 questions and repeat important edge cases.

  • Do not edit the prompt, sources or retrieval settings during the comparison.
  • Use the same language and expected-answer rubric for both candidates.
  • Record transient API errors separately from answer-quality failures.
Change the model candidates and nothing else.
4

Read the complete result: quality, latency, quota and location

The winner must satisfy the whole scorecard, not only answer correctly once.

Wait for the completed result. Compare the aggregate quality score and median latency first, then return to Configuration for the quota multiplier and EU marker. Median latency is more useful than one request because it is less distorted by an unusually slow or fast response.

In the observed run, both models scored 100%. Gemini had 4.0 seconds median latency; Claude had 6.0 seconds. With no quality gain from the challenger, the Arena recommended the current Gemini model for the better quality-speed balance. This is a decision for this workload, not a universal ranking of the two models.

  • Reject any model that fails a hard privacy, provider or tool requirement.
  • Compare quota multipliers at your expected monthly message volume.
  • Repeat latency tests at representative times before setting a strict SLA.
The winner must satisfy the whole scorecard, not only answer correctly once.
5

Inspect individual answers, decide and plan the rollout

A percentage is a summary; the visible answer explains why it earned the score.

Expand at least one correct, one difficult and one failed question. Compare factual content, cited source, missing-information behavior, tone and instruction following. In the captured question, both answers returned NORDSTERN-42. The individual request took 4.6 seconds for Gemini and 6.4 seconds for Claude, matching the aggregate speed direction.

For this measured example, keep the current Gemini model. Select “Use for this bot” only when a challenger wins the agreed scorecard by a meaningful margin. If the switch crosses providers, schedule re-indexing, wait for every source to complete and rerun the full regression set before exposing visitors to the new model. Record the model label, date, score, median latency, quota multiplier and reviewer.

  • Roll out to a non-production assistant first when possible.
  • Monitor negative feedback, unanswered questions and latency after the switch.
  • Keep a fallback decision and the old test result for quick rollback.
A percentage is a summary; the visible answer explains why it earned the score.

Example & result

See the practical test and its result

Every tutorial includes a fixed input, the expected outcome and a transparent record of what was actually verified locally.

Practical example: How to choose the best AI model for your chatbot

This exact scenario was completed with the temporary tutorial account.

Verified end to end

Exact test input

Reuse the five NORDSTERN-42 PDF questions for Vertex Gemini 2.5 Flash and Claude Haiku 4.5.

Expected result

Both models return the correct fact; measured latency, quota and constraints determine whether a switch is justified.

What was actually verified

Both candidates scored 100%. Gemini had 4.0 s median latency versus 6.0 s for Claude; the expanded answer took 4.6 s versus 6.4 s and both returned NORDSTERN-42.

Both candidates scored 100%. Gemini had 4.0 s median latency versus 6.0 s for Claude; the expanded answer took 4.6 s versus 6.4 s and both returned NORDSTERN-42.

Tips & tricks

Make the setup reliable

Test with realistic examples, record your baseline and change one setting at a time. That makes real improvements visible.

Source quality dominates many model differences

If every model misses the same fact, inspect extraction, duplicate sources and retrieved passages before paying for a smarter model.

Include safe non-answer questions

A useful test set asks about facts that are not in the knowledge base. The winning model should decline safely instead of producing a plausible invention.

Estimate quota at production volume

Multiply expected monthly messages by the displayed quota factor. A small per-message difference becomes material at high traffic.

Re-evaluate after meaningful changes

Reuse the same scorecard after a major model release, prompt revision, tool rollout or source migration. Do not restart with a different question set.

When something does not work

Troubleshooting

Check status, permissions and test data systematically before changing the model or prompt.

Both models answer the same fact incorrectly

Search the indexed content for the expected fact, remove outdated duplicates and rerun Knowledge Test. A model comparison cannot repair absent or conflicting evidence.

Model Arena cannot reuse the previous questions

Complete and save a Knowledge Test for the same chatbot first. Then return to Model Arena and select the last test run without changing assistants.

Latency changes sharply between repeated runs

Increase the question count, compare medians and rerun at representative times. Record provider errors separately and avoid deciding from one request.

The desired candidate is unavailable or disabled

Check the current provider and plan requirements in Configuration. Use only candidates shown for this chatbot and document the exact picker state and date.

Answers become unavailable after a provider switch

Re-index every active source with the new provider-compatible embeddings, wait for Completed status and rerun the fixed regression set before launch.

Ready for a production-style test

Save the scorecard, exact questions, current model label, date, quality score, median latency, quota factor and final decision in one record. Monitor negative feedback, unanswered questions and real response time after launch, then reopen the choice only when production evidence or a material model release justifies another controlled test.

Related resources