AI chatbot Knowledge Optimizer: complete tutorial
Use one fixed question set to diagnose the failing layer, make one controlled change and prove whether the answer actually improved.

The AI chatbot Knowledge Optimizer turns a vague complaint about a weak answer into a repeatable test across sources, prompt behavior and model choice.
A weak answer does not automatically mean the AI model is bad. The approved fact may be missing, retrieval may select the wrong passage, the system role may encourage unsupported guesses, or the chosen model may simply be slower without producing a better answer.
This beginner-friendly workflow uses a fixed five-question PDF test as the known-good baseline and a deliberately unsupported 2028 warranty question as the failure case. You will inspect Knowledge Test, Optimize, Role, Model Arena and Audit, change only the layer supported by evidence and reject an AI-generated role proposal when the A/B test proves it is worse.
The important habit is to separate retrieval from answer behavior. First confirm what evidence was found. Then judge whether the answer is accurate, grounded, complete enough and safely refuses missing information. Only after that should latency, model cost or writing style decide between otherwise correct candidates.
Privacy-protected two-click player
AI Chatbot Knowledge Optimizer Tutorial: Fix Weak Answers
Use the AI chatbot Knowledge Optimizer to test weak answers, inspect retrieval, review suggestions, compare models and reject changes that score worse.
The YouTube player stays blocked until you choose Play. Loading it connects your browser to YouTube and may transfer technical data to Google.
Open directly on YouTubeWhat you will have at the end
- A reproducible five-question knowledge baseline
- A known-answer and safe-no-answer regression pair
- A clear decision between source, retrieval, role and model changes
- A verified A/B result instead of an untested AI suggestion
- A practical review routine for Optimize and Audit
- A launch gate that includes quality, latency and regression checks
Before you start
- An indexed knowledge base with one uniquely verifiable fact
- One weak or unanswered real question
- The approved source owner or reviewer for factual changes
- Enough message quota for repeated tests, A/B validation and model comparison
- A saved copy of the current role, source version and model settings
Optimize the failing layer
Use Knowledge Test to reproduce the answer and inspect the evidence. Use Optimize for content gaps found in real conversations, Role for response behavior, Model Arena for a controlled model comparison and Audit for a wider knowledge-base review. The Optimization Inbox keeps every suggestion reviewable instead of changing production knowledge silently.
Diagnose the failing layer before editing: missing approved fact means source work; wrong or weak evidence means retrieval or source structure; correct evidence with an unsafe answer means role or model behavior; equally correct answers with different speed or cost belong in Model Arena.
Keep the questions, expected answers, source version and scoring rules fixed. Change one layer and run the same known-answer and no-answer tests again. A provider change can invalidate embeddings and require re-indexing, so plan it separately instead of mixing it into an ordinary model comparison.
01–06
Set it up step by step
Create a fixed Knowledge Test before changing anything
A comparison is only useful when questions and expected facts remain identical.
Open Knowledge Optimizer → Knowledge Test, choose the intended assistant and select Custom Question. Add five natural paraphrases that should all return the unique PDF fact `NORDSTERN-42`. Include the direct question, a polite variation and wording a real visitor might use.
Write down the expected fact, approved source and pass condition before selecting Start test. A passing answer must contain the exact code, identify the Northstar Services demo PDF and avoid invented warranty, price or availability claims.
Keep the same source version, role, model and temperature. Add a separate unsupported 2028 warranty question to verify safe no-answer behavior. This prevents a fluent answer or a changed test setup from looking like an improvement when it is not.
Read the complete five-out-of-five baseline
Do not stop at the score; inspect every grounded answer.
The recorded run shows 5 answered, 0 partial, 0 unanswered and a 100% score. Open every result and confirm that each answer contains `NORDSTERN-42`, names the Northstar Services demo PDF and adds no invented warranty, price or availability claims.
Evaluate two layers separately. First ask whether the correct passage was retrieved. Then judge whether the answer is accurate, supported by that passage, complete enough for the question and free of unsupported additions. A single overall score can hide a fluent answer that used weak evidence.
This proves that the indexed PDF and retrieval path work for the known fact. It does not prove that the chatbot knows a 2028 warranty policy. Ask that unsupported question separately and expect a safe “no information” response rather than a guess.
Review Optimize and decide when to run an Audit
Conversation suggestions and broader audits answer different questions.
Open Optimize to review knowledge gaps derived from real visitor conversations. The verified tutorial account is empty because controlled Knowledge Tests do not fabricate customer traffic. In production, choose 30 or 90 days, run Analyze, open the original conversation and verify every suggestion against the authoritative owner and source before accepting it.
Use Audit when you need a wider review rather than one known failure. It scans for stale facts, contradictions, thin coverage and weak structure. Treat every finding as a prioritized checklist item: read the affected excerpt, confirm the latest approved fact and decide whether to edit the source or dismiss the finding. Never use Save to Inbox as automatic approval.
Empty Optimize results mean “no conversation-derived suggestion in this range”, not “the knowledge base is perfect”. The published real-click video runs Audit to completion and shows the genuine result. It also verifies that no finding was saved, no role was applied and no model was switched automatically.
Generate a role proposal—but do not apply it yet
Compare the current role, proposed role and rationale side by side.
Open Role and use the unsupported 2028 warranty variants to generate a proposal. The visible suggestion keeps answers short, requires indexed evidence, forbids extrapolation and tells the assistant to say when information is missing. Those are sensible behavioral guardrails.
Still, a role cannot create a warranty policy that does not exist in an approved source. Read Current role, Proposed role and Rationale side by side. Remove any instruction that conflicts with your support process, escalation path, required tone or legal wording.
Leave the proposal unapplied. Select the saved five-question test as the A/B source and compare current and proposed behavior on identical inputs. This isolates the role change; editing the source or model at the same time would make the result impossible to attribute.
Compare answer quality and latency in Model Arena
Use the same five questions, then inspect individual answers.
Reuse the last Knowledge Test so both candidates receive the same five PDF questions. The recorded comparison shows Vertex Gemini 2.5 Flash and Claude Haiku 4.5 at 100% quality; median latency is 2.6 seconds for Gemini and 4.4 seconds for Claude, so the current model is recommended for this test.
Open the per-question answers before deciding. Check exact facts, evidence use, safe refusal, completeness and any unusually slow response. Compare quota or model cost as well as median latency; one fast median should not hide repeated slow outliers or weaker safety behavior.
A green recommendation is not a universal winner. Expand the regression set with long answers, ambiguity, refusal, language and tool-use cases that matter to your chatbot. Changing AI providers can require new embeddings and a full re-index, so back up the current setup and schedule that comparison separately.
Reject the proposed role when the A/B test is worse
Generated changes earn trust only through measured results.
Run the exact five warranty questions against the current and proposed role. The recorded result is 0% better, 3 ties and 2 worse. Open the losing pairs to understand whether the new wording became less accurate, less useful or less safely grounded. Leave Apply role untouched and keep the current role.
If the company truly offers a 2028 warranty, add the approved policy—with owner, effective date, scope and exceptions—to an authoritative source, re-index and repeat the known-answer and safe-no-answer tests. If no policy exists, the correct improvement is a consistent safe fallback or human handoff, not more confident wording.
Before publishing any change, record the accepted source version, role, model, five-question result, unsupported-question result and latency. Review the Optimization Inbox on a schedule and keep suggestions pending until a human owner confirms both the fact and the measured regression result.
Example & result
See the practical test and its result
Every tutorial includes a fixed input, the expected outcome and a transparent record of what was actually verified locally.
Practical example: AI Chatbot Knowledge Optimizer: Find and fix weak answers
This exact scenario was completed with the temporary tutorial account.
Exact test input
What is the unique verification code in the uploaded tutorial PDF?
Expected result
All variants are answered and contain NORDSTERN-42.
What was actually verified
The real Knowledge Test scored 5/5 answered (100%); every answer contained NORDSTERN-42.
Tips & tricks
Make the setup reliable
Test with realistic examples, record your baseline and change one setting at a time. That makes real improvements visible.
Keep a regression set
An improvement for one question must not break other important questions. Re-run a small fixed set before publishing.
Never auto-accept factual rewrites
AI suggestions can simplify away exceptions or dates. Compare every factual change with the authoritative source.
Separate content from behavior
Add a missing policy to an owned source. Use the role only to control tone, evidence rules, fallbacks and escalation.
Version the evidence
Record source version, role, model and test date together so a later regression can be reproduced instead of guessed.
When something does not work
Troubleshooting
Check status, permissions and test data systematically before changing the model or prompt.
The approved fact exists, but the answer still misses it
Confirm that the source is indexed, open the retrieved passage and remove duplicate or outdated documents. Re-index after the approved source changes, then repeat the identical question.
The score is high, but an answer is unsafe
Inspect each answer and its evidence. Add explicit no-answer cases and verify a safe fallback or human handoff instead of relying on the overall percentage.
Optimize shows no suggestions
Check the 30- or 90-day range and whether real visitor conversations exist. Knowledge Tests deliberately do not create conversation-mining suggestions.
Model Arena results change between runs
Keep the question set and source version fixed, repeat the run, inspect individual responses and compare median latency together with slow outliers, quality and quota cost.
An Audit finding sounds plausible but is not verified
Do not save or apply it automatically. Open the affected excerpt, ask the source owner for the latest approved fact and either edit the authoritative source or dismiss the finding.
Ready for a production-style test
Add the verified question to a recurring regression set and review the Optimization Inbox on a fixed schedule instead of making ad-hoc changes.
Related resources
Find unanswered questions and knowledge gaps
Turn one real unsupported visitor question into a safe, trackable improvement workflow.
Choose the best AI model for your chatbot
Compare quality, latency, quota cost and tool behavior with a fixed regression set.
Fix wrong chatbot answers
Add source checks, safe fallbacks and regression tests to your knowledge workflow.
