WebChatAgent data sources: every control explained
Understand the complete Data Sources workspace and maintain a reliable chatbot knowledge base without guesswork.

WebChatAgent data sources control which websites, documents and connected workspaces can support an assistant answer.
Data Sources is where you decide which facts an assistant may retrieve. Websites, files, direct text, Confluence and Notion share one workspace, but they differ in ownership, refresh behavior and the controls available after indexing.
This guide uses a real Premium test assistant named Source Control Center. Its website and PDF are indexed, and a real conversation returns the unique PDF fact NORDSTERN-42. That final check is important: a green status proves indexing finished, while the answer test proves the information can actually be retrieved.
Privacy-protected two-click player
WebChatAgent Data Sources Tutorial: Every Control Explained
Understand WebChatAgent data sources, indexing states, search, filters, metadata, re-indexing and safe deletion with real interface examples.
The YouTube player stays blocked until you choose Play. Loading it connects your browser to YouTube and may transfer technical data to Google.
Open directly on YouTubeWhat you will have at the end
- The right source type for each content owner and update cycle
- A repeatable review of indexing status, character usage and filters
- Safe edit, content-search, re-index and delete procedures
- A real known-answer regression test
Before you start
- A chatbot for which you can manage sources
- At least one small approved test source
- One exact question whose answer appears only in that source
Source lifecycle from input to retrieval
Adding a source starts extraction and indexing. The resulting passages count toward the character-based content quota and become available only after a successful status.
Editing scope or changing provider can require a new index. Deleting a source removes its knowledge from future retrieval.
01–10
Set it up step by step
Open the correct assistant and Data Sources tab
Confirm the assistant name before changing its knowledge.
Select Source Control Center in the left sidebar and open Data Sources in the assistant navigation. The page heading, sidebar selection and source list must belong to the same assistant.
Read the information banner before adding content: sources are split into searchable passages and count toward the plan’s content-character allowance. Changes here affect future answers, so never work from an ambiguous browser tab.
Compare every Add Data Source option
Choose by ownership and update method, not by file name alone.
Open Add Data Source. Website crawls approved public pages; Documents uploads PDF, TXT, Markdown, Word or Excel; Text Input stores a short maintained fact directly. Confluence and Notion connect approved workspace content.
For example, use Website for a public help center, Documents for a signed policy PDF, Text Input for one temporary service notice, and a connector when the content owner already maintains the source in that workspace. Close the dialog without creating duplicates.
Read status, character usage and source counts
A completed source is usable; an unfinished one is not ready for answer testing.
The usage card compares indexed characters with the plan limit. Source counters separate websites, documents, text and connectors, while the table shows each item’s indexing status.
Treat Pending or Processing as a wait state. Failed or Error needs details and a correction. Re-index required means the current vectors are unavailable or outdated, often after a provider change. Completed or Indexed means the job finished, but you still need the answer test in step ten.
Find a source with search, filters and categories
Stable categories make a large source inventory understandable.
Search accepts a source name or website URL. Use the type filter to isolate websites, files, text or connectors, and the category filter to narrow ownership or topic. Reset both filters before concluding that a source is missing.
Source Control Center uses the category Support for the Northstar PDF. In production prefer durable labels such as Support, Products or Legal. Do not encode temporary status such as new or needs review in a category; status and ownership should remain separate.
Inspect website scope before re-indexing
URL, crawl depth, exclusions and schedule decide what enters the index.
Open the website row’s edit action. Verify the canonical start URL, crawl depth and automatic re-index interval. Expand Advanced only when selectors or exclusion patterns are necessary.
A depth of zero indexes only the entered page; a greater depth follows links. Exclude checkout, account and duplicate-language paths before increasing depth. Save only intentional changes, then wait for the new indexing job instead of testing immediately.
Review the pages inside a website source
A source can be green while one unwanted page still consumes quota.
Open Indexed URLs for the website. Search the page list, compare character counts and use View Content to spot navigation noise, cookie text or duplicate pages.
Delete removes only the selected indexed page until the next crawl may discover it again. Exclude adds a durable exclusion for future crawls. Use the confirmation dialog, then verify the total-character count changed as expected.
Edit extracted text, metadata and Search Priority
Review exactly what the chatbot can retrieve from this document.
Open Edit on Northstar Knowledge Base. The editor shows the extracted text used for indexing, plus Document Name, Category and Search Priority. Correct a small extraction error only when you can compare it with the approved original.
Search Priority influences ranking relative to other sources; it is not a quality score. Keep zero unless repeatable questions justify a boost. For a changed policy or larger revision, upload the approved new file, verify it, then remove the obsolete source instead of maintaining a hidden fork in the extracted text.
Search inside the indexed content
Content Search checks the stored passages, not just source titles.
Open Content Search and enter the unique value NORDSTERN-42. The result should point to Northstar Knowledge Base and show the matching indexed passage.
This is the quickest way to separate a source problem from an answer-generation problem. No result means the fact is not in the current index. A matching result with a poor chatbot answer means you should inspect the question wording, source overlap, Search Priority and model behavior.
Use re-index and delete without losing required knowledge
Re-index refreshes a source; delete removes it from future retrieval.
Use the row action to re-index one source, or Re-index All only after a provider change or a deliberately coordinated refresh. Large bulk jobs consume time and can make troubleshooting harder, so record what changed first.
Before Delete, identify every unique fact that lives only in that source. The confirmation warns that removal affects future answers. Cancel during a review, or delete only after a replacement is completed and its fixed regression questions pass.
Verify the source with an exact chatbot question
The observed answer must match the fact in the indexed document.
Open Conversations and select the Source Control Center test session. Exact input: “What is the unique verification code in the uploaded tutorial PDF?” Expected result: the reply names NORDSTERN-42 and does not invent a different code.
Observed result: the real assistant replied, “The unique verification code is NORDSTERN-42.” This example is verified end to end. Keep one known-answer question per critical source and repeat it after file replacement, re-indexing, model changes or source deletion.
Example & result
See the practical test and its result
Every tutorial includes a fixed input, the expected outcome and a transparent record of what was actually verified locally.
Practical example: WebChatAgent Data Sources explained: every control with examples
This exact scenario was completed with the temporary tutorial account.
Exact test input
Ask: “What is the unique verification code in the uploaded tutorial PDF?”
Expected result
The assistant returns NORDSTERN-42 from the indexed PDF instead of guessing.
What was actually verified
After the website and PDF both showed Completed, the real visitor test answered with NORDSTERN-42 and referenced the indexed tutorial knowledge.
Tips & tricks
Make the setup reliable
Test with realistic examples, record your baseline and change one setting at a time. That makes real improvements visible.
Remove duplicates before buying more quota
Navigation copies, print pages and translated duplicates can consume characters and create competing passages. Clean the source scope before upgrading the plan.
Assign an owner to every source
Ownership makes outdated policies and abandoned connectors easier to detect and remove.
Change one retrieval variable at a time
Keep model, source set and test questions stable while comparing Search Priority, category or crawl scope. Otherwise you cannot explain why a result changed.
Use a fast model for routine source checks
A fast economical model is usually enough for deterministic known-answer tests. Compare a stronger model only when the indexed passage is present and the fixed questions still fail.
When something does not work
Troubleshooting
Check status, permissions and test data systematically before changing the model or prompt.
The source is completed, but the answer misses its fact
Find the exact phrase with Content Search. If it is absent, fix extraction or scope and re-index. If it is present, test the exact question, remove competing duplicates and compare Search Priority or model only one change at a time.
A website keeps adding unwanted pages
Use Exclude rather than only Delete, or add a path pattern in the website source settings. Then re-index and verify the indexed URL list and character total.
Re-index required appears after a provider change
Start re-indexing for the affected sources, wait until every job is completed and rerun the fixed known and unknown questions before publishing the change.
Ready for a production-style test
Create a quarterly source inventory with owner, category, refresh method, last verification date and one fixed question for every critical source.
