Step-by-step tutorial Knowledge base

WebChatAgent data sources: every control explained

Understand the complete Data Sources workspace and maintain a reliable chatbot knowledge base without guesswork.

Beginner27 min readJuly 16, 2026
WebChatAgent data sources: every control explained

WebChatAgent data sources control which websites, documents and connected workspaces can support an assistant answer.

Data Sources is where you decide which facts an assistant may retrieve. Websites, files, direct text, Confluence and Notion share one workspace, but they differ in ownership, refresh behavior and the controls available after indexing.

This guide uses a real Premium test assistant named Source Control Center. Its website and PDF are indexed, and a real conversation returns the unique PDF fact NORDSTERN-42. That final check is important: a green status proves indexing finished, while the answer test proves the information can actually be retrieved.

Privacy-protected two-click player

WebChatAgent Data Sources Tutorial: Every Control Explained

Understand WebChatAgent data sources, indexing states, search, filters, metadata, re-indexing and safe deletion with real interface examples.

YouTube · 5:35 · English

The YouTube player stays blocked until you choose Play. Loading it connects your browser to YouTube and may transfer technical data to Google.

Open directly on YouTube

What you will have at the end

  • The right source type for each content owner and update cycle
  • A repeatable review of indexing status, character usage and filters
  • Safe edit, content-search, re-index and delete procedures
  • A real known-answer regression test

Before you start

  • A chatbot for which you can manage sources
  • At least one small approved test source
  • One exact question whose answer appears only in that source

Source lifecycle from input to retrieval

Adding a source starts extraction and indexing. The resulting passages count toward the character-based content quota and become available only after a successful status.

Editing scope or changing provider can require a new index. Deleting a source removes its knowledge from future retrieval.

Add sourceIndex and monitorRefresh or remove

01–10

Set it up step by step

1

Open the correct assistant and Data Sources tab

Confirm the assistant name before changing its knowledge.

Select Source Control Center in the left sidebar and open Data Sources in the assistant navigation. The page heading, sidebar selection and source list must belong to the same assistant.

Read the information banner before adding content: sources are split into searchable passages and count toward the plan’s content-character allowance. Changes here affect future answers, so never work from an ambiguous browser tab.

Confirm the assistant name before changing its knowledge.
2

Compare every Add Data Source option

Choose by ownership and update method, not by file name alone.

Open Add Data Source. Website crawls approved public pages; Documents uploads PDF, TXT, Markdown, Word or Excel; Text Input stores a short maintained fact directly. Confluence and Notion connect approved workspace content.

For example, use Website for a public help center, Documents for a signed policy PDF, Text Input for one temporary service notice, and a connector when the content owner already maintains the source in that workspace. Close the dialog without creating duplicates.

Choose by ownership and update method, not by file name alone.
3

Read status, character usage and source counts

A completed source is usable; an unfinished one is not ready for answer testing.

The usage card compares indexed characters with the plan limit. Source counters separate websites, documents, text and connectors, while the table shows each item’s indexing status.

Treat Pending or Processing as a wait state. Failed or Error needs details and a correction. Re-index required means the current vectors are unavailable or outdated, often after a provider change. Completed or Indexed means the job finished, but you still need the answer test in step ten.

A completed source is usable; an unfinished one is not ready for answer testing.
4

Find a source with search, filters and categories

Stable categories make a large source inventory understandable.

Search accepts a source name or website URL. Use the type filter to isolate websites, files, text or connectors, and the category filter to narrow ownership or topic. Reset both filters before concluding that a source is missing.

Source Control Center uses the category Support for the Northstar PDF. In production prefer durable labels such as Support, Products or Legal. Do not encode temporary status such as new or needs review in a category; status and ownership should remain separate.

Stable categories make a large source inventory understandable.
5

Inspect website scope before re-indexing

URL, crawl depth, exclusions and schedule decide what enters the index.

Open the website row’s edit action. Verify the canonical start URL, crawl depth and automatic re-index interval. Expand Advanced only when selectors or exclusion patterns are necessary.

A depth of zero indexes only the entered page; a greater depth follows links. Exclude checkout, account and duplicate-language paths before increasing depth. Save only intentional changes, then wait for the new indexing job instead of testing immediately.

URL, crawl depth, exclusions and schedule decide what enters the index.
6

Review the pages inside a website source

A source can be green while one unwanted page still consumes quota.

Open Indexed URLs for the website. Search the page list, compare character counts and use View Content to spot navigation noise, cookie text or duplicate pages.

Delete removes only the selected indexed page until the next crawl may discover it again. Exclude adds a durable exclusion for future crawls. Use the confirmation dialog, then verify the total-character count changed as expected.

A source can be green while one unwanted page still consumes quota.
7

Edit extracted text, metadata and Search Priority

Review exactly what the chatbot can retrieve from this document.

Open Edit on Northstar Knowledge Base. The editor shows the extracted text used for indexing, plus Document Name, Category and Search Priority. Correct a small extraction error only when you can compare it with the approved original.

Search Priority influences ranking relative to other sources; it is not a quality score. Keep zero unless repeatable questions justify a boost. For a changed policy or larger revision, upload the approved new file, verify it, then remove the obsolete source instead of maintaining a hidden fork in the extracted text.

Review exactly what the chatbot can retrieve from this document.
9

Use re-index and delete without losing required knowledge

Re-index refreshes a source; delete removes it from future retrieval.

Use the row action to re-index one source, or Re-index All only after a provider change or a deliberately coordinated refresh. Large bulk jobs consume time and can make troubleshooting harder, so record what changed first.

Before Delete, identify every unique fact that lives only in that source. The confirmation warns that removal affects future answers. Cancel during a review, or delete only after a replacement is completed and its fixed regression questions pass.

Re-index refreshes a source; delete removes it from future retrieval.
10

Verify the source with an exact chatbot question

The observed answer must match the fact in the indexed document.

Open Conversations and select the Source Control Center test session. Exact input: “What is the unique verification code in the uploaded tutorial PDF?” Expected result: the reply names NORDSTERN-42 and does not invent a different code.

Observed result: the real assistant replied, “The unique verification code is NORDSTERN-42.” This example is verified end to end. Keep one known-answer question per critical source and repeat it after file replacement, re-indexing, model changes or source deletion.

The observed answer must match the fact in the indexed document.

Example & result

See the practical test and its result

Every tutorial includes a fixed input, the expected outcome and a transparent record of what was actually verified locally.

Practical example: WebChatAgent Data Sources explained: every control with examples

This exact scenario was completed with the temporary tutorial account.

Verified end to end

Exact test input

Ask: “What is the unique verification code in the uploaded tutorial PDF?”

Expected result

The assistant returns NORDSTERN-42 from the indexed PDF instead of guessing.

What was actually verified

After the website and PDF both showed Completed, the real visitor test answered with NORDSTERN-42 and referenced the indexed tutorial knowledge.

After the website and PDF both showed Completed, the real visitor test answered with NORDSTERN-42 and referenced the indexed tutorial knowledge.

Tips & tricks

Make the setup reliable

Test with realistic examples, record your baseline and change one setting at a time. That makes real improvements visible.

Remove duplicates before buying more quota

Navigation copies, print pages and translated duplicates can consume characters and create competing passages. Clean the source scope before upgrading the plan.

Assign an owner to every source

Ownership makes outdated policies and abandoned connectors easier to detect and remove.

Change one retrieval variable at a time

Keep model, source set and test questions stable while comparing Search Priority, category or crawl scope. Otherwise you cannot explain why a result changed.

Use a fast model for routine source checks

A fast economical model is usually enough for deterministic known-answer tests. Compare a stronger model only when the indexed passage is present and the fixed questions still fail.

When something does not work

Troubleshooting

Check status, permissions and test data systematically before changing the model or prompt.

The source is completed, but the answer misses its fact

Find the exact phrase with Content Search. If it is absent, fix extraction or scope and re-index. If it is present, test the exact question, remove competing duplicates and compare Search Priority or model only one change at a time.

A website keeps adding unwanted pages

Use Exclude rather than only Delete, or add a path pattern in the website source settings. Then re-index and verify the indexed URL list and character total.

Re-index required appears after a provider change

Start re-indexing for the affected sources, wait until every job is completed and rerun the fixed known and unknown questions before publishing the change.

Ready for a production-style test

Create a quarterly source inventory with owner, category, refresh method, last verification date and one fixed question for every critical source.

Related resources