Crawl a website and its sitemap
Import useful public pages cleanly without indexing technical or private areas.
Updated 7 August 2026What crawling imports — and what it must ignore
The “Full website” source explores public pages on the same domain and extracts text the assistant can use. It expands the lightweight onboarding analysis: onboarding prepares the profile, while crawling builds a page-level knowledge base.
Allowed volume depends on the plan and monthly URL allowance. Wavize discovers addresses, queues them, processes them in the background and displays discovered, successful, excluded and failed counts. A large sitemap does not override your plan limit.
Security boundary
Use only public pages your organization is authorized to process. Wavize blocks private or unsafe destinations and must not bypass login, firewall or anti-bot challenges.
Create a full website source
Under Knowledge base, choose Full website, enter a recognizable title and the canonical URL.
Choose the public root
Prefer https://example.com or its final www variant. Avoid a cart URL, temporary campaign, private subdomain or redirect leaving the domain.
Keep “Train now” enabled
The checkbox is selected by default. It asks Wavize to start discovery and indexing after saving. Clear it only when preparing a source without running it.
Wait for processing
Do not repeatedly refresh or recreate the source. Status moves through queued and processing states; progress changes as background jobs complete.
Understand discovery, sitemap and selection
Wavize may use sitemaps and follow internal links, while keeping safety rules and priorities.
High-value pages
Home, services, categories, products, help, shipping, returns and contact are generally high priority. A page needs useful text, not only animation or blocks loaded after interaction.
Technical URLs
Exclude account, login, cart, checkout, internal search, tracking parameters, previews and infinite filters. They create duplicates or expose screens with no answer value.
Quota and partial selection
When discovery exceeds capacity, focus on the most useful pages or exclude URL families. The absolute safety cap never guarantees importing every URL from a very large site.
Manage URLs after the first crawl
The source detail lets you search discovered pages, persistently exclude them, restore them and retrain.
Exclude without losing control
A persistent exclusion prevents the page returning on the next crawl. Record the reason internally, especially for an old policy or local legal area.
Restore then retrain
Restoring the URL makes it eligible again; then retrain and wait for processing before testing information contained there.
Interpret failure carefully
The table stores the failed-page total but does not promise detailed diagnostics for every URL. Test the page publicly, then check redirects, HTTP status, firewall and content before contacting support.
Crawl quality check
A completed crawl is useful only when its scope matches what the assistant should know.
- Coverage — Find priority commercial and support pages in the list.
- Cleanliness — No cart, account or massive filter variant dominates results.
- Tested answer — A question returns an explicit fact from the processed page without invented additions.
Still need help?
Our team can help from your client area.