Setting up a web crawler to index external content

Plan requirement

Subscription Suite Growth or higher, Guide Professional or higher
Access Admin

The crawler indexes any site you point it at, which makes it the general answer for content with no direct connector.

Set one up

  1. Open the search settings in Knowledge.
  2. Add a web crawler and give it the starting address.
  3. Add URL filters to scope what it takes.
  4. Set the language of the content.
  5. Run it, then check what it found.

Filters are the whole job

Without them a crawler takes everything it can reach: your marketing pages, your careers page, old campaign microsites. All of that then appears in help centre search.

Restrict it to the section that holds documentation, and check the result rather than assuming.

Set the language

A crawler indexing English content marked as Dutch produces results that appear for the wrong readers. Where a site has several languages, that usually means a crawler per language with its own filters.

It only sees public pages

Anything behind a login is invisible to it. That is a limitation and a safeguard: it cannot accidentally index something private.

Check it after site changes

A crawler is tied to the structure of the site it reads. A redesign that changes addresses leaves it indexing nothing or indexing the wrong pages, and nothing warns you.

See also

Was this article helpful?

0

Still stuck?

Our support team will take a look with you.

Comments

0 comments

Article is closed for comments.