Skip to main content
Search

Importing files, links and sitemaps

Load PDF, Markdown, text and HTML files, web pages and whole sitemaps into the knowledge base and keep them current without duplicates.

Press Add knowledge on the Knowledge base page to load files (PDF, Markdown, plain text, HTML) and web pages, or a whole website through its sitemap. The import runs in the background, and importing the same file or link again updates the article it created instead of adding a copy.

Open the import window

  1. Go to Admin > AI > Knowledge base and press Add knowledge.
  2. Stay on the Files and links tab. Choose files under Files, paste links under Links, or do both.
  3. Press Import.

Each file and each web page becomes one article. The other tab of the window, Confluence, connects a wiki: see Confluence connector.

Files

  • Accepted types: PDF, Markdown (.md, .markdown), plain text (.txt) and HTML (.html, .htm). You can pick several files at once.
  • A file can be up to 25 MB. If one of the chosen files is bigger, the import does not start and asks you to split that file: cut it into parts and upload those.
  • Word and other formats are not accepted. Save a Word document as PDF, and export pages from Notion or Google Docs as Markdown or HTML first.
  • Text in a PDF is read from its text layer. A scanned PDF is a set of pictures and gives "No readable text found."; the same message appears for any file with less than 40 characters of text.
  • The article title is built from the file name and the first line of text: returns.pdf that starts with "How to return an item" becomes "returns - How to return an item". An HTML file takes the title of its page.
  • Put one link per line. Each page becomes its own article with the readable text of the page; the site's menus, headers and footers are left out.
  • Only web pages can be imported by link. A link to a PDF or another file is refused: download the file and add it under Files.
  • Import does not sign in to websites. A link with a login and password inside it is refused, and a page that opens only after sign-in cannot be read: the import fails, or you get the text of the sign-in page instead. A Confluence wiki is connected with the Confluence connector instead.
  • Links to addresses inside a private network, such as a company intranet, are refused.
  • A page that shows its text only after scripts run in the browser can come out empty, and the log then reports no readable text. Save such content as a file and upload it.
  • E-mail addresses that a site hides with Cloudflare e-mail protection are recovered, so the bot can name them.
  • Questions of collapsible FAQ blocks, where the question is a button that opens the answer, are kept as section headings, so the bot finds the answer by its question.
  • Email addresses and phone numbers on the page, the footer included, also go into one contacts article per website and language, named like Contacts of Example Shop: your company name as the site shows it in page titles, otherwise the domain. It updates itself when the pages are imported again. See Check your knowledge base with your AI.
  • When a customer asks for a link, the bot can give the address of the page its answer came from.

Sitemaps

A link that ends in .xml or has sitemap in it is read as a sitemap, for example https://example.com/help/sitemap.xml:

  • every page listed in it becomes its own article, while the sitemap itself does not;
  • a sitemap that lists other sitemaps is followed into them;
  • one sitemap link brings in up to 300 pages, and the rest are skipped. If the log says "300 pages found", the site may have more: add the sitemap of each section on its own line, or import the remaining pages as lists of links.

If such a link is not a sitemap, the log says "That link is not a sitemap.", and an empty one gives "The sitemap lists no pages." For the same reason, an ordinary web page whose address contains sitemap cannot be imported by link.

A help center you run on this site needs no import: let the bot use its articles directly, see Help center. A help center on another website is imported through its sitemap.

Watch the import

The import runs on the server. The window shows a counter (successful items out of the total, and the number that failed) and a log with one line for every file and page:

Log line Meaning
Added: name A new article was created.
Updated: name The article from this source got the new text.
Unchanged: name The source has not changed since the last import.
Failed: name This file or page was not imported; the reason follows the name.
link: N pages found A sitemap was opened and its pages were queued.

At the end the log says "Import completed" with the numbers of successful and failed items. A sitemap of a few hundred pages takes minutes. You can close the window and keep working: the import goes on, and the article list refreshes when it ends (or reload the page later). Only one import runs at a time; a second one is refused with "Import already in progress".

Importing again updates, not duplicates

Every imported article remembers its source: a file by its name, a web page by its address. When you import the same file name or the same address again, the existing article is updated, or left as it is when nothing changed, and no copy appears.

  • A file uploaded under a new name becomes a new article. Delete the old one by hand.
  • Two different files with the same name count as one source: the one imported later replaces the article of the other. Rename such files before you upload them.
  • Importing again keeps an article's Enabled state, and new articles start enabled.
  • In the list, articles from web pages carry the Imported label (hover to see the address), and articles from files carry From file.

Refresh from the source

  • For a web page, choose Refresh from source in the article's row menu: the page is read again right away ("Source re-read.").
  • For a file, upload the new version under the same file name.

Websites are not re-read on a schedule. After you change your site, run the import again, for example with the same sitemap link. Pages removed from the site stay in the knowledge base until you delete their articles; your own AI can compare the knowledge base with your sitemap and remove what is gone, see Connect your AI (MCP). Only a Confluence connection adds, updates and removes pages by itself.

Manual edits and conflicts

You can edit an imported article in the panel, but there the source always wins: the next import of the same file or link, or Refresh from source, replaces your edit with the text from the source. To make a lasting change, change the source itself (the page on your site or the document) and import it again.

Imports made by your own AI through Connect your AI protect manual edits:

  • if the article was edited in the panel and its source has not changed since the last import, your edit stays;
  • if both the article and its source have changed, nothing is written. The AI reports a conflict for that article and asks you which version to keep.

The hint on the From file and Synced labels is a reminder of this: a manual edit may conflict with the next sync. Articles from Confluence are edited only in Confluence.

Limits

What Limit
File types PDF, Markdown, plain text, HTML
File size 25 MB per file
Pages from one sitemap link 300
Text in a file or page at least 40 characters
One web page must answer within 20 seconds; only its first 3 MB are read
Imports at the same time one

How much the whole knowledge base can hold depends on your plan: Plans, limits and billing.

Other sources

For Zendesk, Notion or a folder of documents, your own AI can do the work: it reads the source, shows a plan of what it will add, update and delete, and only then writes, and any import can be rolled back. See Connect your AI (MCP).

Was this article helpful?