Short answer
Diffbot is a crawler for Diffbot Knowledge Graph and web data products; Diffbot does not describe it as foundation-model training.
AI crawler guide
Diffbot is a crawler for Diffbot Knowledge Graph and web data products; Diffbot does not describe it as foundation-model training.
Short answer
Diffbot is a crawler for Diffbot Knowledge Graph and web data products; Diffbot does not describe it as foundation-model training.
Crawler facts
User-agent token
Diffbot
Operator
KI-Console group
Autonomous crawler: allow access, keep public content readable, and monitor logs.
robots.txt
Documentation
Open source documentationDiffbot crawls web content for search, Knowledge Graph, and extraction services. Diffbot separates the general Diffbot crawler from user-driven Diffbot-User fetches.
The KI-Console group is training because the bot crawls autonomously. Bot-specifically, the documented purpose is closer to index, Knowledge Graph, and web search.
Blocking Diffbot can reduce how systems from Diffbot see, index, train on, or retrieve your public content. Allowing it improves technical access, but it never guarantees crawling, ranking, citation, or a positive AI answer.
Diffbot documents robots.txt compliance including Disallow and Crawl-delay, while also noting exceptions for partnerships or agreements.
Use the exact token Diffbot in your robots.txt. Keep private paths blocked separately if only part of the site should be available.
Allow :bot
User-agent: Diffbot
Allow: /
Block :bot
User-agent: Diffbot
Disallow: /
Straight answers
KI-Console does not sell magic visibility. It makes your website more readable for AI systems and shows evidence where it can.
Diffbot is a crawler or robots.txt control token operated by Diffbot. The facts above show its purpose, robots.txt behavior, and exact allow or block rules.
Use User-agent: Diffbot and Disallow: /. robots.txt is a directive for compliant crawlers, not a technical access barrier.
It can reduce visibility in systems from Diffbot if those systems can no longer crawl, index, train on, or retrieve your content. The effect depends on the bot type and product.
Check your own website
Run the free scan first. With an account, you can verify the domain, keep history, generate files, and document crawler visits over time.