Monitor discovery, downloads, and document intelligence
Crawler manager
Choose a public website, define hard limits, and specify exactly which documents should be kept.
Use a public HTTP or HTTPS website you have permission to crawl.
Optional. Projects group related jobs without changing crawl behavior.
Only selected formats will be downloaded.
Keywords, file size, pacing, timeout, and retry behavior.
Hard boundaries keep discovery predictable.
Link levels from start.
Maximum pages opened.
Maximum files saved.
Same domain only
Do not follow links to external websites
Enable JavaScript
Render pages that load links dynamically
Respect robots.txt
Honor the website publisher's crawl policy
Retry failed downloads
Retry transient network and server failures
Private-network and local targets are blocked. Use conservative limits and crawl only content you are authorized to access.
Configuration and runtime settings are saved before execution