Crawler7 min
Filters and crawl safety
Control what is followed and downloaded while respecting website policies.
Domain and JavaScript controls
- Same domain only prevents discovery from expanding to external websites.
- Enable JavaScript renders dynamic pages that do not expose links in their initial HTML.
- Respect robots.txt checks the publisher's machine-readable crawling policy before visiting pages.
Document filters
- File types restrict downloads by supported extension.
- Include keywords require at least one comma-separated term in the document URL.
- Exclude keywords skip any URL containing a blocked term.
- Filename contains adds a strict filename substring check.
- Minimum and maximum size are evaluated when the server provides a valid size or after download validation.
Built-in protection
The backend rejects localhost, private-network, link-local, and other non-public targets. Redirects are validated again. This reduces server-side request forgery risk but does not replace organizational authorization, legal review, or responsible rate limits.