Crawler5 min
Crawl limits explained
Understand depth, page, file, retry, and pacing limits before running a job.
Depth
Depth is the number of link levels followed from the starting page. Depth 0 inspects only the starting page. Depth 1 also opens pages linked directly from it. Larger depth values can expand the crawl dramatically.
Pages
Pages is the maximum number of HTML pages the crawler may inspect. When the limit is reached, remaining queued pages are not opened.
Files
Files is the maximum number of matching documents saved during the run. It is a hard cap for new downloads, not a performance setting.
Timeout, delay, and retries
- Timeout: maximum seconds allowed for one browser or download request.
- Delay: minimum pause between page visits; increase it to reduce pressure on the target.
- Retries: additional attempts after the first failed download. Zero still means one initial attempt.
Recommended starting profile
Depth: 1
Pages: 20
Files: 10
Delay: 0.5 seconds
Retries: 2Increase limits only after verifying the target structure and your authorization to crawl it.