Start here6 min
Quick start
Start the backend and frontend, confirm the connection, and run a safe first crawl.
What this application does
UDC Pro discovers pages on a public website, identifies supported document links, downloads matching files to local storage, and records searchable metadata. The web dashboard controls jobs while the Python backend performs crawling and persistence.
1. Start the backend
cd C:\Users\musta\universal-document-crawler
.\venv\Scripts\Activate.ps1
python -m uvicorn backend.main:app --reload --host 127.0.0.1 --port 8010Keep this PowerShell window open. A healthy backend prints “Application startup complete.”
2. Start the frontend
cd C:\Users\musta\universal-document-crawler\frontend
npm install
npm run dev -- --port 3001Use a second PowerShell window. Open http://localhost:3001 after Next.js reports Ready.
3. Confirm connection
- Open Overview.
- Check the bottom-left workspace indicator. It must say the API is connected.
- If it is disconnected, verify the backend terminal and NEXT_PUBLIC_API_URL in frontend/.env.local.
4. Create a first crawl
- Open New crawl.
- Enter a descriptive job name and a public HTTPS URL.
- Keep conservative limits such as depth 1, pages 20, and files 10 for the first run.
- Select PDF and any other required types.
- Keep Respect robots.txt enabled, then start the crawl.
- Watch live progress in Jobs and downloaded results in Documents.