A web crawler that uses Firefox and js injection to interact with webpages and crawl their content, written in nodejs.
-
Updated
Aug 24, 2023 - JavaScript
A web crawler that uses Firefox and js injection to interact with webpages and crawl their content, written in nodejs.
Cross-platform Python crawler that finds and verifies downloadable media, documents, and other files, then creates wget-ready URL lists for fast bulk downloading. It also scans sitemap trees and generates validated text or XML sitemaps, with HTTP, HTTPS, FTP, persistent SQLite history, resumable crawls, robots support, and no pip dependencies.
Web crawler designed to efficiently retrieve unique href, script and form links from a web application.
Generate a list of file links you can feed to wget for easy downloading! Mainly used for spidering web folders with lots of files. Can even generate a sitemap.txt or XML file for your website!
🕷️ Crawl websites efficiently with this Bash script, producing a clean list of URLs while respecting `robots.txt` and staying within specified domains.
To associate your repository with the web-spidering topic, visit your repo's landing page and select "manage topics."