Supercharge your AI agents with data from the web and beyond. Building the library for superintelligence. 🔥
-
Updated
Oct 8, 2026 - TypeScript
Supercharge your AI agents with data from the web and beyond. Building the library for superintelligence. 🔥
⬇️ A simple all-in-one CLI tool to download EVERYTHING from a URL (like youtube-dl/yt-dlp, forum-dl, gallery-dl, simpler ArchiveBox). 🎭 Uses headless Chrome to get HTML, JS, CSS, images/video/audio/subtitles, PDFs, screenshots, article text, git repos, and more...
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDceNQ and follow here for daily tips and tricks: https://x.com/Scrapling_dev
Continuously updated catalogue of open-source OSINT tools, skills, plugins, MCP servers, and agentic integrations.
Extract business leads and emails from Google Maps automatically with this Python tool featuring a simple web interface for efficient data collection.
Search the web from your coding agent with concise mini-briefings, ratings, and source links—no tab switching required
🎧 Download audio from YouTube and more with ease using this simple command-line tool that simplifies common audio extraction tasks.
The definitive list of the latest libraries, tools, APIs and providers for web scraping. The only daily-updated collection of web scraping resources.
High-performance web scraping engine that converts any web page into clean markdown --- with 3-layer fallback (Cheerio --> Playwright --> Abrasio) and AI-powered structured extraction
Extract data from websites using a fast, intelligent web scraping library designed for modern HTML structures and dynamic content.
Extract text from images using a robust OCR model designed for accuracy and efficiency in varied visual contexts.
The official Node.js SDK for Spidra.
Python scraper based on AI
Enhance your Scrap Mechanic experience with advanced cheats, unlimited resources, and instant building hacks for 2026.
A web-based email scraper for 2026, enabling fast and easy email data collection directly from your browser.
Scrape local events and activities directly in your browser with this simple HTML scraper. Find nearby happenings fast.
Python, Javascript, and Rust libraries for the Spider Cloud API.
AnyCrawl 🚀: A Node.js/TypeScript crawler that turns websites into LLM-ready data and extracts structured SERP results from Google/Bing/Baidu/etc. Native multi-threading for bulk processing.
A Powerful web scraper powered by LLM | OpenAI, Gemini & Ollama
To associate your repository with the ai-scraping topic, visit your repo's landing page and select "manage topics."