← Back to all projects

Scrapling

Scraping that survives redesigns: adaptive selectors, three grades of fetcher, and a Scrapy-style spider

BSD-3-Clause
Stars
76.9k
Forks
7.7k
Open issues
4
Last commit
25 Aug 2026

What is Scrapling?

You pick a fetcher per job — plain HTTP with TLS fingerprint impersonation when speed matters, a Playwright-driven browser for JavaScript pages, or a hardened stealth browser that patches headless tells and solves Cloudflare Turnstile — then select elements six ways, from CSS and XPath to text and structural similarity. Adaptive mode fingerprints matched elements so the same call finds them again after a site redesign, and the spider added in 0.4 crawls with deduplication, per-domain limits, pause-and-resume and JSON or CSV export. An MCP server lets a coding agent scrape conversationally, trimming pages by selector before they reach the model. It is a structured-extraction library for one machine: whole-site-to-markdown corpus building is Crawl4AI's territory, and the Turnstile-solving stealth sits in open tension with the README's own education-and-research disclaimer.

What can you do with Scrapling?

  • Three fetchers, one decision: speed or stealth — The plain fetcher does browser-like headers by default, TLS fingerprint impersonation and HTTP/3 at full speed; the dynamic one drives Chromium through Playwright and can attach to a real browser; the stealthy one adds headless-detection patches, canvas-fingerprint noise and automatic Turnstile solving — at the price of a browser's memory and pace. Sessions keep state and tab pools across requests.
  • Elements found again after the redesign — With adaptive and auto-save enabled, a matched element's tag, text, attributes, siblings and ancestor path are stored per domain; when the layout changes, the same call re-finds the element by similarity. Only the first match per selector is remembered — a stated limit worth knowing.
  • Six ways to say which element — CSS with text and attribute pseudo-elements, XPath, find-by-text, regular expressions, BeautifulSoup-style filters and find-similar for structurally alike elements — and it can generate a selector from any element it found, so a one-off discovery becomes a stable script.
  • A spider for when one page stops being enough — Scrapy-like async spiders bring start URLs, a priority queue with URL deduplication, per-domain concurrency limits, a robots.txt option, blocked-request detection with retries, checkpoint-based pause and resume, and JSON, JSONL, CSV or XML export — plus a proxy rotator and crawl templates.
  • Scraping as a conversation — The bundled MCP server exposes fetch and stealth-fetch tools with bulk variants, narrows pages by CSS selector before content reaches the model, and sanitises hidden DOM content against prompt injection; an interactive shell and terminal extract commands cover page-to-markdown without writing Python.

Before you choose Scrapling

  • Anti-bot bypass is the product — automatic Turnstile solving included — while the README limits use to education and research and defers to sites' terms; squaring those is your call to make before depending on it.
  • The API is still settling — 0.3 was a rewrite and 0.4 broke interfaces again — and the commit history is effectively one person's.

Star history

19 Aug to 28 Aug · +1.8k

75.2k76.9k

Frequently asked questions

Is Scrapling free for commercial use?

Scrapling is released under the BSD-3-Clause licence — OSI-approved open source, which permits commercial use.

How can Scrapling be deployed?

Scrapling is available as Runs locally.

Documentation

Reproduced from the D4Vinci/Scrapling README, published under BSD-3-Clause. Read the original ↗

Scrapling is an adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl.

Its parser learns from website changes and automatically relocates your elements when pages update. Its fetchers bypass anti-bot systems like Cloudflare Turnstile out of the box. And its spider framework lets you scale up to concurrent, multi-session crawls with pause/resume and automatic proxy rotation - all in a few lines of Python. One library, zero compromises.

Blazing fast crawls with real-time stats and streaming. Built by Web Scrapers for Web Scrapers and regular users, there’s something for everyone.

from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher
StealthyFetcher.adaptive = True
p = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True)  # Fetch website under the radar!
products = p.css('.product', auto_save=True)                                        # Scrape data that survives website design changes!
products = p.css('.product', adaptive=True)                                         # Later, if the website structure changes, pass `adaptive=True` to find them!

Or scale up to full crawls

from scrapling.spiders import Spider, Response

class MySpider(Spider):
  name = "demo"
  start_urls = ["https://example.com/"]

  async def parse(self, response: Response):
      for item in response.css('.product'):
          yield {"title": item.css('h2::text').get()}

MySpider().start()

Platinum Sponsors

Do you want to show your ad here? Click here

Key Features

Spiders - A Full Crawling Framework

  • 🕷️ Scrapy-like Spider API: Define spiders with start_urls, async parse callbacks, and Request/Response objects.
  • ⚡ Concurrent Crawling: Configurable concurrency limits, per-domain throttling, and download delays.
  • 🔄 Multi-Session Support: Unified interface for HTTP requests, and stealthy headless browsers in a single spider - route requests to different sessions by ID.
  • 💾 Pause & Resume: Checkpoint-based crawl persistence. Press Ctrl+C for a graceful shutdown; restart to resume from where you left off.
  • 📡 Streaming Mode: Stream scraped items as they arrive via async for item in spider.stream() with real-time stats - ideal for UI, pipelines, and long-running crawls.
  • 🛡️ Blocked Request Detection: Automatic detection and retry of blocked requests with customizable logic.
  • 🚦 AutoThrottle: Stop guessing delays. The spider tunes the delay of each domain on its own from how fast the website responds, then doubles it (or waits what Retry-After asks) whenever the website starts blocking or rate-limiting you, and speeds back up once it stops.
  • 🤖 Robots.txt Compliance: Optional robots_txt_obey flag that respects Disallow, Crawl-delay, and Request-rate directives with per-domain caching.
  • 🧪 Development Mode: Cache responses to disk on the first run and replay them on subsequent runs - iterate on your parse() logic without re-hitting the target servers.
  • 🧩 Ready-made Spider Templates: Skip the boilerplate with CrawlSpider for rule-based link following, SitemapSpider for sitemap/robots.txt-driven crawls, XMLFeedSpider/CSVFeedSpider for iterating XML/RSS and CSV feeds, and ShopifySpider to pull every product out of any Shopify store through its JSON API, one item per variant.
  • 🔗 Link Extraction: A standalone LinkExtractor primitive with allow/deny patterns, domain filters, CSS/XPath scoping, extension filtering, and canonicalization - use it inside the templates or on its own.
  • 📦 Built-in Export: Export results through hooks and your own pipeline or the built-in JSON/JSONL/CSV/XML exporters with result.items.to_json(), to_jsonl(), to_csv(), and to_xml().

Advanced Websites Fetching with Session Support

  • HTTP Requests: Fast and stealthy HTTP requests with the Fetcher class. Can impersonate browsers’ TLS fingerprint, headers, and use HTTP/3.
  • Dynamic Loading: Fetch dynamic websites with full browser automation through the DynamicFetcher class supporting Playwright’s Chromium and Google’s Chrome.
  • Anti-bot Bypass: Advanced stealth capabilities with StealthyFetcher and fingerprint spoofing. Can easily bypass all types of Cloudflare’s Turnstile/Interstitial with automation.
  • Session Management: Persistent session support with FetcherSession, StealthySession, and DynamicSession classes for cookie and state management across requests.
  • Proxy Rotation: Built-in ProxyRotator with cyclic or custom rotation strategies across all session types, plus per-request proxy overrides.
  • Domain & Ad Blocking: Block requests to specific domains (and their subdomains) or enable built-in ad blocking (~3,500 known ad/tracker domains) in browser-based fetchers.
  • DNS Leak Prevention: Optional DNS-over-HTTPS support to route DNS queries through Cloudflare’s DoH, preventing DNS leaks when using proxies.
  • Remote Browsers: Instead of launching a browser locally, connect to one that’s already running through CDP with cdp_url, whether it’s on the same machine, another host, or a managed browser provider. You can also point any browser fetcher at your own Chromium build with executable_path.
  • Background API Capture: Pass a URL pattern to capture_xhr, and all matching XHR/fetch responses the page makes while loading are collected for you as Response objects in response.captured_xhr - grab a site’s API data without reverse-engineering the requests yourself.
  • Async Support: Complete async support across all fetchers and dedicated async session classes.

Adaptive Scraping & AI Integration

  • 🔄 Smart Element Tracking: Relocate elements after website changes using intelligent similarity algorithms.
  • 🎯 Smart Flexible Selection: CSS selectors, XPath selectors, filter-based search, text search, regex search, and more.
  • 🔍 Find Similar Elements: Automatically locate elements similar to found elements.
  • 🤖 MCP Server to be used with AI: Built-in MCP server for AI-assisted Web Scraping and data extraction. The MCP server features powerful, custom capabilities that leverage Scrapling to extract targeted content before passing it to the AI (Claude/Cursor/etc), thereby speeding up operations and reducing costs by minimizing token usage. (demo video) It can also keep browser sessions open across calls, take page screenshots, and drive remote browsers over CDP.
  • 🧠 Agent Skill: A ready-to-install Agent Skill that teaches coding agents the whole library, so the code they write with Scrapling matches the current API instead of guessing.

High-Performance & battle-tested Architecture

  • 🚀 Lightning Fast: Optimized performance outperforming most Python scraping libraries.
  • 🔋 Memory Efficient: Optimized data structures and lazy loading for a minimal memory footprint.
  • ⚡ Fast JSON Serialization: 10x faster than the standard library.
  • 🏗️ Battle tested: Not only does Scrapling have 92% test coverage and full type hints coverage, but it has been used daily by hundreds of Web Scrapers over the past year.

Developer/Web Scraper Friendly Experience

  • 🎯 Interactive Web Scraping Shell: Optional built-in IPython shell with Scrapling integration, shortcuts, and new tools to speed up Web Scraping scripts development, like converting curl requests to Scrapling requests and viewing requests results in your browser.
  • 🚀 Use it directly from the Terminal: Optionally, you can use Scrapling to scrape a URL without writing a single line of code!
  • 🛠️ Rich Navigation API: Advanced DOM traversal with parent, sibling, and child navigation methods.
  • 🧬 Enhanced Text Processing: Built-in regex, cleaning methods, and optimized string operations.
  • 📝 Auto Selector Generation: Generate robust CSS/XPath selectors for any element.
  • 🔌 Familiar API: Similar to Scrapy/BeautifulSoup with the same pseudo-elements used in Scrapy/Parsel.
  • 🤝 Drop-in Scrapy Integration: Already invested in Scrapy? Decorate any callback with scrapling_response to parse the responses you already fetch with Scrapling’s parser, no rewrite needed.
  • 📘 Complete Type Coverage: Full type hints for excellent IDE support and code completion. The entire codebase is automatically scanned with PyRight and MyPy with each change.
  • 🔋 Ready Docker image: With each release, a Docker image containing all browsers is automatically built and pushed.

Getting Started

Let’s give you a quick glimpse of what Scrapling can do without deep diving.

Basic Usage

HTTP requests with session support

from scrapling.fetchers import Fetcher, FetcherSession

with FetcherSession(impersonate='chrome') as session:  # Use latest version of Chrome's TLS fingerprint
    page = session.get('https://quotes.toscrape.com/', stealthy_headers=True)
    quotes = page.css('.quote .text::text').getall()

# Or use one-off requests
page = Fetcher.get('https://quotes.toscrape.com/')
quotes = page.css('.quote .text::text').getall()

Advanced stealth mode

from scrapling.fetchers import StealthyFetcher, StealthySession

with StealthySession(headless=True, solve_cloudflare=True) as session:  # Keep the browser open until you finish
    page = session.fetch('https://nopecha.com/demo/cloudflare', google_search=False)
    data = page.css('#padded_content a').getall()

# Or use one-off request style, it opens the browser for this request, then closes it after finishing
page = StealthyFetcher.fetch('https://nopecha.com/demo/cloudflare')
data = page.css('#padded_content a').getall()

Full browser automation

from scrapling.fetchers import DynamicFetcher, DynamicSession

with DynamicSession(headless=True, disable_resources=False, network_idle=True) as session:  # Keep the browser open until you finish
    page = session.fetch('https://quotes.toscrape.com/', load_dom=False)
    data = page.xpath('//span[@class="text"]/text()').getall()  # XPath selector if you prefer it

# Or use one-off request style, it opens the browser for this request, then closes it after finishing
page = DynamicFetcher.fetch('https://quotes.toscrape.com/')
data = page.css('.quote .text::text').getall()

This README has been shortened. The full version is on GitHub. Read the original ↗