← プロジェクト一覧に戻る

Scrapling

サイト改修を生き延びるスクレイピング。要素を記憶して再発見する適応モードと、速度と偽装の3段階フェッチャーが柱です

BSD-3-Clause
スター
76.9k
フォーク
7.7k
オープンIssue
4
最終コミット
2026年8月25日

Scraplingとは

仕事ごとにフェッチャーを選びます。速度重視ならTLSフィンガープリントを偽装する素のHTTP、JavaScriptが要るページならPlaywright駆動のブラウザ、検出対策が要る相手ならヘッドレスの痕跡を消しCloudflare Turnstileも突破する強化ブラウザです。要素の選択はCSS・XPath・テキスト・構造的類似など6通りで、適応モードを有効にすると一致した要素の特徴を記録し、改修後のサイトでも同じ呼び出しで再発見します。0.4で加わったスパイダーは、重複排除・ドメイン別の同時実行制限・中断と再開・JSONやCSVへの出力を備えたクロールを担います。MCPサーバー経由ならコーディングエージェントが対話的にスクレイピングでき、ページはセレクタで絞ってからモデルに渡るのでトークンも節約されます。位置づけは1台で動く構造化抽出のライブラリです。サイト全体をMarkdown化するコーパス作りはCrawl4AIの領分で、Turnstile突破の機能とREADME自身の「教育・研究目的」という但し書きは緊張関係にあります。

Scraplingで何ができますか?

  • 3つのフェッチャーで、速度と偽装を選ぶ — 素のフェッチャーは既定でブラウザ風のヘッダーを送り、TLSフィンガープリントの偽装とHTTP/3を全速で使えます。動的フェッチャーはPlaywright経由でChromiumを駆動し、実ブラウザへの接続もできます。強化版はヘッドレス検出の痕跡消し、キャンバス指紋のノイズ付与、Turnstileの自動突破まで担い、その分ブラウザ相応のメモリと速度を要します。セッションは状態とタブのプールを維持します。
  • 改修後のサイトで要素をもう一度見つける — 適応モードと自動保存を有効にすると、一致した要素のタグ・テキスト・属性・兄弟・祖先の並びがドメイン単位で記録されます。レイアウトが変わっても、同じ呼び出しが類似度で同じ要素を再発見します。記憶されるのはセレクタごとに最初の一致だけ、という制限も明記されています。
  • 要素の指定は6通り — テキストや属性の擬似要素つきCSS、XPath、テキスト一致、正規表現、BeautifulSoup風のフィルタ、構造的に似た要素を探すfind_similarが使えます。見つけた要素からセレクタを自動生成できるので、一度の探索がそのまま安定したスクリプトになります。
  • 1ページで足りなくなったらスパイダー — Scrapy風の非同期スパイダーが、開始URL群、重複排除つきの優先度キュー、ドメイン別の同時実行制限、robots.txtの尊重オプション、ブロック検知と再試行、チェックポイントによる中断と再開、JSON・JSONL・CSV・XMLへの出力を備えます。プロキシのローテーターとクロールの雛形も付属します。
  • スクレイピングを会話でやる — 付属のMCPサーバーは取得・強化取得のツールを一括版も含めて公開し、内容をCSSセレクタで絞ってからモデルへ渡し、隠されたDOMの中身はプロンプトインジェクション対策として無害化します。対話シェルとターミナルの抽出コマンドを使えば、Pythonを書かずにページのMarkdown化もできます。

Scraplingを選ぶ前に

  • Turnstileの自動突破まで含む検出回避が製品の核である一方、READMEは用途を教育・研究に限ると述べ、各サイトの規約の尊重を求めています。この両者の整理は導入側に委ねられています。
  • APIはまだ固まっていません。0.3で全面的に書き直され、0.4でも互換性が壊れました。コミット履歴も実質1人のものです。

スター推移

8月19日〜8月28日 · +1.8k

75.2k76.9k

よくある質問

Scraplingは商用利用できますか?

ScraplingはBSD-3-Clauseライセンスで公開されています。OSI承認のオープンソースライセンスで、商用利用が認められています。

Scraplingはどの形で使えますか?

Scraplingはローカル実行の形で利用できます。

ドキュメント

D4Vinci/Scrapling のREADMEより転載(BSD-3-Clause)。 原文を読む ↗

Scrapling is an adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl.

Its parser learns from website changes and automatically relocates your elements when pages update. Its fetchers bypass anti-bot systems like Cloudflare Turnstile out of the box. And its spider framework lets you scale up to concurrent, multi-session crawls with pause/resume and automatic proxy rotation - all in a few lines of Python. One library, zero compromises.

Blazing fast crawls with real-time stats and streaming. Built by Web Scrapers for Web Scrapers and regular users, there’s something for everyone.

from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher
StealthyFetcher.adaptive = True
p = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True)  # Fetch website under the radar!
products = p.css('.product', auto_save=True)                                        # Scrape data that survives website design changes!
products = p.css('.product', adaptive=True)                                         # Later, if the website structure changes, pass `adaptive=True` to find them!

Or scale up to full crawls

from scrapling.spiders import Spider, Response

class MySpider(Spider):
  name = "demo"
  start_urls = ["https://example.com/"]

  async def parse(self, response: Response):
      for item in response.css('.product'):
          yield {"title": item.css('h2::text').get()}

MySpider().start()

Platinum Sponsors

Do you want to show your ad here? Click here

Key Features

Spiders - A Full Crawling Framework

  • 🕷️ Scrapy-like Spider API: Define spiders with start_urls, async parse callbacks, and Request/Response objects.
  • ⚡ Concurrent Crawling: Configurable concurrency limits, per-domain throttling, and download delays.
  • 🔄 Multi-Session Support: Unified interface for HTTP requests, and stealthy headless browsers in a single spider - route requests to different sessions by ID.
  • 💾 Pause & Resume: Checkpoint-based crawl persistence. Press Ctrl+C for a graceful shutdown; restart to resume from where you left off.
  • 📡 Streaming Mode: Stream scraped items as they arrive via async for item in spider.stream() with real-time stats - ideal for UI, pipelines, and long-running crawls.
  • 🛡️ Blocked Request Detection: Automatic detection and retry of blocked requests with customizable logic.
  • 🚦 AutoThrottle: Stop guessing delays. The spider tunes the delay of each domain on its own from how fast the website responds, then doubles it (or waits what Retry-After asks) whenever the website starts blocking or rate-limiting you, and speeds back up once it stops.
  • 🤖 Robots.txt Compliance: Optional robots_txt_obey flag that respects Disallow, Crawl-delay, and Request-rate directives with per-domain caching.
  • 🧪 Development Mode: Cache responses to disk on the first run and replay them on subsequent runs - iterate on your parse() logic without re-hitting the target servers.
  • 🧩 Ready-made Spider Templates: Skip the boilerplate with CrawlSpider for rule-based link following, SitemapSpider for sitemap/robots.txt-driven crawls, XMLFeedSpider/CSVFeedSpider for iterating XML/RSS and CSV feeds, and ShopifySpider to pull every product out of any Shopify store through its JSON API, one item per variant.
  • 🔗 Link Extraction: A standalone LinkExtractor primitive with allow/deny patterns, domain filters, CSS/XPath scoping, extension filtering, and canonicalization - use it inside the templates or on its own.
  • 📦 Built-in Export: Export results through hooks and your own pipeline or the built-in JSON/JSONL/CSV/XML exporters with result.items.to_json(), to_jsonl(), to_csv(), and to_xml().

Advanced Websites Fetching with Session Support

  • HTTP Requests: Fast and stealthy HTTP requests with the Fetcher class. Can impersonate browsers’ TLS fingerprint, headers, and use HTTP/3.
  • Dynamic Loading: Fetch dynamic websites with full browser automation through the DynamicFetcher class supporting Playwright’s Chromium and Google’s Chrome.
  • Anti-bot Bypass: Advanced stealth capabilities with StealthyFetcher and fingerprint spoofing. Can easily bypass all types of Cloudflare’s Turnstile/Interstitial with automation.
  • Session Management: Persistent session support with FetcherSession, StealthySession, and DynamicSession classes for cookie and state management across requests.
  • Proxy Rotation: Built-in ProxyRotator with cyclic or custom rotation strategies across all session types, plus per-request proxy overrides.
  • Domain & Ad Blocking: Block requests to specific domains (and their subdomains) or enable built-in ad blocking (~3,500 known ad/tracker domains) in browser-based fetchers.
  • DNS Leak Prevention: Optional DNS-over-HTTPS support to route DNS queries through Cloudflare’s DoH, preventing DNS leaks when using proxies.
  • Remote Browsers: Instead of launching a browser locally, connect to one that’s already running through CDP with cdp_url, whether it’s on the same machine, another host, or a managed browser provider. You can also point any browser fetcher at your own Chromium build with executable_path.
  • Background API Capture: Pass a URL pattern to capture_xhr, and all matching XHR/fetch responses the page makes while loading are collected for you as Response objects in response.captured_xhr - grab a site’s API data without reverse-engineering the requests yourself.
  • Async Support: Complete async support across all fetchers and dedicated async session classes.

Adaptive Scraping & AI Integration

  • 🔄 Smart Element Tracking: Relocate elements after website changes using intelligent similarity algorithms.
  • 🎯 Smart Flexible Selection: CSS selectors, XPath selectors, filter-based search, text search, regex search, and more.
  • 🔍 Find Similar Elements: Automatically locate elements similar to found elements.
  • 🤖 MCP Server to be used with AI: Built-in MCP server for AI-assisted Web Scraping and data extraction. The MCP server features powerful, custom capabilities that leverage Scrapling to extract targeted content before passing it to the AI (Claude/Cursor/etc), thereby speeding up operations and reducing costs by minimizing token usage. (demo video) It can also keep browser sessions open across calls, take page screenshots, and drive remote browsers over CDP.
  • 🧠 Agent Skill: A ready-to-install Agent Skill that teaches coding agents the whole library, so the code they write with Scrapling matches the current API instead of guessing.

High-Performance & battle-tested Architecture

  • 🚀 Lightning Fast: Optimized performance outperforming most Python scraping libraries.
  • 🔋 Memory Efficient: Optimized data structures and lazy loading for a minimal memory footprint.
  • ⚡ Fast JSON Serialization: 10x faster than the standard library.
  • 🏗️ Battle tested: Not only does Scrapling have 92% test coverage and full type hints coverage, but it has been used daily by hundreds of Web Scrapers over the past year.

Developer/Web Scraper Friendly Experience

  • 🎯 Interactive Web Scraping Shell: Optional built-in IPython shell with Scrapling integration, shortcuts, and new tools to speed up Web Scraping scripts development, like converting curl requests to Scrapling requests and viewing requests results in your browser.
  • 🚀 Use it directly from the Terminal: Optionally, you can use Scrapling to scrape a URL without writing a single line of code!
  • 🛠️ Rich Navigation API: Advanced DOM traversal with parent, sibling, and child navigation methods.
  • 🧬 Enhanced Text Processing: Built-in regex, cleaning methods, and optimized string operations.
  • 📝 Auto Selector Generation: Generate robust CSS/XPath selectors for any element.
  • 🔌 Familiar API: Similar to Scrapy/BeautifulSoup with the same pseudo-elements used in Scrapy/Parsel.
  • 🤝 Drop-in Scrapy Integration: Already invested in Scrapy? Decorate any callback with scrapling_response to parse the responses you already fetch with Scrapling’s parser, no rewrite needed.
  • 📘 Complete Type Coverage: Full type hints for excellent IDE support and code completion. The entire codebase is automatically scanned with PyRight and MyPy with each change.
  • 🔋 Ready Docker image: With each release, a Docker image containing all browsers is automatically built and pushed.

Getting Started

Let’s give you a quick glimpse of what Scrapling can do without deep diving.

Basic Usage

HTTP requests with session support

from scrapling.fetchers import Fetcher, FetcherSession

with FetcherSession(impersonate='chrome') as session:  # Use latest version of Chrome's TLS fingerprint
    page = session.get('https://quotes.toscrape.com/', stealthy_headers=True)
    quotes = page.css('.quote .text::text').getall()

# Or use one-off requests
page = Fetcher.get('https://quotes.toscrape.com/')
quotes = page.css('.quote .text::text').getall()

Advanced stealth mode

from scrapling.fetchers import StealthyFetcher, StealthySession

with StealthySession(headless=True, solve_cloudflare=True) as session:  # Keep the browser open until you finish
    page = session.fetch('https://nopecha.com/demo/cloudflare', google_search=False)
    data = page.css('#padded_content a').getall()

# Or use one-off request style, it opens the browser for this request, then closes it after finishing
page = StealthyFetcher.fetch('https://nopecha.com/demo/cloudflare')
data = page.css('#padded_content a').getall()

Full browser automation

from scrapling.fetchers import DynamicFetcher, DynamicSession

with DynamicSession(headless=True, disable_resources=False, network_idle=True) as session:  # Keep the browser open until you finish
    page = session.fetch('https://quotes.toscrape.com/', load_dom=False)
    data = page.xpath('//span[@class="text"]/text()').getall()  # XPath selector if you prefer it

# Or use one-off request style, it opens the browser for this request, then closes it after finishing
page = DynamicFetcher.fetch('https://quotes.toscrape.com/')
data = page.css('.quote .text::text').getall()

このREADMEは一部を省略しています。全文はGitHubにあります。 原文を読む ↗

Scrapling
AIに聞く
GitHub