Scrapling 是什麼?

Scrapling 是一個自適應的 Python 網頁爬蟲框架,由開發者 D4Vinci 建立。它在 2026 年 GitHub Trending 上迅速走紅,獲得超過 57,000 個 Star,成為新一代爬蟲工具的代表。

為什麼需要 Scrapling?

傳統爬蟲框架(如 Scrapy、BeautifulSoup)面臨三大痛點:

  1. 網站結構變化導致程式碼失效:CSS 選擇器或 XPath 一旦頁面改版就全部崩潰
  2. 反爬系統攔截:Cloudflare Turnstile、Akamai 等防護讓普通請求直接返回 403
  3. 擴充性差:從小規模抓取到大規模並行爬取,需要重寫大量程式碼

Scrapling 的核心設計理念是 "One library, zero compromises"(一個庫,零妥協):

  • 自適應解析器:自動學習網站結構,頁面更新後自動重新定位元素
  • 內建反反爬:開箱即用,繞過 Cloudflare Turnstile 等主流防護
  • Scrapy-like API:熟悉 Scrapy 的開發者可以無縫遷移
  • 從單請求到全量爬取:同一套 API 支援簡單抓取和大規模並行爬蟲

Scrapling vs 傳統爬蟲工具

特性 BeautifulSoup Scrapy Scrapling
學習曲線 低 中 中
自適應解析 ❌ ❌ ✅
繞過 Cloudflare ❌ 需外掛 ✅ 內建
並行爬取 ❌ ✅ ✅
動態頁面支援 ❌ 需中介軟體 ✅ 內建 Playwright
暫停/恢復 ❌ 需擴充套件 ✅ 內建
串流輸出 ❌ ❌ ✅

安裝 Scrapling

環境要求

  • Python 3.8+
  • pip 套件管理器

快速安裝

BASH
pip install scrapling

驗證安裝

PYTHON
from scrapling.fetchers import Fetcher

# 測試基本功能
p = Fetcher.fetch('https://example.com')
print(p.title)  # 輸出頁面標題

可選依賴

如果需要動態頁面渲染(JavaScript 網站),需要安裝 Playwright:

BASH
pip install playwright
playwright install chromium

快速上手:第一個爬蟲

範例 1:簡單頁面抓取

讓我們從一個簡單的例子開始——抓取 Hacker News 的頭條新聞。

PYTHON
from scrapling.fetchers import Fetcher

# 發起 HTTP 請求
page = Fetcher.fetch('https://news.ycombinator.com/')

# 使用 CSS 選擇器提取資料
stories = page.css('.titleline > a')

for story in stories[:5]:  # 只取前 5 條
    title = story.text
    link = story.attrs.get('href', '')
    print(f"標題: {title}")
    print(f"連結: {link}")
    print("-" * 40)

輸出範例:

標題: Show HN: I built a real-time code collaboration tool
連結: https://github.com/example/collab-tool
----------------------------------------
標題: Ask HN: What's your favorite Python library in 2026?
連結: https://news.ycombinator.com/item?id=123456
----------------------------------------

範例 2:自適應解析(Auto-save)

Scrapling 的核心特性是自適應解析。當你第一次抓取資料時,可以啟用 auto_save=True,Scrapling 會學習頁面結構並儲存特徵。當網站改版後,只需傳入 adaptive=True,它就能自動找到目標元素。

PYTHON
from scrapling.fetchers import Fetcher

page = Fetcher.fetch('https://quotes.toscrape.com/')

# 第一次抓取:啟用 auto_save
quotes = page.css('.quote', auto_save=True)

for quote in quotes[:3]:
    text = quote.css('.text::text').get()
    author = quote.css('.author::text').get()
    print(f"{text} — {author}")

如果網站結構變化了:

PYTHON
# 後續抓取:傳入 adaptive=True
page = Fetcher.fetch('https://quotes.toscrape.com/')
quotes = page.css('.quote', adaptive=True)  # 自動適應新結構!

for quote in quotes[:3]:
    text = quote.css('.text::text').get()
    author = quote.css('.author::text').get()
    print(f"{text} — {author}")

💡 工作原理:Scrapling 會記錄元素的多種特徵(標籤類型、附近文字、屬性模式等),即使 CSS 類名改變,它也能透過其他特徵重新定位元素。


進階功能

1. 繞過 Cloudflare Turnstile

很多網站使用 Cloudflare Turnstile 或其他反爬系統。Scrapling 的 StealthyFetcher 可以自動處理這些防護。

PYTHON
from scrapling.fetchers import StealthyFetcher

# 啟用自適應模式
StealthyFetcher.adaptive = True

# 自動繞過 Cloudflare
page = StealthyFetcher.fetch(
    'https://example-protected-site.com',
    headless=True,       # 無頭瀏覽器模式
    network_idle=True    # 等待網路空閒
)

# 正常提取資料
products = page.css('.product-item')
for product in products:
    name = product.css('.name::text').get()
    price = product.css('.price::text').get()
    print(f"{name}: {price}")

關鍵參數說明: - headless=True:使用無頭瀏覽器,模擬真實使用者行為 - network_idle=True:等待所有網路完成後再提取(適合 SPA 應用) - adaptive=True:啟用自適應解析

2. 動態頁面渲染(Playwright)

對於需要 JavaScript 渲染的網站,使用 DynamicFetcher。

PYTHON
from scrapling.fetchers import DynamicFetcher

page = DynamicFetcher.fetch(
    'https://spa-example.com',
    wait_for='.content-loaded',  # 等待特定元素出現
    timeout=30000                # 逾時時間(毫秒)
)

# 提取動態載入的內容
articles = page.css('article')
for article in articles:
    title = article.css('h2::text').get()
    summary = article.css('.summary::text').get()
    print(f"{title}\n{summary}\n")

3. 非同步並行抓取

使用 AsyncFetcher 實現高並行抓取。

PYTHON
import asyncio
from scrapling.fetchers import AsyncFetcher

async def fetch_multiple_pages():
    urls = [
        'https://example.com/page/1',
        'https://example.com/page/2',
        'https://example.com/page/3',
        'https://example.com/page/4',
        'https://example.com/page/5',
    ]

    # 並行發起請求
    pages = await AsyncFetcher.fetch_many(urls, concurrency=3)

    for url, page in zip(urls, pages):
        if page:
            title = page.title
            print(f"{url}: {title}")
        else:
            print(f"{url}: 請求失敗")

asyncio.run(fetch_multiple_pages())

Spider 框架:大規模爬取

Scrapling 提供了類似 Scrapy 的 Spider 框架,支援大規模並行爬取。

基礎 Spider

PYTHON
from scrapling.spiders import Spider, Response

class QuoteSpider(Spider):
    name = "quotes"
    start_urls = ["https://quotes.toscrape.com/"]

    async def parse(self, response: Response):
        # 提取目前頁面的名言
        for quote in response.css('.quote'):
            yield {
                "text": quote.css('.text::text').get(),
                "author": quote.css('.author::text').get(),
                "tags": quote.css('.tag::text').getall()
            }

        # 翻頁
        next_page = response.css('.next a::attr(href)').get()
        if next_page:
            yield self.follow(next_page, callback=self.parse)

# 啟動爬蟲
QuoteSpider().start()

並行設定

PYTHON
class MultiPageSpider(Spider):
    name = "multi-page"
    start_urls = [f"https://example.com/page/{i}" for i in range(1, 101)]

    # 設定並行數
    custom_settings = {
        "concurrency": 5,           # 最大並行數
        "download_delay": 1,        # 下載延遲(秒)
        "robots_txt_obey": True,    # 遵守 robots.txt
    }

    async def parse(self, response: Response):
        title = response.css('h1::text').get()
        yield {"url": response.url, "title": title}

MultiPageSpider().start()

暫停與恢復

Scrapling 支援 checkpoint-based 持久化,按 Ctrl+C 優雅退出後,下次啟動會自動恢復。

PYTHON
class LongRunningSpider(Spider):
    name = "long-crawl"
    start_urls = ["https://large-site.com/"]

    custom_settings = {
        "checkpoint_dir": "./checkpoints",  # 檢查點目錄
    }

    async def parse(self, response: Response):
        # 提取資料...
        yield {"data": "..."}

        # 繼續爬取
        for link in response.css('a::attr(href)').getall():
            yield self.follow(link, callback=self.parse)

LongRunningSpider().start()

實戰案例

PYTHON
from scrapling.fetchers import Fetcher

def scrape_github_trending():
    page = Fetcher.fetch('https://github.com/trending')

    repos = page.css('.Box-row')

    trending = []
    for repo in repos[:10]:
        name = repo.css('h2 a::text').get('').strip()
        description = repo.css('p.col-9::text').get('').strip()
        stars = repo.css('[href$=stargazers] span::text').get('').strip()
        language = repo.css('[itemprop=programmingLanguage]::text').get('').strip()

        trending.append({
            "name": name,
            "description": description,
            "stars": stars,
            "language": language
        })

    return trending

if __name__ == "__main__":
    results = scrape_github_trending()
    for repo in results:
        print(f"📦 {repo['name']}")
        print(f"   {repo['description'][:80]}...")
        print(f"   ⭐ {repo['stars']} | 📝 {repo['language']}")
        print()

案例 2:電商產品價格監控

PYTHON
from scrapling.fetchers import StealthyFetcher
import json
from datetime import datetime

def monitor_prices():
    urls = [
        "https://amazon.com/dp/B08N5WRWNW",
        "https://amazon.com/dp/B0BSHF7WHW",
        "https://amazon.com/dp/B09G9FPHY6",
    ]

    results = []

    for url in urls:
        page = StealthyFetcher.fetch(url, headless=True)

        title = page.css('#productTitle::text').get('').strip()
        price = page.css('.a-price .a-offscreen::text').get('').strip()

        results.append({
            "url": url,
            "title": title,
            "price": price,
            "timestamp": datetime.now().isoformat()
        })

        print(f"✅ {title[:50]}... - {price}")

    # 儲存到 JSON
    with open('price_monitor.json', 'w', encoding='utf-8') as f:
        json.dump(results, f, ensure_ascii=False, indent=2)

    print(f"\n📊 已儲存 {len(results)} 條價格資料到 price_monitor.json")

if __name__ == "__main__":
    monitor_prices()

案例 3:串流輸出(即時處理)

對於長時間執行的爬蟲,可以使用串流模式即時處理資料。

PYTHON
from scrapling.spiders import Spider, Response

class StreamingSpider(Spider):
    name = "streaming-demo"
    start_urls = ["https://quotes.toscrape.com/"]

    async def parse(self, response: Response):
        for quote in response.css('.quote'):
            item = {
                "text": quote.css('.text::text').get(),
                "author": quote.css('.author::text').get()
            }
            yield item  # 立即產出,無需等待全部完成

        next_page = response.css('.next a::attr(href)').get()
        if next_page:
            yield self.follow(next_page, callback=self.parse)

# 串流消費
async def main():
    spider = StreamingSpider()

    async for item in spider.stream():
        # 即時處理每個 item
        print(f"收到: {item['text'][:50]}... by {item['author']}")
        # 可以立即存入資料庫、傳送到訊息佇列等

import asyncio
asyncio.run(main())

進階技巧

1. 代理輪換

PYTHON
from scrapling.spiders import Spider, Response

class ProxySpider(Spider):
    name = "proxy-spider"
    start_urls = ["https://httpbin.org/ip"]

    custom_settings = {
        "proxy_list": [
            "http://proxy1.example.com:8080",
            "http://proxy2.example.com:8080",
            "http://proxy3.example.com:8080",
        ],
        "proxy_rotation": "per_request",  # 每個請求輪換代理
    }

    async def parse(self, response: Response):
        ip = response.json().get('origin')
        print(f"目前 IP: {ip}")

2. 自訂匯出管道

PYTHON
from scrapling.spiders import Spider, Response
import csv

class CSVSpider(Spider):
    name = "csv-export"
    start_urls = ["https://quotes.toscrape.com/"]

    def __init__(self, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.csv_file = open('quotes.csv', 'w', newline='', encoding='utf-8')
        self.writer = csv.writer(self.csv_file)
        self.writer.writerow(['Text', 'Author', 'Tags'])

    async def parse(self, response: Response):
        for quote in response.css('.quote'):
            text = quote.css('.text::text').get()
            author = quote.css('.author::text').get()
            tags = ', '.join(quote.css('.tag::text').getall())

            self.writer.writerow([text, author, tags])

        next_page = response.css('.next a::attr(href)').get()
        if next_page:
            yield self.follow(next_page, callback=self.parse)

    def close(self, reason):
        self.csv_file.close()
        print(f"✅ 資料已儲存到 quotes.csv")

3. 開發模式(快取回應)

在除錯解析邏輯時,避免重複請求伺服器。

PYTHON
from scrapling.spiders import Spider, Response

class DevSpider(Spider):
    name = "dev-mode"
    start_urls = ["https://example.com"]

    custom_settings = {
        "dev_mode": True,              # 啟用開發模式
        "cache_dir": "./http_cache",   # 快取目錄
    }

    async def parse(self, response: Response):
        # 第一次執行:快取回應到磁碟
        # 後續執行:直接從磁碟讀取,無需網路請求
        title = response.css('h1::text').get()
        yield {"title": title}

常見問題

Q1: Scrapling 和 Scrapy 有什麼區別?

A: Scrapling 可以看作是 Scrapy 的現代化增強版: - 自適應解析:Scrapy 的選擇器在頁面改版後會失效,Scrapling 能自動適應 - 內建反反爬:Scrapy 需要額外中介軟體才能繞過 Cloudflare,Scrapling 開箱即用 - 更簡潔的 API:Scrapling 的 Fetcher API 更適合小規模快速抓取

如果你已經熟悉 Scrapy,遷移到 Scrapling 幾乎沒有學習成本。

Q2: 如何處理登入後的頁面?

A: 使用 StealthyFetcher 或 DynamicFetcher 模擬登入:

PYTHON
from scrapling.fetchers import DynamicFetcher

page = DynamicFetcher.fetch(
    'https://example.com/login',
    headless=True,
    wait_for='#dashboard'  # 等待登入後跳轉
)

# 執行登入操作(透過 Playwright)
page.page.fill('#username', 'your_username')
page.page.fill('#password', 'your_password')
page.page.click('#login-button')
page.page.wait_for_selector('#dashboard')

# 現在可以抓取登入後的內容
data = page.css('.private-data::text').getall()

Q3: 如何限制爬取速度,避免被封?

A: 在 Spider 中設定下載延遲和並行限制:

PYTHON
class PoliteSpider(Spider):
    custom_settings = {
        "concurrency": 2,           # 降低並行數
        "download_delay": 2,        # 每個請求間隔 2 秒
        "robots_txt_obey": True,    # 遵守 robots.txt
    }

Q4: Scrapling 支援哪些選擇器?

A: 支援 CSS 選擇器和 XPath:

PYTHON
# CSS 選擇器
page.css('.class-name::text').get()
page.css('#id-name').getall()

# XPath
page.xpath('//div[@class="example"]/text()').get()

總結

Scrapling 是 2026 年最值得關注的 Python 爬蟲框架。它將自適應解析、反反爬能力和Scrapy-like API完美結合,讓開發者可以用最少的程式碼實現最穩定的爬蟲。

核心優勢回顧: - ✅ 自適應解析器:頁面改版後自動重新定位元素 - ✅ 內建反反爬:開箱即用繞過 Cloudflare Turnstile - ✅ 從單請求到全量爬取:同一套 API 覆蓋所有場景 - ✅ 並行、暫停/恢復、串流輸出:生產級特性一應俱全

資源連結: - GitHub 儲存庫 - 官方文件 - Discord 社群

如果你覺得這篇文章有幫助,歡迎分享給更多開發者!