Scrapling 是什麼?
Scrapling 是一個自適應的 Python 網頁爬蟲框架,由開發者 D4Vinci 建立。它在 2026 年 GitHub Trending 上迅速走紅,獲得超過 57,000 個 Star,成為新一代爬蟲工具的代表。
為什麼需要 Scrapling?
傳統爬蟲框架(如 Scrapy、BeautifulSoup)面臨三大痛點:
- 網站結構變化導致程式碼失效:CSS 選擇器或 XPath 一旦頁面改版就全部崩潰
- 反爬系統攔截:Cloudflare Turnstile、Akamai 等防護讓普通請求直接返回 403
- 擴充性差:從小規模抓取到大規模並行爬取,需要重寫大量程式碼
Scrapling 的核心設計理念是 "One library, zero compromises"(一個庫,零妥協):
- 自適應解析器:自動學習網站結構,頁面更新後自動重新定位元素
- 內建反反爬:開箱即用,繞過 Cloudflare Turnstile 等主流防護
- Scrapy-like API:熟悉 Scrapy 的開發者可以無縫遷移
- 從單請求到全量爬取:同一套 API 支援簡單抓取和大規模並行爬蟲
Scrapling vs 傳統爬蟲工具
| 特性 | BeautifulSoup | Scrapy | Scrapling |
|---|---|---|---|
| 學習曲線 | 低 | 中 | 中 |
| 自適應解析 | ❌ | ❌ | ✅ |
| 繞過 Cloudflare | ❌ | 需外掛 | ✅ 內建 |
| 並行爬取 | ❌ | ✅ | ✅ |
| 動態頁面支援 | ❌ | 需中介軟體 | ✅ 內建 Playwright |
| 暫停/恢復 | ❌ | 需擴充套件 | ✅ 內建 |
| 串流輸出 | ❌ | ❌ | ✅ |
安裝 Scrapling
環境要求
- Python 3.8+
- pip 套件管理器
快速安裝
pip install scrapling
驗證安裝
from scrapling.fetchers import Fetcher
# 測試基本功能
p = Fetcher.fetch('https://example.com')
print(p.title) # 輸出頁面標題
可選依賴
如果需要動態頁面渲染(JavaScript 網站),需要安裝 Playwright:
pip install playwright
playwright install chromium
快速上手:第一個爬蟲
範例 1:簡單頁面抓取
讓我們從一個簡單的例子開始——抓取 Hacker News 的頭條新聞。
from scrapling.fetchers import Fetcher
# 發起 HTTP 請求
page = Fetcher.fetch('https://news.ycombinator.com/')
# 使用 CSS 選擇器提取資料
stories = page.css('.titleline > a')
for story in stories[:5]: # 只取前 5 條
title = story.text
link = story.attrs.get('href', '')
print(f"標題: {title}")
print(f"連結: {link}")
print("-" * 40)
輸出範例:
標題: Show HN: I built a real-time code collaboration tool
連結: https://github.com/example/collab-tool
----------------------------------------
標題: Ask HN: What's your favorite Python library in 2026?
連結: https://news.ycombinator.com/item?id=123456
----------------------------------------
範例 2:自適應解析(Auto-save)
Scrapling 的核心特性是自適應解析。當你第一次抓取資料時,可以啟用 auto_save=True,Scrapling 會學習頁面結構並儲存特徵。當網站改版後,只需傳入 adaptive=True,它就能自動找到目標元素。
from scrapling.fetchers import Fetcher
page = Fetcher.fetch('https://quotes.toscrape.com/')
# 第一次抓取:啟用 auto_save
quotes = page.css('.quote', auto_save=True)
for quote in quotes[:3]:
text = quote.css('.text::text').get()
author = quote.css('.author::text').get()
print(f"{text} — {author}")
如果網站結構變化了:
# 後續抓取:傳入 adaptive=True
page = Fetcher.fetch('https://quotes.toscrape.com/')
quotes = page.css('.quote', adaptive=True) # 自動適應新結構!
for quote in quotes[:3]:
text = quote.css('.text::text').get()
author = quote.css('.author::text').get()
print(f"{text} — {author}")
💡 工作原理:Scrapling 會記錄元素的多種特徵(標籤類型、附近文字、屬性模式等),即使 CSS 類名改變,它也能透過其他特徵重新定位元素。
進階功能
1. 繞過 Cloudflare Turnstile
很多網站使用 Cloudflare Turnstile 或其他反爬系統。Scrapling 的 StealthyFetcher 可以自動處理這些防護。
from scrapling.fetchers import StealthyFetcher
# 啟用自適應模式
StealthyFetcher.adaptive = True
# 自動繞過 Cloudflare
page = StealthyFetcher.fetch(
'https://example-protected-site.com',
headless=True, # 無頭瀏覽器模式
network_idle=True # 等待網路空閒
)
# 正常提取資料
products = page.css('.product-item')
for product in products:
name = product.css('.name::text').get()
price = product.css('.price::text').get()
print(f"{name}: {price}")
關鍵參數說明:
- headless=True:使用無頭瀏覽器,模擬真實使用者行為
- network_idle=True:等待所有網路完成後再提取(適合 SPA 應用)
- adaptive=True:啟用自適應解析
2. 動態頁面渲染(Playwright)
對於需要 JavaScript 渲染的網站,使用 DynamicFetcher。
from scrapling.fetchers import DynamicFetcher
page = DynamicFetcher.fetch(
'https://spa-example.com',
wait_for='.content-loaded', # 等待特定元素出現
timeout=30000 # 逾時時間(毫秒)
)
# 提取動態載入的內容
articles = page.css('article')
for article in articles:
title = article.css('h2::text').get()
summary = article.css('.summary::text').get()
print(f"{title}\n{summary}\n")
3. 非同步並行抓取
使用 AsyncFetcher 實現高並行抓取。
import asyncio
from scrapling.fetchers import AsyncFetcher
async def fetch_multiple_pages():
urls = [
'https://example.com/page/1',
'https://example.com/page/2',
'https://example.com/page/3',
'https://example.com/page/4',
'https://example.com/page/5',
]
# 並行發起請求
pages = await AsyncFetcher.fetch_many(urls, concurrency=3)
for url, page in zip(urls, pages):
if page:
title = page.title
print(f"{url}: {title}")
else:
print(f"{url}: 請求失敗")
asyncio.run(fetch_multiple_pages())
Spider 框架:大規模爬取
Scrapling 提供了類似 Scrapy 的 Spider 框架,支援大規模並行爬取。
基礎 Spider
from scrapling.spiders import Spider, Response
class QuoteSpider(Spider):
name = "quotes"
start_urls = ["https://quotes.toscrape.com/"]
async def parse(self, response: Response):
# 提取目前頁面的名言
for quote in response.css('.quote'):
yield {
"text": quote.css('.text::text').get(),
"author": quote.css('.author::text').get(),
"tags": quote.css('.tag::text').getall()
}
# 翻頁
next_page = response.css('.next a::attr(href)').get()
if next_page:
yield self.follow(next_page, callback=self.parse)
# 啟動爬蟲
QuoteSpider().start()
並行設定
class MultiPageSpider(Spider):
name = "multi-page"
start_urls = [f"https://example.com/page/{i}" for i in range(1, 101)]
# 設定並行數
custom_settings = {
"concurrency": 5, # 最大並行數
"download_delay": 1, # 下載延遲(秒)
"robots_txt_obey": True, # 遵守 robots.txt
}
async def parse(self, response: Response):
title = response.css('h1::text').get()
yield {"url": response.url, "title": title}
MultiPageSpider().start()
暫停與恢復
Scrapling 支援 checkpoint-based 持久化,按 Ctrl+C 優雅退出後,下次啟動會自動恢復。
class LongRunningSpider(Spider):
name = "long-crawl"
start_urls = ["https://large-site.com/"]
custom_settings = {
"checkpoint_dir": "./checkpoints", # 檢查點目錄
}
async def parse(self, response: Response):
# 提取資料...
yield {"data": "..."}
# 繼續爬取
for link in response.css('a::attr(href)').getall():
yield self.follow(link, callback=self.parse)
LongRunningSpider().start()
實戰案例
案例 1:抓取 GitHub Trending 專案
from scrapling.fetchers import Fetcher
def scrape_github_trending():
page = Fetcher.fetch('https://github.com/trending')
repos = page.css('.Box-row')
trending = []
for repo in repos[:10]:
name = repo.css('h2 a::text').get('').strip()
description = repo.css('p.col-9::text').get('').strip()
stars = repo.css('[href$=stargazers] span::text').get('').strip()
language = repo.css('[itemprop=programmingLanguage]::text').get('').strip()
trending.append({
"name": name,
"description": description,
"stars": stars,
"language": language
})
return trending
if __name__ == "__main__":
results = scrape_github_trending()
for repo in results:
print(f"📦 {repo['name']}")
print(f" {repo['description'][:80]}...")
print(f" ⭐ {repo['stars']} | 📝 {repo['language']}")
print()
案例 2:電商產品價格監控
from scrapling.fetchers import StealthyFetcher
import json
from datetime import datetime
def monitor_prices():
urls = [
"https://amazon.com/dp/B08N5WRWNW",
"https://amazon.com/dp/B0BSHF7WHW",
"https://amazon.com/dp/B09G9FPHY6",
]
results = []
for url in urls:
page = StealthyFetcher.fetch(url, headless=True)
title = page.css('#productTitle::text').get('').strip()
price = page.css('.a-price .a-offscreen::text').get('').strip()
results.append({
"url": url,
"title": title,
"price": price,
"timestamp": datetime.now().isoformat()
})
print(f"✅ {title[:50]}... - {price}")
# 儲存到 JSON
with open('price_monitor.json', 'w', encoding='utf-8') as f:
json.dump(results, f, ensure_ascii=False, indent=2)
print(f"\n📊 已儲存 {len(results)} 條價格資料到 price_monitor.json")
if __name__ == "__main__":
monitor_prices()
案例 3:串流輸出(即時處理)
對於長時間執行的爬蟲,可以使用串流模式即時處理資料。
from scrapling.spiders import Spider, Response
class StreamingSpider(Spider):
name = "streaming-demo"
start_urls = ["https://quotes.toscrape.com/"]
async def parse(self, response: Response):
for quote in response.css('.quote'):
item = {
"text": quote.css('.text::text').get(),
"author": quote.css('.author::text').get()
}
yield item # 立即產出,無需等待全部完成
next_page = response.css('.next a::attr(href)').get()
if next_page:
yield self.follow(next_page, callback=self.parse)
# 串流消費
async def main():
spider = StreamingSpider()
async for item in spider.stream():
# 即時處理每個 item
print(f"收到: {item['text'][:50]}... by {item['author']}")
# 可以立即存入資料庫、傳送到訊息佇列等
import asyncio
asyncio.run(main())
進階技巧
1. 代理輪換
from scrapling.spiders import Spider, Response
class ProxySpider(Spider):
name = "proxy-spider"
start_urls = ["https://httpbin.org/ip"]
custom_settings = {
"proxy_list": [
"http://proxy1.example.com:8080",
"http://proxy2.example.com:8080",
"http://proxy3.example.com:8080",
],
"proxy_rotation": "per_request", # 每個請求輪換代理
}
async def parse(self, response: Response):
ip = response.json().get('origin')
print(f"目前 IP: {ip}")
2. 自訂匯出管道
from scrapling.spiders import Spider, Response
import csv
class CSVSpider(Spider):
name = "csv-export"
start_urls = ["https://quotes.toscrape.com/"]
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.csv_file = open('quotes.csv', 'w', newline='', encoding='utf-8')
self.writer = csv.writer(self.csv_file)
self.writer.writerow(['Text', 'Author', 'Tags'])
async def parse(self, response: Response):
for quote in response.css('.quote'):
text = quote.css('.text::text').get()
author = quote.css('.author::text').get()
tags = ', '.join(quote.css('.tag::text').getall())
self.writer.writerow([text, author, tags])
next_page = response.css('.next a::attr(href)').get()
if next_page:
yield self.follow(next_page, callback=self.parse)
def close(self, reason):
self.csv_file.close()
print(f"✅ 資料已儲存到 quotes.csv")
3. 開發模式(快取回應)
在除錯解析邏輯時,避免重複請求伺服器。
from scrapling.spiders import Spider, Response
class DevSpider(Spider):
name = "dev-mode"
start_urls = ["https://example.com"]
custom_settings = {
"dev_mode": True, # 啟用開發模式
"cache_dir": "./http_cache", # 快取目錄
}
async def parse(self, response: Response):
# 第一次執行:快取回應到磁碟
# 後續執行:直接從磁碟讀取,無需網路請求
title = response.css('h1::text').get()
yield {"title": title}
常見問題
Q1: Scrapling 和 Scrapy 有什麼區別?
A: Scrapling 可以看作是 Scrapy 的現代化增強版: - 自適應解析:Scrapy 的選擇器在頁面改版後會失效,Scrapling 能自動適應 - 內建反反爬:Scrapy 需要額外中介軟體才能繞過 Cloudflare,Scrapling 開箱即用 - 更簡潔的 API:Scrapling 的 Fetcher API 更適合小規模快速抓取
如果你已經熟悉 Scrapy,遷移到 Scrapling 幾乎沒有學習成本。
Q2: 如何處理登入後的頁面?
A: 使用 StealthyFetcher 或 DynamicFetcher 模擬登入:
from scrapling.fetchers import DynamicFetcher
page = DynamicFetcher.fetch(
'https://example.com/login',
headless=True,
wait_for='#dashboard' # 等待登入後跳轉
)
# 執行登入操作(透過 Playwright)
page.page.fill('#username', 'your_username')
page.page.fill('#password', 'your_password')
page.page.click('#login-button')
page.page.wait_for_selector('#dashboard')
# 現在可以抓取登入後的內容
data = page.css('.private-data::text').getall()
Q3: 如何限制爬取速度,避免被封?
A: 在 Spider 中設定下載延遲和並行限制:
class PoliteSpider(Spider):
custom_settings = {
"concurrency": 2, # 降低並行數
"download_delay": 2, # 每個請求間隔 2 秒
"robots_txt_obey": True, # 遵守 robots.txt
}
Q4: Scrapling 支援哪些選擇器?
A: 支援 CSS 選擇器和 XPath:
# CSS 選擇器
page.css('.class-name::text').get()
page.css('#id-name').getall()
# XPath
page.xpath('//div[@class="example"]/text()').get()
總結
Scrapling 是 2026 年最值得關注的 Python 爬蟲框架。它將自適應解析、反反爬能力和Scrapy-like API完美結合,讓開發者可以用最少的程式碼實現最穩定的爬蟲。
核心優勢回顧: - ✅ 自適應解析器:頁面改版後自動重新定位元素 - ✅ 內建反反爬:開箱即用繞過 Cloudflare Turnstile - ✅ 從單請求到全量爬取:同一套 API 覆蓋所有場景 - ✅ 並行、暫停/恢復、串流輸出:生產級特性一應俱全
資源連結: - GitHub 儲存庫 - 官方文件 - Discord 社群
如果你覺得這篇文章有幫助,歡迎分享給更多開發者!