跳到正文

kloxeld

xscrape

Async Python client for collecting public data from X (Twitter) - search, profiles, threads, replies, and media. Built-in account rotation, adaptive rate limiting, and export to JSON/CSV/SQLite.

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

xscrape

Asynchronous Python client for collecting public data from X (Twitter).

图片:Python 图片:License 图片:Build 图片:Version 图片:Docker

xscrape is a lightweight library for asynchronously collecting public data from X (Twitter): posts, profiles, threads, replies, and media. It supports account rotation, a built-in rate limiter, and export to JSON / CSV / SQLite.


Table of Contents

  • Features
  • Installation
  • Quick Start
  • How It Works
  • Configuration
  • CLI
  • Examples
  • Docker
  • Testing
  • Project Layout
  • Roadmap
  • Contributing
  • Disclaimer
  • License

Features

  • 🔍 Post search — by keywords, hashtags, and operators (from:, since:, until:)
  • 👤 User profiles — metadata, followers, activity counters
  • 🧵 Threads and replies — reconstruction of conversation chains
  • 🖼 Media — extraction of image and video links
  • 🔄 Account rotation — session pool with automatic failover
  • ⏱ Rate limiting — adaptive control of request frequency
  • 💾 Storage backends — JSON, CSV, SQLite out of the box
  • 🧩 Plugins — custom handlers and exporters
  • 🖥 CLI — ready-to-use command line interface
  • 🐳 Docker-ready — single command deployment
  • 📊 Structured logging — JSON logs with request tracing

📦 Installation

From source

git clone https://github.com/kloxeld/xscrape.git
cd xscrape
pip install -e .

Requirements

  • Python 3.10+
  • aiohttp, pydantic, tenacity, orjson

Optional

pip install "xscrape[socks]"    # SOCKS proxy support
pip install "xscrape[dev]"      # development tools
pip install "xscrape[docs]"     # documentation builders

Quick Start


# Search posts
xscrape search "python asyncio" --limit 50 --out tweets.json

# User profile
xscrape user elonmusk

# User timeline
xscrape timeline elonmusk --limit 200 --out timeline.csv

# Reconstruct a thread
xscrape thread 1234567890123456789 --out thread.json

# Collect by hashtag into SQLite
xscrape hashtag "#opensource" --limit 1000 --db hashtag.db

# Multi-account pool
XSCRAPE_POOL=accounts.json xscrape search "data engineering" --limit 2000

Run xscrape --help for the full command reference.

How It Works

xscrape talks to public GraphQL endpoints of X using session cookies. The pipeline looks like this:

┌────────────┐   ┌──────────────┐   ┌────────────┐   ┌────────────┐
│  Client    │──▶│  AuthPool    │──▶│  Fetcher   │──▶│  Parser    │
└────────────┘   └──────────────┘   └────────────┘   └────────────┘
                                            │                │
                                            ▼                ▼
                                     ┌────────────┐   ┌────────────┐
                                     │ RateLimiter│   │  Storage   │
                                     └────────────┘   └────────────┘
  1. Client — public interface (search, user, thread, replies).
  2. AuthPool — session pool; picks a free account, handles 429/401.
  3. Fetcher — low-level HTTP requests with retries and exponential backoff.
  4. Parser — normalizes raw responses into typed models (Tweet, User).
  5. RateLimiter — per-account token bucket plus a global cap.
  6. Storage — serialization of results into the chosen format.

See docs/ARCHITECTURE.md for details.


Configuration

All options are read from environment variables (see .env.example):

VariableDescriptionDefault
XSCRAPE_COOKIESCookie string (auth_token, ct0)—
XSCRAPE_POOLPath to JSON with account poolNone
XSCRAPE_CONCURRENCYMax parallel requests4
XSCRAPE_TIMEOUTRequest timeout (seconds)20
XSCRAPE_RETRIESNumber of retries on error3
XSCRAPE_USER_AGENTCustom User-Agentbuilt-in
XSCRAPE_PROXYProxy (http://user:pass@host:port)None
XSCRAPE_LOG_LEVELLogging levelINFO
XSCRAPE_LOG_FORMATtext or jsontext

Full reference: docs/CONFIGURATION.md.


Examples

The examples/ directory contains ready-to-run scripts:

  • search_tweets.py — search with pagination and filters
  • user_timeline.py — collect a user’s timeline
  • export_to_csv.py — dump results to CSV
  • export_to_sqlite.py — persist results into SQLite
  • thread_dump.py — reconstruct a full thread
  • hashtag_monitor.py — long-running hashtag watcher
  • multi_account_pool.py — usage of an account pool

Run:

python examples/search_tweets.py --query "openai" --limit 200

🐳 Docker

Build

docker build -f docker/Dockerfile -t xscrape:latest .

Run with docker-compose

cp .env.example .env
docker compose up --build

Development stack

docker compose -f docker-compose.dev.yml up --build

See docs/EXAMPLES.md for advanced Docker workflows.


Testing

pytest -q                    # run everything
pytest tests/unit            # unit tests only
pytest tests/integration     # integration tests only

Coverage:

pytest --cov=xscrape --cov-report=html

Project Layout

xscrape/
├── xscrape/          # library source
│   ├── client.py     # public client
│   ├── auth.py       # session pool
│   ├── parser.py     # response parsing
│   ├── ratelimit.py  # rate limiter
│   ├── storage.py    # storage backends
│   ├── plugins/      # plugin system
│   └── exporters/    # pluggable exporters
├── tests/            # unit + integration tests
├── examples/         # ready-to-run scripts
├── docs/             # documentation
├── docker/           # Dockerfiles
├── scripts/          # helper shell scripts
└── .github/          # CI workflows, templates

Roadmap

  • Search and profiles
  • Account pool and rate limiter
  • JSON / CSV / SQLite export
  • CLI
  • Docker support
  • Plugin exporter system
  • Media download support
  • Webhook notifications
  • Web monitoring dashboard
  • Prometheus metrics
  • GraphQL query cache

Full roadmap: docs/ROADMAP.md.


Contributing

We welcome contributions. Please read CONTRIBUTING.md and CODE_OF_CONDUCT.md before opening a PR.


⚠️ Disclaimer

This project is intended for educational purposes and work with public data only. Use it in accordance with the laws of your jurisdiction and the platform’s rules. The authors are not responsible for any consequences of use.


License

MIT — see LICENSE.

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。