kloxeld
xscrape
Async Python client for collecting public data from X (Twitter) - search, profiles, threads, replies, and media. Built-in account rotation, adaptive rate limiting, and export to JSON/CSV/SQLite.
Documentation snapshot
README 快照
翻译暂时拿不到。
机器翻译的项目简介,仅供参考。原文在下方,也可以直接用浏览器自带的整页翻译 (Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」)。
下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。
本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。
xscrape
Asynchronous Python client for collecting public data from X (Twitter).
图片:Python 图片:License 图片:Build 图片:Version 图片:Docker
xscrape is a lightweight library for asynchronously collecting public data from X (Twitter): posts, profiles, threads, replies, and media. It supports account rotation, a built-in rate limiter, and export to JSON / CSV / SQLite.
Table of Contents
- Features
- Installation
- Quick Start
- How It Works
- Configuration
- CLI
- Examples
- Docker
- Testing
- Project Layout
- Roadmap
- Contributing
- Disclaimer
- License
Features
- 🔍 Post search — by keywords, hashtags, and operators (
from:,since:,until:) - 👤 User profiles — metadata, followers, activity counters
- 🧵 Threads and replies — reconstruction of conversation chains
- 🖼 Media — extraction of image and video links
- 🔄 Account rotation — session pool with automatic failover
- ⏱ Rate limiting — adaptive control of request frequency
- 💾 Storage backends — JSON, CSV, SQLite out of the box
- 🧩 Plugins — custom handlers and exporters
- 🖥 CLI — ready-to-use command line interface
- 🐳 Docker-ready — single command deployment
- 📊 Structured logging — JSON logs with request tracing
📦 Installation
From source
git clone https://github.com/kloxeld/xscrape.git
cd xscrape
pip install -e .
Requirements
- Python 3.10+
aiohttp,pydantic,tenacity,orjson
Optional
pip install "xscrape[socks]" # SOCKS proxy support
pip install "xscrape[dev]" # development tools
pip install "xscrape[docs]" # documentation builders
Quick Start
# Search posts
xscrape search "python asyncio" --limit 50 --out tweets.json
# User profile
xscrape user elonmusk
# User timeline
xscrape timeline elonmusk --limit 200 --out timeline.csv
# Reconstruct a thread
xscrape thread 1234567890123456789 --out thread.json
# Collect by hashtag into SQLite
xscrape hashtag "#opensource" --limit 1000 --db hashtag.db
# Multi-account pool
XSCRAPE_POOL=accounts.json xscrape search "data engineering" --limit 2000
Run xscrape --help for the full command reference.
How It Works
xscrape talks to public GraphQL endpoints of X using session cookies. The pipeline looks like this:
┌────────────┐ ┌──────────────┐ ┌────────────┐ ┌────────────┐
│ Client │──▶│ AuthPool │──▶│ Fetcher │──▶│ Parser │
└────────────┘ └──────────────┘ └────────────┘ └────────────┘
│ │
▼ ▼
┌────────────┐ ┌────────────┐
│ RateLimiter│ │ Storage │
└────────────┘ └────────────┘
- Client — public interface (
search,user,thread,replies). - AuthPool — session pool; picks a free account, handles 429/401.
- Fetcher — low-level HTTP requests with retries and exponential backoff.
- Parser — normalizes raw responses into typed models (
Tweet,User). - RateLimiter — per-account token bucket plus a global cap.
- Storage — serialization of results into the chosen format.
See docs/ARCHITECTURE.md for details.
Configuration
All options are read from environment variables (see .env.example):
| Variable | Description | Default |
|---|---|---|
XSCRAPE_COOKIES | Cookie string (auth_token, ct0) | — |
XSCRAPE_POOL | Path to JSON with account pool | None |
XSCRAPE_CONCURRENCY | Max parallel requests | 4 |
XSCRAPE_TIMEOUT | Request timeout (seconds) | 20 |
XSCRAPE_RETRIES | Number of retries on error | 3 |
XSCRAPE_USER_AGENT | Custom User-Agent | built-in |
XSCRAPE_PROXY | Proxy (http://user:pass@host:port) | None |
XSCRAPE_LOG_LEVEL | Logging level | INFO |
XSCRAPE_LOG_FORMAT | text or json | text |
Full reference: docs/CONFIGURATION.md.
Examples
The examples/ directory contains ready-to-run scripts:
search_tweets.py— search with pagination and filtersuser_timeline.py— collect a user’s timelineexport_to_csv.py— dump results to CSVexport_to_sqlite.py— persist results into SQLitethread_dump.py— reconstruct a full threadhashtag_monitor.py— long-running hashtag watchermulti_account_pool.py— usage of an account pool
Run:
python examples/search_tweets.py --query "openai" --limit 200
🐳 Docker
Build
docker build -f docker/Dockerfile -t xscrape:latest .
Run with docker-compose
cp .env.example .env
docker compose up --build
Development stack
docker compose -f docker-compose.dev.yml up --build
See docs/EXAMPLES.md for advanced Docker workflows.
Testing
pytest -q # run everything
pytest tests/unit # unit tests only
pytest tests/integration # integration tests only
Coverage:
pytest --cov=xscrape --cov-report=html
Project Layout
xscrape/
├── xscrape/ # library source
│ ├── client.py # public client
│ ├── auth.py # session pool
│ ├── parser.py # response parsing
│ ├── ratelimit.py # rate limiter
│ ├── storage.py # storage backends
│ ├── plugins/ # plugin system
│ └── exporters/ # pluggable exporters
├── tests/ # unit + integration tests
├── examples/ # ready-to-run scripts
├── docs/ # documentation
├── docker/ # Dockerfiles
├── scripts/ # helper shell scripts
└── .github/ # CI workflows, templates
Roadmap
- Search and profiles
- Account pool and rate limiter
- JSON / CSV / SQLite export
- CLI
- Docker support
- Plugin exporter system
- Media download support
- Webhook notifications
- Web monitoring dashboard
- Prometheus metrics
- GraphQL query cache
Full roadmap: docs/ROADMAP.md.
Contributing
We welcome contributions. Please read CONTRIBUTING.md and CODE_OF_CONDUCT.md before opening a PR.
⚠️ Disclaimer
This project is intended for educational purposes and work with public data only. Use it in accordance with the laws of your jurisdiction and the platform’s rules. The authors are not responsible for any consequences of use.
License
MIT — see LICENSE.
Official distribution
获取与安装
暂未发现可确认的官方软件包地址
当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。
本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。
Before installing
使用前核验
本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。