跳到正文

vancyland

DataClaw0

DataClaw: Agentic Tailoring Multimodal Data from Raw Streams — coming soon (code, weights, dataset & DataClaw-val upon acceptance).

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

Actively refining and structuring raw multimodal data to align with diverse user and downstream intents.

图片:Paper 图片:Project Page 图片:Hugging Face 图片:License 图片:Status

[!NOTE] The code, model weights, dataset, and the DataClaw-val benchmark will be released upon paper acceptance. In the meantime, read the method in the paper and explore the qualitative cases on the project page.


📰 News

  • 2026-06 — DataClaw v1 paper released on arXiv 📄
  • 2026-06 — Project page with qualitative cases across five domains is live 🌐
  • Upcoming — Code, weights, dataset, and the DataClaw-val benchmark will be released upon acceptance ⏳

🐾 Overview

Massive unstructured multimodal streams suffer from high “data entropy,” impeding both efficient human knowledge acquisition and high-quality AI post-training. Existing passive annotation paradigms — heavily reliant on heuristic rules or general VLMs — are costly, monotonous, and fail to unlock the deep procedural logic embedded in raw data.

DataClaw elevates data processing to a learnable, high-order capability. We propose a paradigm shift towards Agentic Data Tailoring: given a user intent or downstream objective, a 9B tailoring agent filters redundant signal from long videos, GUI traces, embodied trajectories, and editing sequences, then reorganizes the residual into dense, verifiable, application-specific supervision.

Key ideas:

  • Bottom-up Factual Anchors → Top-down Semantic Synthesis. A two-stage pipeline grounds generative semantic synthesis in deterministic factual anchors, yielding a large-scale dataset spanning five core physical and digital domains.
  • SFT + rule-driven GRPO. DataClaw-9B synergizes Supervised Fine-Tuning with Group Relative Policy Optimization to robustly align with complex refinement and tailoring intents.
  • Two deployment paradigms. A single unified Omni model (DataClaw-O) or a panel of domain Experts (DataClaw-E).
  • DataClaw-val. The first benchmark dedicated to data refinement, scoring outputs by JSON validity and schema-aware Field / Semantic / Sequence metrics.
  • Downstream post-training as the ultimate touchstone. Validated on video generation, real-world VQA, and GUI navigation under volume-aligned training budgets.

DataClaw pipeline: bottom-up factual anchor extraction and top-down semantic synthesis, followed by training under the Omni and Expert paradigms, inference, and downstream utilization.


🎬 Qualitative Cases

Interactive case replays across five domains — daily life, education, GUI agents, embodied, and AIGC — with input videos and the agent’s reasoning, are available on the project page →


📊 Results (from the v1 paper)

DataClaw-E is the routed expert configuration; DataClaw-O is the unified omni model.

DataClaw-val — structured-output quality (Field / Semantic / Sequence)

ModelFieldSemanticSequence
Gemini-3.1-Pro98.1273.8558.50
GPT-4o97.2775.1549.43
DataClaw-E (Ours)97.5374.9448.86
DataClaw-O (Ours)87.6562.4644.82

Targeted Refinement — downstream SFT (same raw streams, same budget, only the annotator changes)

Downstream taskMetricSFT on Gemini-3.1-ProSFT on DataClaw
GUI navigation (AgentNet)SSR ↑ / TSR ↑39.5 / 14.238.2 / 15.6
Action video gen (Ego4D)FVD ↓ / Contact mAP ↑295.4 / 48.5288.6 / 51.2
Spatio-temporal VQA (ReMoT)Partial / Overall ↑53.4 / 31.552.1 / 33.2

DataClaw-E matches frontier VLMs on schema quality and leads on end-to-end downstream task success. Full tables, ablations, scaling curves, and t-SNE diversity analysis are in the paper.


🗺️ Release Roadmap

Everything below ships upon paper acceptance (v2).

  • 📄 Paper (v1) on arXiv
  • 🌐 Project page with qualitative cases & demos
  • 🧩 Code — training (SFT + GRPO) and inference
  • 🏋️ Model weights — DataClaw-O (Omni) & DataClaw-E (Experts)
  • 📦 Dataset — five domains (daily life, education, embodied, GUI agents, AIGC)
  • 📐 DataClaw-val benchmark + evaluation scripts
  • 📝 Reproduction recipes for downstream SFT tasks

⭐ Star / watch this repo to be notified the moment the code and data drop.


📌 Citation

If you find DataClaw useful for your research, please consider citing:

@article{wan2026dataclaw,
  title   = {DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams},
  author  = {Wan, Cong and Guo, Zeyu and Cai, Zijian and Li, Jiangyang and
             Dong, SongLin and Peng, Lin and Luo, Xiangyang and Ma, Zhiheng and Gong, Yihong},
  journal = {arXiv preprint arXiv:2606.21337},
  year    = {2026}
}

📬 Contact

Questions, collaboration, or follow-up? Open an issue or reach the authors via the contacts listed on the paper.

License

The license for the code and released artifacts will be announced together with the open-source release upon acceptance.

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。