跳到正文

Lutschippi

DEHUB

The Ultimate Data Engineering Hub — 500+ Resources, 50+ Tools, Roadmaps & Community for Data Engineers Worldwide

README 已保存到本站,可直接阅读

Documentation snapshot

README 快照

这篇是英文原文

下面正文是项目自己的英文 README。想读全文就用浏览器自带的整页翻译: Chrome / Edge 点地址栏右侧的翻译图标,或用右键菜单里的「翻译成中文」; 手机浏览器一般在菜单里。

本页保存的是公开项目资料快照,阅读过程不需要连接 GitHub。

图片:GitHub Stars 图片:GitHub Forks 图片:GitHub Watchers 图片:License 图片:Website 图片:PRs Welcome 图片:Contributions

$ pip install dehub-knowledge --upgrade
✓ Loaded: 500+ Resources | 50+ Tools | 10+ Roadmaps | 1 Community

The command center for data engineers worldwide. From zero to petabyte-scale — everything you need to become an elite data engineer.

🚀 Get Started • 📚 Resources • 🗺️ Roadmap • 🛠️ Tools • 🌐 Website • 🤝 Contribute


📡 Live Terminal Preview

╔══════════════════════════════════════════════════════════════╗
║  dehub@engineer:~$                                           ║
╠══════════════════════════════════════════════════════════════╣
║                                                              ║
║  $ dbt run --select +orders_mart                             ║
║  [INFO] Running with dbt=1.8.0                               ║
║  [INFO] Found 47 models, 312 tests                          ║
║  ✓ Completed successfully                                    ║
║                                                              ║
║  $ spark-submit --master yarn pipeline.py                   ║
║  [INFO] SparkContext initialized                             ║
║  [INFO] Processing 2.4B records...                          ║
║  ✓ Job completed: 847s elapsed                              ║
║                                                              ║
║  $ kafka-topics --create --topic events --partitions 12     ║
║  ✓ Created topic "events"                                   ║
║                                                              ║
║  $ SELECT COUNT(*) FROM iceberg.prod.events;                ║
║  > 1,427,893,004 rows                                       ║
║                                                              ║
╚══════════════════════════════════════════════════════════════╝

📋 Table of Contents

Click to expand

  • 🚀 Getting Started
  • 🗺️ Roadmap
  • 📚 Books
  • 🎓 Courses & Bootcamps
  • 🛠️ Data Engineering Ecosystem
  • 💼 Projects & Portfolio
  • 🌐 Communities
  • 🎙️ Podcasts
  • 📰 Newsletters
  • 🎥 YouTube Channels
  • 💼 Interview Preparation
  • 📊 Data Engineering Salary
  • 🤝 Contributing
  • 📄 License

🚀 Getting Started

New to Data Engineering? Follow this path:

📍 You are here
    │
    ▼
[1] 📖 Read "Fundamentals of Data Engineering" (Joe Reis & Matt Housley)
    │
    ▼
[2] 🐍 Learn Python + SQL to professional level
    │
    ▼
[3] 🌊 Build your first pipeline (Batch → Streaming)
    │
    ▼
[4] ☁️  Get cloud certified (AWS/GCP/Azure)
    │
    ▼
[5] 🚀 Contribute to open-source, get hired

Experienced Engineer? Jump to:

  • Advanced Tools & Frameworks
  • System Design Resources
  • Senior-level Interview Prep
  • Open Source Projects

🗺️ Roadmap

Beginner Path (0–6 months)

SkillResourcesStatus
Python for Data EngineeringPython for Data Analysis🔥 Essential
SQL MasteryMode SQL Tutorial🔥 Essential
Linux & BashThe Linux Command Line✅ Required
Git & Version ControlPro Git Book✅ Required
Basic ETL ConceptsFundamentals of Data Engineering🔥 Essential
Docker BasicsDocker Official Docs✅ Required

Intermediate Path (6–18 months)

SkillResourcesStatus
Apache SparkSpark: The Definitive Guide🔥 Essential
Apache AirflowAstronomer Guides🔥 Essential
dbt (Data Build Tool)dbt Learn🔥 Essential
Cloud PlatformsAWS/GCP/Azure Certifications✅ Required
Data ModelingKimball - Data Warehouse Toolkit🔥 Essential
Kafka/StreamingConfluent Kafka Tutorials⭐ Recommended

Advanced Path (18+ months)

SkillResourcesStatus
Data ArchitectureDesigning Data-Intensive Applications🔥 Essential
Apache IcebergApache Iceberg: The Definitive Guide🔥 Essential
Real-Time StreamingStreaming Systems⭐ Recommended
Data MeshData Mesh (Zhamak Dehghani)⭐ Recommended
ML EngineeringDesigning Machine Learning Systems⭐ Recommended
System DesignByteByteGo✅ Required

Expert Path

SkillFocus
Data Platform ArchitectureDesign multi-cloud, multi-region data platforms
Cost OptimizationFinOps for data, query optimization at scale
Team LeadershipBuilding and mentoring data engineering teams
Open Source ContributionBuild tooling the community relies on

📚 Books — Must Read

Top 3 Essential Books

#BookAuthorLevel
1Fundamentals of Data EngineeringJoe Reis & Matt HousleyAll Levels
2Designing Data-Intensive ApplicationsMartin KleppmannIntermediate+
3Designing Machine Learning SystemsChip HuyenAdvanced

Complete Book List

View all 35+ books

Data Engineering Core
Apache Spark
Streaming
Cloud & Lakehouse
dbt & Transformation
Data Architecture
Machine Learning & AI
Analytics & Python

🎓 Courses & Bootcamps

Free Bootcamps

NameProviderLevelDuration
Data Engineering ZoomcampDataTalks.ClubBeginner9 weeks
Beginner Data Engineering BootcampDataExpert.ioBeginner4 weeks
Intermediate BootcampDataExpert.ioIntermediate6 weeks
DE FundamentalsDatabricksAll LevelsSelf-paced

Premium Courses

NameProviderLevel
Data Expert CoursesZach WilsonAll Levels
dbt Fundamentalsdbt LabsBeginner
Astronomer CertificationAstronomerIntermediate
Databricks Certified AssociateDatabricksIntermediate
Snowflake SnowPro CoreSnowflakeIntermediate
Google Cloud Professional Data EngineerGoogleAdvanced
AWS Data Analytics SpecialtyAWSAdvanced

🛠️ Data Engineering Ecosystem

Orchestration

ToolStarsDescription
Apache Airflow图片:StarsIndustry-standard workflow orchestrator
Dagster图片:StarsData-aware orchestration platform
Prefect图片:StarsModern workflow automation
Mage图片:StarsModern data pipeline tool
Kestra图片:StarsDeclarative orchestration
Hamilton图片:StarsFunction-based DAG framework

Data Lake / Lakehouse

ToolStarsDescription
Apache Iceberg图片:StarsOpen table format for huge datasets
Delta Lake图片:StarsACID transactions for big data
Apache Hudi图片:StarsIncremental data processing
Apache Polaris图片:StarsOpen catalog for Apache Iceberg
DuckLake-SQL-native lakehouse

Data Warehouse

ToolDescription
SnowflakeCloud-native data warehouse
Google BigQueryServerless data warehouse
DatabricksUnified analytics platform
Amazon RedshiftAWS data warehouse
FireboltUltra-fast cloud warehouse
DatabendOpen-source cloud DW
ClickHouseReal-time analytics

Processing Engines

ToolStarsDescription
Apache Spark图片:StarsUnified analytics engine
Apache Flink图片:StarsStream & batch processing
DuckDB图片:StarsIn-process analytical DB
Trino图片:StarsDistributed SQL query engine
Polars图片:StarsFast DataFrame library

Transformation

ToolStarsDescription
dbt图片:StarsSQL-first transformation
SQLMesh图片:StarsNext-gen dbt alternative
Coalesce-Cloud-native transformation

Data Quality

ToolStarsDescription
Great Expectations图片:StarsData quality framework
Soda图片:StarsData quality platform
dbt tests-Built-in dbt testing
Metaplane-Data observability
DQOps图片:StarsAutomated data quality

Streaming & Messaging

ToolDescription
Apache KafkaDistributed event streaming
Apache PulsarCloud-native messaging
RedpandaKafka-compatible streaming
ConfluentManaged Kafka platform
AWS KinesisReal-time data streaming

Data Catalog & Governance

ToolDescription
Apache AtlasData governance & metadata
DataHubModern metadata platform
OpenMetadataOpen-source data catalog
AmundsenData discovery & metadata

Ingestion & Integration

ToolDescription
AirbyteOpen-source data integration
FivetranAutomated data movement
DebeziumCDC (Change Data Capture)
Apache NiFiData flow automation
dltPython data load tool

Visualization

ToolDescription
Apache SupersetOpen-source BI
MetabaseBusiness intelligence
GrafanaObservability & analytics
RedashQuery & visualization

💼 Projects & Portfolio

Build these projects to demonstrate real-world data engineering skills:

Beginner Projects

  1. End-to-End NYC Taxi Data Pipeline

    • Tools: Python, BigQuery, Looker Studio
    • Skills: ETL, cloud storage, BI visualization
  2. Weather Data Pipeline

    • Tools: Airflow, PostgreSQL, dbt
    • Skills: Orchestration, scheduling, transformation
  3. Extract YouTube Metadata

    • GitHub Project
    • Tools: AWS Lambda, S3, Free Tier
    • Skills: Serverless, cloud storage, API ingestion

Intermediate Projects

  1. Real Estate Data Platform

    • GitHub
    • Tools: S3, Spark, Delta Lake, Dagster, Superset
    • Skills: Lakehouse, orchestration, visualization
  2. Azure End-to-End Analytics Platform

    • Tools: ADF, ADLS, Databricks, Synapse, Power BI
    • Skills: Azure ecosystem, medallion architecture
  3. LLM Data Pipeline

    • Lecture
    • Tools: OpenAI API, vector databases, Airflow
    • Skills: AI integration, vector search

Advanced Projects

  1. Real-Time Streaming Analytics

    • Tools: Kafka, Flink, Iceberg, Grafana
    • Skills: Event-driven architecture, stream processing
  2. SQL Query Engine with LLMs

    • Tutorial
    • Tools: LangChain, LLMs, databases
    • Skills: AI-powered tooling
  3. Multi-Cloud Lakehouse

    • Tools: Apache Iceberg, AWS + GCP, dbt, Airflow
    • Skills: Cloud-agnostic architecture

🌐 Communities

Join these communities to learn, network, and grow as a data engineer.

Discord Communities

CommunityMembersFocus
DataExpert.io Discord10,000+Data Engineering
AdalFlow-AI/ML Engineering
Chip Huyen MLOps10,000+ML Operations

Slack Communities

CommunityFocus
Data Talks ClubData Science & Engineering
dbt Communitydbt, Analytics Engineering
Great ExpectationsData Quality
PrefectWorkflow Orchestration

Online Communities

CommunityPlatformFocus
Data Engineer ThingsNewsletter/CommunityData Engineering
r/dataengineeringRedditData Engineering
r/apachesparkRedditApache Spark
LinkedIn DE CommunityLinkedInProfessional Networking

🎙️ Podcasts

PodcastHostTopics
The Data Engineering ShowDatabandDE tools & practices
Data Engineering PodcastTobias MaceyOpen-source data tools
DataTopics-Data engineering trends
DataWareAscend.ioData pipelines
The Datastack Show-Modern data stack
Analytics Power Hour-Analytics & data
Drill to DetailMark RittmanAnalytics engineering

📰 Newsletters

NewsletterAuthorTopics
DataEngineer.io NewsletterZach WilsonDE career & tech
The Developing DevRyan PetermanEngineering growth
Data Engineering WeeklyAnanth PackkilduraiDE news
Benn Stancil’s NewsletterBenn StancilData strategy
Ahead of the Trend-Data trends
Seattle Data GuyBen RogojanDE tips

🎥 YouTube Channels

ChannelFocusSubscribers
Zach WilsonData Engineering Career50,000+
Seattle Data GuyDE Interviews & Tips50,000+
Andreas KretzData Engineering School50,000+
ByteByteGoSystem Design1,000,000+
Alex The AnalystData Analysis700,000+

💼 Interview Preparation

Data Engineering Interview Topics

Technical Interview Topics:
├── SQL
│   ├── Window Functions (ROW_NUMBER, RANK, LAG, LEAD)
│   ├── CTEs and Recursive CTEs
│   ├── Query Optimization & EXPLAIN plans
│   └── Aggregations & Subqueries
├── Python
│   ├── PySpark DataFrames
│   ├── Pandas/Polars operations
│   ├── Generators & Iterators
│   └── OOP for data pipelines
├── System Design
│   ├── Design a data warehouse
│   ├── Design a real-time analytics system
│   ├── Design a CDC pipeline
│   └── Design a data lake
├── Data Modeling
│   ├── Star Schema vs Snowflake
│   ├── Slowly Changing Dimensions (SCD)
│   ├── Data Vault 2.0
│   └── Kimball vs Inmon
└── Infrastructure
    ├── Docker & Kubernetes
    ├── Cloud platforms (AWS/GCP/Azure)
    ├── CI/CD for data pipelines
    └── Monitoring & Alerting

Interview Resources

Resume Tips

  • Quantify impact: “Reduced pipeline runtime by 67%” beats “improved pipeline”
  • Include GitHub links to real projects
  • Mention data volumes (TB, PB scale)
  • List certifications (AWS, GCP, Databricks, dbt)

📊 Data Engineering Salary

LevelYoEUSA Salary Range
Junior DE0–2$80k–$120k
Mid-level DE2–5$120k–$170k
Senior DE5–8$160k–$220k
Staff DE8–12$200k–$280k
Principal DE12+$260k–$400k+

Source: levels.fyi, LinkedIn Salary, Glassdoor (2024)


🌟 Top Influencers to Follow

LinkedIn

NameProfileFollowers
Zach WilsonEcZachly100,000+
Seattle Data GuySeattleDataGuy50,000+
Andreas KretzAndreas Kretz50,000+
Lior GavishLior Gavish30,000+

Twitter/X

NameHandleFollowers
ByteByteGo@alexxubyte500,000+
Dan Kornas@dankornas66,000+
Zach Wilson@EcZachly30,000+
Seattle Data Guy@SeattleDataGuy10,000+

🤝 Contributing

Contributions are what make the open source community amazing! Any contributions you make are greatly appreciated.

  1. Fork the Project
  2. Create your Feature Branch (git checkout -b feature/AmazingResource)
  3. Commit your Changes (git commit -m 'Add some AmazingResource')
  4. Push to the Branch (git push origin feature/AmazingResource)
  5. Open a Pull Request

See CONTRIBUTING.md for detailed guidelines.


👥 Contributors


📄 License

Distributed under the MIT License. See LICENSE for more information.


⭐ Star History

图片:Star History Chart


Built with ❤️ for the Data Engineering Community

🌐 Website • ⭐ Star this repo • 🐛 Report Bug • 💡 Request Feature

Official distribution

获取与安装

暂未发现可确认的官方软件包地址

当前 README 快照没有出现 npm、PyPI、Crates.io、pub.dev 等官方包页链接。本站不会根据仓库名称猜测下载地址。

本站不托管项目文件;需要安装时,请以项目维护者发布的官方文档为准。

使用前核验

本站保存公开资料用于阅读,不代表安全审计或功能背书。安装前请核对许可证、依赖来源和发布签名,不要直接运行来源不明的二进制文件或高权限脚本。