The course repo on 2026-10-08 lists the next live cohort as January 2027; all materials are open for self-paced study at any time. source
A free, cohort-based course that walks you through a modern batch and streaming pipeline on Google Cloud and ends with a peer-reviewed capstone. It suits self-directed learners with some coding and SQL who want a portfolio project rather than a credential or placement help. Check that you can commit 5 to 15 hours a week from January to April, because the certificate only exists inside the live cohort.
The facts
| Category | Data engineering and analytics |
|---|---|
| Website | github.com/DataTalksClub/data-engineering-zoomcamp |
| Based | Remote |
| Format | online-live, online-self-paced |
| Programs | Live cohort (January start, homework deadlines, peer-reviewed capstone, certificate), Self-paced (no certificate) |
| Length | 9 weeks (DataTalks.Club describes a fixed 9-week module schedule plus a capstone; the FAQ says the cohort runs roughly January to April at 5 to 15 hours per week.) verified |
| Tuition | $0 Free. Modules use Google Cloud; the FAQ says the $300 GCP free-trial credit covers the course if you tear down resources, and a credit card is required to open the trial. verified |
| Financing | free No tuition, no guarantee, no refunds to consider. Certificate requires completing and peer-reviewing capstone projects during a live cohort. |
| Admission | open Prerequisites per the repo: basic coding experience, familiarity with SQL; Python helpful but not required. |
| Stack | Docker, Terraform, Google Cloud, PostgreSQL, Kestra, dlt, BigQuery, dbt, Bruin, Apache Spark, Kafka, Python, SQL |
Outcomes
This program does not publish job-placement or salary outcomes. Treat any number you hear from admissions as unverifiable.
Interview-prep depth
0 / 5
The module list (infrastructure, orchestration, warehouse, analytics engineering, batch, streaming, project) contains no interview content.
Whatever the program covers, the interview itself is free to prepare for here: coding patterns, system design, practice prompts, company guides.
What a careful reader should know
Run by DataTalks.Club with Alexey Grigorev as primary author and community contributors. Reddit threads in r/dataengineering describe it as hard for true beginners (environment setup, Docker, homework that assumes some background) while still recommending it because it is free and ends with a real project. Tool choices in several modules (Kestra, dlt, Bruin) change year to year.
Third-party signals: Reddit sentiment mixed. Graduation-time star ratings are collected by the schools; weigh them accordingly.
Sources (4)
This directory takes no lead fees or placement payments from any program. Where a program runs an affiliate scheme we say so on its page and mark the link. Facts carry a badge: verified means we read it on a primary source on the date shown, reported means a third party said it, unknown means nobody we trust has published it.
Useful next steps:
