Getting started with dbt
| Source: Towards Data Science
Tags: dbt, data engineering, SQL, data transformation, DuckDB, analytics
A practical walkthrough of dbt Core — the free, Apache 2.0-licensed SQL transformation tool that adds software engineering discipline (testing, documentation, lineage, CI/CD) to analytics pipelines running on Snowflake, BigQuery, Redshift, and DuckDB.
Details
This Towards Data Science article provides a hands-on introduction to dbt (data build tool), written by a contract data engineer motivated by seeing it appear repeatedly in job listings. The piece covers what dbt is, why it exists, and how to use it at a basic level. dbt Core is free and Apache 2.0 licensed, running locally via CLI without a dbt account. The managed paid version, dbt Platform, is also available for teams wanting scheduled jobs and CI/CD pipelines. dbt transforms data already in a warehouse by executing user-supplied SQL to create tables or views, but adds capabilities that raw SQL scripts lack: automated testing, dataset documentation, lineage tracking, macro reuse, and environment management across dev/test/prod. The article focuses on models and sources — the building blocks of dbt — and shows how testing and documentation are layered on top. It uses DuckDB for local examples, making it easy to follow without a cloud warehouse. The piece is an entry point, not an exhaustive reference; readers wanting advanced features like custom materializations or cross-project dependencies will need additional resources. It's most useful for data engineers, analysts, or ML practitioners encountering dbt for the first time.