A presentation at DDD Perth in in Perth WA, Australia by Sierra McDowell

The pursuit of spending $0 Building a zero-cost cloud data platform By Sierra McDowell
3-ish year journey 2023/ 2024 2025 2026
Background and Context
Sell jewellery (B2C) Background and Context People who like jewellery Business process Turn into jewellery (in- house) Miner Facet rocks Buy rocks Sell faceted rocks (B2C) People who like rocks Jewellers People who like rocks with preexisting rocks Sell faceted rocks (B2B)
Background and Context B2C B2B Products Services Physical products - cut gemstones Gemstone cutting as a service (GaS) Points of Sale Invoicing (Square & PayPal) Website (Squarespace) Multiple sources of sales data In Person Sales
Background and Context 20 Google Sheets Running the entire business in 2023
Sell jewellery (B2C) Background and Context Which stones should I cut and source more of? Miner Buy rocks Facet rocks People who like jewellery Turn into jewellery (in- house) Sell faceted rocks (B2C) People who like rocks Jewellers People Sell faceted who like rocks (B2B) rocks with preexisting Which stones are selling and which rocks are left in stock?
Considerations What I normally work with What I was dealing with Big money $0 Tech budget Tech budget Lots of people 1 person Engineering + business teams + stakeholders Business size of 1 person, non-technical Big data Small data Big data, solutions need to scale Small data, and likely won’t ever be big data
The Goal The Aim Current State Sierra running a script on her laptop Overkill and $$$
The Goal 1 2 3 Redundancy Better decisions Built to extend Historical record of gemstone data, reducing risk of the business running entirely off spreadsheets Serve analytic use cases so decisions are no longer made off vibes Designed so future use cases can be served from it without total rework
Tooling Choices
Free or Open Source: No limited-time free trials Tech Tooling Guiding Principles No Vendor Lock in: Must store data in a way that makes re-platforming simple if required. No ClickOps Against my Will: Everything able to be defined as code. Low Maintenance: Runs the background Learning Experience: Overkill is ok if it’s for learning.
Required Components 1 2 3 Ingest Store Transform Something to ingest the data Something to store the data Something to transform the data from raw data to curated data products 4 5 6 Orchestrate Visualise IaC backend Something to orchestrate the data pipelines Something to visualise the data For anything in Terraform (IaC) — something to store the Terraform backend
What Didn’t Make the Cut
What Didn’t Make the Cut No Free Tier Free Trial Only *At the time of design AWS S3
What Didn’t Make the Cut Update for 2026 – now has a free tier No Free Tier Free Trial Only *At the time of design AWS S3
What Didn’t Make the Cut Tooling Type How free? Freeness Details (at Design stage) Freeness details (2026) Why it wasn’t chosen BigQuery Storing Data Free Tier 1 TB of querying per month 10 GB of storage No change A strong contender Databricks community version Storing Data Free Separate product to Databricks, with limited features No longer exists Very limited features PostgresSQL Storing Data Open source Entirely free - just have to run it somewhere Yes Didn’t want something running on my laptop Looker Studio Data Visualisation Free Tier Fully free just doesn’t have enterprise features Renamed to Data Studio but same Freeness I would rather die
What was Chosen
Fivetran Ingestion Ingestion as a Service Manged Service Declarative Just define the connector type, and destination, sync frequency Many Connector Types Databases, apps etc
Ingestion Tooling How free? Freeness Details (at Design stage) All the features of Standard Plan Up to 500,000 monthly active rows for connections Freeness details (2026) • Fivetran Forever Free Tier • Same plus some additional features (not relevant to me)
MotherDuck Storage duckDB in the Cloud Manged Service Serverless Ducklings Read/Write tables to Parquet S3, GCS, Azure blob
Storage Tooling How free? MotherDuck Forever Free Tier Freeness Details (at Design stage) ● ● ● 10 GB of Storage 10 Compute Unit Hours Per Month 5 Collaborators Freeness details (2026) ● ● ● ● 10 GB of Storage 10 Pulse Compute Unit Hours Per Month Up to 3 internal active users and 2 service accounts Role-based access control (preset roles)
Storage Tooling How free? MotherDuck Forever Free Tier Freeness Details (at Design stage) ● ● ● 10 GB of Storage 10 Compute Unit Hours Per Month 5 Collaborators Freeness details (2026) ● ● ● ● Google Cloud Storage Forever Free Tier 5 GB-months of regional storage (US regions only) per month 10 GB of Storage 10 Pulse Compute Unit Hours Per Month Up to 3 internal active users and 2 service accounts Role-based access control (preset roles) No change
dbt Transformation SQL Transformation Compiles SQL and executes it against data warehouse. DAG Builds a dependency graph between models Command Line Execute dbt transformations with the command line
Transformation Tooling How free? dbt core Open Source Freeness Details (at Design stage) Entirely free - just have to run it somewhere Freeness details (2026) No change
Streamlit Visualisation Streamlit Open source data visualization framework written in Python Streamlit Community Cloud SaS offering for hosting streamlit apps
Visualisation Tooling Streamlit Community Cloud How free? Freeness Details (at Design stage) ● Free Version ● Deploy 1 private app (private Github repo) Deploy unlimited public apps Freeness details (2026) No change
DevOps Tooling Github Actions HCL Terraform Component How free? Orchestration Forever Free Tier Freeness Details (at Design stage) ● ● ● IaC Backend Forever Free Tier ● Freeness details (2026) 500 MB Storage 2,000 Minutes (per month) No change 500 resources Per Month Remote state storage encrypted at rest No change
Architecture Data Ingestion
Data Ingestion Motherduck Supported Fivetran includes Motherduck as a destination Connectors Fivetran includes connectors for Google Sheets, Squarespace, Square and Paypal
Challenges and Hiccups
Challenges and Hiccups
Data Ingestion Original Ingestion Plan Long Live Google Sheets Each Google Sheet as a Connector Fivetran connectors deployed with Terraform Terraform Backend State stored in HCP Terraform
Data Ingestion
Trade-offs Trade-off Reasoning Using Google sheets as a data source Didn’t want to make operational running of the business overly complicated. Fivetran MotherDuck destination is Partner Built and (was) Private preview This is a low-stakes project, ok with rolling the dice on this.
Architecture Data transformation and storage
Bronze Layer Staging layer Intermediate layer converting all columns to string History layer Track changes using incremental mode ‘SCD2 lite’ Idempotent from Bronze Designed so everything downstream from bronze is idempotent
Bronze Layer
Bronze Layer
Bronze Layer
Bronze Layer
Silver Layer Cleaned Historical Layer Union bronze tables into one table Data type casting and cleaning Calculate date of sale using history Cleaned Current layer Intermediate layer getting current records from history Curated Layer Split products and services into separate tables
Gold Layer
Trade-offs Trade-off Reasoning MotherDuck is currently hosted in Amazon AWS region us-east-1 *This was the only region available in 2023 but no customer data and is still pretty fast. Australia (apsoutheast-2) is now available in 2026 GCS free tier is restricted to us-east1, uswest1, and us-central1 No customer data but I probably should have chosen useast1 instead of us-west1 DuckDB (and MotherDuck) do have concurrency restrictions Might cause problems in a team but for this purpose can get around it by coordinating timing of jobs, no problems so far. Lack of fine grain access control, SSO and private networking in Motherduck Free Tier These features are important for an enterprise deployment but not a dealbreaker for this project.
Architecture Orchestration
Orchestration - Ingestion
Orchestration - Transformation
What About Dev? Component Prod Dev Database MotherDuck duckdb Ingestion Fivetran Python script Transformation dbt core dbt core Orchestration Github Actions Bash commands Infra Terraform (HCP Terraform backend) Terraform (local backend) Analytics Streamlit deployed with Streamlit Community Cloud Streamlit run locally
What About Dev?
Data Quality There were some data quality issues.. Tasmania has declared independence? These are meant to be IDs
Data Quality Framework Design Tests or Tests +
Data Quality Framework Design
Data Quality Framework Design
Data Quality Framework Design
Dashboard
Changes in the Tech World
2025 – Databricks Free Edition is Released
2025 – Databricks Free Edition is Released
2025 – Databricks Free Edition is Released Google Analytics Data Core Operational Data + KPIs Gemstone Design Files and Images Product Sales and Inventory Genie Spaces Professional Services Potential parallel use cases
2026 – Initial Use Cases in Databricks Explored
Sell jewellery (B2C) What Next? Turn into jewellery (in- house) Miner Facet rocks Buy rocks People who like jewellery TODO Sell faceted rocks (B2C) People who like rocks TODO Jewellers People who like rocks with preexisting rocks Sell faceted rocks (B2B) Covered
Have We Stayed Free? $0 Total Spend so Far
Have We Stayed Incident Free? 1 Fivetran Metadata Connector Bug Metadata sync issue only, didn’t break prod Fivetran Connector Incident 1 dbt Pipeline Incident I forgot TRY_CAST Dodgy data came through and broke everything
Summary 100% Data Driven Gem Cutting 3 yrs 1 Platform has been Live Times Prod Broke 117 $0 Tables in Prod Total Spend
This talk walks through an attempt to build a modern cloud data platform while spending absolutely nothing. As a personal challenge and passion project, I pulled together a stack using MotherDuck, dbt Core, GitHub Actions, and other forever-free or open-source tools (no free trials allowed) I’ll share the architecture, how I approached the build, and the lessons I learnt along the way, including what worked, what didn’t, and the trade-offs that came with keeping costs at zero.