The pursuit of spending $0 Building a zero-cost cloud data platform By Sierra McDowell

  1. Why? 2. Tooling Choices Agenda
  2. Building the Stack 4. Changes and Future Plans
  3. The Result

3-ish year journey 2023/ 2024 2025 2026

Background and Context

Sell jewellery (B2C) Background and Context People who like jewellery Business process Turn into jewellery (in- house) Miner Facet rocks Buy rocks Sell faceted rocks (B2C) People who like rocks Jewellers People who like rocks with preexisting rocks Sell faceted rocks (B2B)

Background and Context B2C B2B Products Services Physical products - cut gemstones Gemstone cutting as a service (GaS) Points of Sale Invoicing (Square & PayPal) Website (Squarespace) Multiple sources of sales data In Person Sales

Background and Context 20 Google Sheets Running the entire business in 2023

Sell jewellery (B2C) Background and Context Which stones should I cut and source more of? Miner Buy rocks Facet rocks People who like jewellery Turn into jewellery (in- house) Sell faceted rocks (B2C) People who like rocks Jewellers People Sell faceted who like rocks (B2B) rocks with preexisting Which stones are selling and which rocks are left in stock?

Considerations What I normally work with What I was dealing with Big money $0 Tech budget Tech budget Lots of people 1 person Engineering + business teams + stakeholders Business size of 1 person, non-technical Big data Small data Big data, solutions need to scale Small data, and likely won’t ever be big data

The Goal The Aim Current State Sierra running a script on her laptop Overkill and $$$

The Goal 1 2 3 Redundancy Better decisions Built to extend Historical record of gemstone data, reducing risk of the business running entirely off spreadsheets Serve analytic use cases so decisions are no longer made off vibes Designed so future use cases can be served from it without total rework

Tooling Choices

Free or Open Source: No limited-time free trials Tech Tooling Guiding Principles No Vendor Lock in: Must store data in a way that makes re-platforming simple if required. No ClickOps Against my Will: Everything able to be defined as code. Low Maintenance: Runs the background Learning Experience: Overkill is ok if it’s for learning.

Required Components 1 2 3 Ingest Store Transform Something to ingest the data Something to store the data Something to transform the data from raw data to curated data products 4 5 6 Orchestrate Visualise IaC backend Something to orchestrate the data pipelines Something to visualise the data For anything in Terraform (IaC) — something to store the Terraform backend

What Didn’t Make the Cut

What Didn’t Make the Cut No Free Tier Free Trial Only *At the time of design AWS S3

What Didn’t Make the Cut Update for 2026 – now has a free tier No Free Tier Free Trial Only *At the time of design AWS S3

What Didn’t Make the Cut Tooling Type How free? Freeness Details (at Design stage) Freeness details (2026) Why it wasn’t chosen BigQuery Storing Data Free Tier 1 TB of querying per month 10 GB of storage No change A strong contender Databricks community version Storing Data Free Separate product to Databricks, with limited features No longer exists Very limited features PostgresSQL Storing Data Open source Entirely free - just have to run it somewhere Yes Didn’t want something running on my laptop Looker Studio Data Visualisation Free Tier Fully free just doesn’t have enterprise features Renamed to Data Studio but same Freeness I would rather die

What was Chosen

Fivetran Ingestion Ingestion as a Service Manged Service Declarative Just define the connector type, and destination, sync frequency Many Connector Types Databases, apps etc

Ingestion Tooling How free? Freeness Details (at Design stage) All the features of Standard Plan Up to 500,000 monthly active rows for connections Freeness details (2026) • Fivetran Forever Free Tier • Same plus some additional features (not relevant to me)

MotherDuck Storage duckDB in the Cloud Manged Service Serverless Ducklings Read/Write tables to Parquet S3, GCS, Azure blob

Storage Tooling How free? MotherDuck Forever Free Tier Freeness Details (at Design stage) ● ● ● 10 GB of Storage 10 Compute Unit Hours Per Month 5 Collaborators Freeness details (2026) ● ● ● ● 10 GB of Storage 10 Pulse Compute Unit Hours Per Month Up to 3 internal active users and 2 service accounts Role-based access control (preset roles)

Storage Tooling How free? MotherDuck Forever Free Tier Freeness Details (at Design stage) ● ● ● 10 GB of Storage 10 Compute Unit Hours Per Month 5 Collaborators Freeness details (2026) ● ● ● ● Google Cloud Storage Forever Free Tier 5 GB-months of regional storage (US regions only) per month 10 GB of Storage 10 Pulse Compute Unit Hours Per Month Up to 3 internal active users and 2 service accounts Role-based access control (preset roles) No change

dbt Transformation SQL Transformation Compiles SQL and executes it against data warehouse. DAG Builds a dependency graph between models Command Line Execute dbt transformations with the command line

Transformation Tooling How free? dbt core Open Source Freeness Details (at Design stage) Entirely free - just have to run it somewhere Freeness details (2026) No change

Streamlit Visualisation Streamlit Open source data visualization framework written in Python Streamlit Community Cloud SaS offering for hosting streamlit apps

Visualisation Tooling Streamlit Community Cloud How free? Freeness Details (at Design stage) ● Free Version ● Deploy 1 private app (private Github repo) Deploy unlimited public apps Freeness details (2026) No change

DevOps Tooling Github Actions HCL Terraform Component How free? Orchestration Forever Free Tier Freeness Details (at Design stage) ● ● ● IaC Backend Forever Free Tier ● Freeness details (2026) 500 MB Storage 2,000 Minutes (per month) No change 500 resources Per Month Remote state storage encrypted at rest No change

Architecture Data Ingestion

Data Ingestion Motherduck Supported Fivetran includes Motherduck as a destination Connectors Fivetran includes connectors for Google Sheets, Squarespace, Square and Paypal

Challenges and Hiccups

Challenges and Hiccups

Data Ingestion Original Ingestion Plan Long Live Google Sheets Each Google Sheet as a Connector Fivetran connectors deployed with Terraform Terraform Backend State stored in HCP Terraform

Data Ingestion

Trade-offs Trade-off Reasoning Using Google sheets as a data source Didn’t want to make operational running of the business overly complicated. Fivetran MotherDuck destination is Partner Built and (was) Private preview This is a low-stakes project, ok with rolling the dice on this.

Architecture Data transformation and storage

Bronze Layer Staging layer Intermediate layer converting all columns to string History layer Track changes using incremental mode ‘SCD2 lite’ Idempotent from Bronze Designed so everything downstream from bronze is idempotent

Bronze Layer

Bronze Layer

Bronze Layer

Bronze Layer

Silver Layer Cleaned Historical Layer Union bronze tables into one table Data type casting and cleaning Calculate date of sale using history Cleaned Current layer Intermediate layer getting current records from history Curated Layer Split products and services into separate tables

Gold Layer

Trade-offs Trade-off Reasoning MotherDuck is currently hosted in Amazon AWS region us-east-1 *This was the only region available in 2023 but no customer data and is still pretty fast. Australia (apsoutheast-2) is now available in 2026 GCS free tier is restricted to us-east1, uswest1, and us-central1 No customer data but I probably should have chosen useast1 instead of us-west1 DuckDB (and MotherDuck) do have concurrency restrictions Might cause problems in a team but for this purpose can get around it by coordinating timing of jobs, no problems so far. Lack of fine grain access control, SSO and private networking in Motherduck Free Tier These features are important for an enterprise deployment but not a dealbreaker for this project.

Architecture Orchestration

Orchestration - Ingestion

Orchestration - Transformation

What About Dev? Component Prod Dev Database MotherDuck duckdb Ingestion Fivetran Python script Transformation dbt core dbt core Orchestration Github Actions Bash commands Infra Terraform (HCP Terraform backend) Terraform (local backend) Analytics Streamlit deployed with Streamlit Community Cloud Streamlit run locally

What About Dev?

Data Quality There were some data quality issues.. Tasmania has declared independence? These are meant to be IDs

Data Quality Framework Design Tests or Tests +

Data Quality Framework Design

Data Quality Framework Design

Data Quality Framework Design

Dashboard

Changes in the Tech World

2025 – Databricks Free Edition is Released

2025 – Databricks Free Edition is Released

2025 – Databricks Free Edition is Released Google Analytics Data Core Operational Data + KPIs Gemstone Design Files and Images Product Sales and Inventory Genie Spaces Professional Services Potential parallel use cases

2026 – Initial Use Cases in Databricks Explored

Sell jewellery (B2C) What Next? Turn into jewellery (in- house) Miner Facet rocks Buy rocks People who like jewellery TODO Sell faceted rocks (B2C) People who like rocks TODO Jewellers People who like rocks with preexisting rocks Sell faceted rocks (B2B) Covered

Have We Stayed Free? $0 Total Spend so Far

Have We Stayed Incident Free? 1 Fivetran Metadata Connector Bug Metadata sync issue only, didn’t break prod Fivetran Connector Incident 1 dbt Pipeline Incident I forgot TRY_CAST Dodgy data came through and broke everything

Summary 100% Data Driven Gem Cutting 3 yrs 1 Platform has been Live Times Prod Broke 117 $0 Tables in Prod Total Spend