Data Platform Architecture
Lakehouse and warehouse design from zero: BigQuery, Databricks, Snowflake, dbt. Medallion layering, schema evolution, data modeling, and the quality standards that keep a platform trustworthy as the team grows.
Layland Consulting, LLC
Over twenty years building the realtime pipelines, infrastructure, schemas, and governance that AI runs on — I’ve scaled it to 1M events a second, terabytes a day, under HIPAA, GDPR, and CPRA.
I'm Steve Layland, a hands-on data leader with roots in Silicon Valley. I've spent more than two decades across streaming, healthcare, and adtech — building the realtime pipelines and architecture that process 1M requests per second, writing terabytes daily into a petabyte-scale data lake.
Most recently I've been Principal Engineer on data platforms at Tubi, where I architected and scaled the realtime ingestion platform, grew the data org to 25 engineers across four teams, and served as the primary engineer accountable for compliance across GDPR, CCPA, and Executive Order 14117. Before that: engineering leadership at Nuna in a HIPAA-regulated environment, platform work at Metric Insights, and data infrastructure at Linden Lab and Wolfram Research.
Through Layland Consulting I take on a small number of engagements at a time. Current work is with Birches Health, a telehealth startup operating under HIPAA, where I designed and built their data platform on GCP, including a bespoke Fivetran alternative that ingests data from multiple third-party service providers into the data lake hourly.
I write code daily make AI write code daily, mostly in Python, Rust & Scala,
and I enjoy building the data foundations high-growth startups run on.
Engagements usually start with one of these and grow into the others.
Lakehouse and warehouse design from zero: BigQuery, Databricks, Snowflake, dbt. Medallion layering, schema evolution, data modeling, and the quality standards that keep a platform trustworthy as the team grows.
High-throughput ingestion on Kafka, Kinesis, Flink, and Spark. Event schema design, exactly-once semantics, and the operational practice — monitoring, cost control, incident response — that keeps streams healthy at peak.
HIPAA, GDPR, CCPA/CPRA, COPPA, VPPA, and Executive Order 14117. Privacy-safe identity spines, inline de-identification, surrogate IDs, deletion and retention workflows, and audit-ready lineage.
Self-service analytics that actually answers questions: semantic layer configuration, MCP servers, and tool design that let agents and analysts query the warehouse without inventing their own definitions of revenue.
AWS, GCP, and Azure; Kubernetes, Cloud Run, Terraform, Bazel, and CI/CD. Remote build caching and executors, infrastructure-as-code, and cost structures that hold up as volume grows.
Feature stores, model training, versioning, and serving. Realtime features and the pipelines behind them, plus the ML ops scaffolding that gets models from a notebook into production reliably.
First- and last-touch attribution, MMP integrations (Adjust, Kochava, AppsFlyer), and CRM integrations (Braze, HubSpot, Kustomer) — wired through a privacy layer so coverage goes up without your exposure going up with it.
Experiment and cohort analysis, and the data plumbing behind it — integrating Statsig, PostHog, or LaunchDarkly with your warehouse so results are reproducible rather than trapped in a vendor dashboard.
Let’s fix your data problems.
Reach out with a brief overview of what you’re working on.
consulting@layland.xyz