Data Platform Engineer

Hubble NetworkSan Francisco, CA
$153,000 - $250,000Onsite

About The Position

This role will be based in San Francisco, CA. This is the founding data hire. You will build Hubble's data platform from the warehouse layer up, make decisions on architecture, determine processes for managing schema and data models, and work with many stakeholders. We know what we want: a layered lakehouse with Landing, Staging, Warehouse, Mart stages. Each layer has an owner, a purpose, and a quality bar. You’ll work with the Infrastructure team to manage the pipes: Fivetran, Redshift, orchestration, monitoring. You own everything downstream of raw: the transformations, the models, the metric definitions, the quality gates, and the performance bar. The guiding principle is data democratization. Every person at Hubble should be able to answer their own questions. Success is not how many queries you run for other people, it's how few they need you to run. You'll be a peer to Platform Engineering and to the team building our internal operations platform. You'll be an embedded consultant to Finance, Product, and Customer Success. You’ll think about how to present a strong story from data. We expect you to use AI agents as part of how you work, and data work rewards this more than most: schema exploration, model scaffolding, test generation, and documentation are all places where agentic tooling earns its keep if you review the output critically.

Requirements

  • 7+ years building production data systems, with 2+ years at senior or staff scope owning architecture rather than executing someone else's.
  • Deep SQL and dimensional modeling: star schemas, slowly changing dimensions, grain discipline. You can explain fact tables and grains.
  • dbt (or similar) in production: model organization, macros, incremental strategies, tests as contracts, and CI that blocks bad merges.
  • Cloud data warehouse depth, ideally Redshift (or similar): distribution and sort keys, vacuum and analyze behavior, workload management, and the ability to take a slow query apart and make it fast.
  • Python or Scala for the specialized pipelines: custom connectors against proprietary APIs, backfills, and reconciliation jobs.
  • Event and product analytics instrumentation: you've defined a tracking plan, argued about event taxonomy, and built funnel tables that survive contact with a changing product.
  • Data quality as engineering: anomaly detection, freshness monitoring, and incident response with real escalation paths. You've been paged for a broken pipeline and you've fixed the class of bug, not the instance.
  • Fluency with AI agents in your own workflow: you already use agentic coding tools to ship real work, and you have a point of view on where they help and where they quietly introduce errors that only show up three models downstream.

Nice To Haves

  • Interest or experience in IoT, satellite systems, telemetry, or geospatial data.
  • Billing, usage-based pricing, or revenue reconciliation and recognition workflows.
  • SOC 2 or similar compliance work: classification frameworks, access reviews, audit evidence.
  • Open source contributions or public technical writing.

Responsibilities

  • Own the transformation layer end-to-end: Defined models across Staging, Warehouse, and Mart. Star-schema design, SCD2, and data contracts enforced by tests at every layer.
  • Build the data layer our internal tools run on: mart tables designed for specific operational workflows, sitting behind a generated internal metrics API. You own the mart and the API's performance characteristics; Engineering owns the framework and the frontend. Hold the line at p95 under 5 seconds on mart queries and under 500ms on metric endpoints.
  • Make freshness a commitment, not an aspiration: core data and critical operational metrics streamed in real time with Apache Iceberg, with monitoring that catches breakage before a stakeholder does.
  • Instrument the funnels: Design event collection across our different product surfaces, unified into a customer 360 spanning authentication, usage, and billing. Acquisition, activation, engagement, retention, revenue.
  • Data in space: We collect large amounts of telemetry from our satellites which is critical to mission operations - help our Mission Operations group land this telemetry in dashboards and monitoring systems so they can take action on data immediately.
  • Automate revenue reconciliation: Daily comparison of Stripe Platform billing against Hubble usage, with discrepancies flagged before month-end and drill-down to the device and packet level. Turn a multi-day manual close into a few hours.
  • Define metrics once: Active devices, packet volume, MRR, contract utilization. Every number has a canonical definition, an owner, and visible lineage from source to dashboard. “Which revenue number is correct?” should stop being a question anyone asks.
  • Implement classification in the pipeline: Security defines the tiers and the guardrails. You tag datasets by sensitivity, enforce retention, masking, and anonymization, and keep lineage and access logs auditable. Hubble does not store PII, and the pipeline is where that commitment either holds or quietly fails.
  • Build for self-service: Semantic layers that let business users work with customers and revenue rather than join keys, curated datasets, a data catalog, and documentation good enough that people stop asking you.
  • Choose boring technology: Fivetran, dbt, Redshift, Metabase are all options - we lean towards buy over build. Bring a strong opinion about which is which, and be able to defend it.

Benefits

  • Health, Dental, Vision, HSA/FSA options
  • Unlimited PTO
  • Commuter benefits (if working from HQ)
  • Learning & Development allowance
  • Health & Wellness stipend
  • Sabbatical program
  • Work on state-of-the-art satellite systems
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service