Senior Data Platform Engineer

GeminiMiami, FL
$126,000 - $180,000Hybrid

About The Position

As a Senior Data Platform Engineer, you'll be an SRE for our data infrastructure: responsible for the reliability, automation, and operability of the database and datastore fleet, not just tuning any single engine. You'll work closely with data engineering and product engineering teams to build self-service, automated foundations that let teams provision, scale, and operate their own datastores safely. Your role centers on deep relational database expertise (PostgreSQL, Aurora) - internals, replication, failover, backup/recovery, and query performance - paired with the automation and tooling to run that expertise at scale rather than by hand. You'll extend that same automation-first discipline to the other datastores in our fleet - NoSQL/document, columnar, key-value, and streaming systems - so that scaling, failover, and provisioning are repeatable and largely self-service. You'll drive improvements to observability, incident response, and uptime posture through proactive ownership of the platform, and you'll bring SRE practices (SLOs, error budgets, toil reduction, blameless postmortems) to how we run data infrastructure. This role is ideal for someone who thinks in systems and automation first, thrives on cross-team collaboration, and wants to reduce operational toil across a diverse set of database technologies in a fast-paced, cloud-native environment.

Requirements

  • 5 years of experience in the field.
  • Deep, specialist-level experience managing and scaling relational databases - cloud-native systems like PostgreSQL, Amazon Aurora, or similar - including replication, failover, backup/recovery, and query/engine performance tuning.
  • Demonstrated SRE mindset: experience building automation, self-service tooling, and guardrails that eliminate manual, repetitive database operations rather than performing them by hand.
  • Hands-on experience with at least one non-relational paradigm in production (e.g., NoSQL, columnar, document, key-value), and working knowledge of when to apply each.
  • Familiarity with cloud-based data platforms and services such as AWS RDS, Redshift, EMR, Google BigQuery, or Databricks.
  • Experience in an infrastructure as code environment (Terraform), developing automated solutions to solve support and operational issues.
  • Proficiency writing scripts, CLIs, or services that increase developer productivity and reduce operational toil, in languages like Python, Go, etc.
  • Understanding of CI/CD, observability tooling, SLOs/error budgets, and incident response in production environments.
  • Experience integrating with data pipelines and real-time messaging systems like Kafka or Kinesis.
  • Comfortable participating in on-call rotations and owning uptime and recovery responsibilities across multiple database technologies.
  • Strong communication and collaboration skills; able to work effectively across infrastructure, data, and product teams.

Responsibilities

  • Build Infrastructure as Code (IaC), CLI tools, and CI/CD-driven automation that make database provisioning, scaling, failover, and deployment self-service, consistent, and repeatable across environments - this is the core of the role, not a supporting activity.
  • Serve as the team's depth on relational database systems (e.g., Amazon Aurora, PostgreSQL) - replication topologies, failover, backup/recovery, and query/engine-level performance - ensuring high performance and availability under growing workloads.
  • Extend that operational rigor to the rest of the datastore fleet - document, key-value, and columnar systems - applying the right paradigm to the right workload and building common tooling and guardrails across all of them.
  • Define and track SLOs/error budgets, build proactive monitoring and alerting, implement high-availability architectures, and participate in the on-call rotation to troubleshoot and resolve production issues quickly.
  • Collaborate with data and product engineering teams to integrate with upstream and downstream pipelines - both real-time and batch - via message queues (e.g., Kafka), ETL workflows, and processing frameworks.
  • Identify and resolve performance bottlenecks at both the query and infrastructure levels across engines. Establish alerting, observability, and incident response procedures that reduce MTTR and maintain service health.
  • Continuously identify and automate away repetitive operational work; contribute to shared documentation, incident retrospectives, and platform playbooks to improve team effectiveness and reliability of operations.

Benefits

  • Competitive starting pay
  • A discretionary annual bonus
  • Long-term incentive in the form of a new hire equity grant
  • Comprehensive health plans
  • 401K with company matching
  • Paid Parental Leave
  • Flexible time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service