About The Position

Sunset operates customer-facing SaaS products, connector and ingestion services, asynchronous workers, high-volume data pipelines, model-backed systems, review tools, and customer-delivery paths. These workloads have different shapes, but they need a coherent foundation for infrastructure, delivery, observability, recovery, access, and cost. You will build and operate the shared platform that lets our product, data, and AI teams ship reliable, secure, observable, and cost-aware systems without manual infrastructure work or operational risk growing linearly. You will write software and infrastructure, improve real engineering workflows, lead through incidents, and create paved roads teams can use without waiting on you. This is not a deployment-operator or internal-IT role. Product, data, and ML teams remain responsible for the systems they build. You will give them the runtime, delivery, visibility, recovery, and operating patterns to own those systems well. You will partner closely with our Security Lead, but you will not be expected to run the entire security or compliance program.

Requirements

  • You have personally owned production cloud infrastructure and delivery or reliability systems across multiple services, including an asynchronous, batch-data, or model-backed workload
  • You are a strong software engineer who is comfortable changing application, platform, and infrastructure code and operating the result in production
  • You can reason from user impact through dependencies, state, telemetry, incident response, recovery, and durable remediation
  • You have built paved roads other engineers adopted because they made real work easier, not because a platform team required them
  • You understand both long-running services and high-volume or scheduled workloads and know where their reliability models should differ
  • You can make pragmatic tradeoffs among delivery speed, least privilege, isolation, recovery, developer experience, and unit cost
  • You are effective in an early-stage environment where the first step is often to establish ownership and a trustworthy baseline
  • You can lead calmly through ambiguous incidents, communicate clearly, and leave the system and operating model stronger afterward
  • You use modern AI engineering tools fluently and verify generated infrastructure, queries, code, and operational conclusions before they affect production

Nice To Haves

  • Experience as an early platform or SRE hire at a fast-growing company
  • Experience with AWS, Terraform, container runtimes, workflow orchestration, and observability systems
  • Experience with high-volume data processing, model serving, evaluation jobs, GPU workloads, or machine-learning platforms
  • Experience improving developer environments, preview systems, CI/CD, progressive delivery, or internal developer platforms
  • Experience with replayable pipelines, backup and restore, disaster recovery, capacity planning, or cloud-cost allocation
  • Experience implementing technical controls and automated evidence for SOC 2 or enterprise customer requirements

Responsibilities

  • Establish Sunset's current platform, workload, reliability, ownership, toil, recovery, cost, and technical-control baseline
  • Build reusable infrastructure-as-code modules, runtime templates, deployment workflows, environment contracts, and operational tooling
  • Create supported paths for customer-facing services, asynchronous and batch jobs, data pipelines, and model-backed workloads
  • Improve deploy safety, workload visibility, backup and recovery, incident response, replay, rollback, and durable remediation
  • Work with engineering teams to define useful service and pipeline objectives, ownership, escalation, and recovery paths
  • Build self-service for common infrastructure, environment, access, deploy, debugging, and recovery work without becoming a central approval queue
  • Make cloud and vendor cost understandable by service and workload and improve efficiency within explicit reliability and security bounds
  • Partner with Security on cloud identity, secrets, isolation, audit logging, vulnerability response, incident readiness, and automated control evidence
  • Support employees and contractors through bounded access, safe environments, release controls, documentation, and timely removal of authority
  • Use AI tools deeply for platform engineering and operations while verifying generated code, plans, queries, state changes, and incident conclusions
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service