Staff Software Engineer, Devops

Slate Auto•Remote WA, WA
•$155,854 - $233,781

About The Position

Slate is looking for a Staff DevOps Engineer to set the technical direction for how we build, run and operate our infrastructure, and to build a large share of it yourself. The team that builds the platform also runs it, carries the pager, and is accountable for its reliability. Our AWS foundation is in place: a governed multi-account environment, private connectivity between the cloud and our sites, and infrastructure as code throughout. Next is a hybrid Kubernetes platform that spans Amazon EKS and self-managed clusters in our own facilities, serving our digital, data, manufacturing and enterprise teams. You will design that platform and help ship it. This is a hands-on role. You'll lead through architecture, code and example rather than through direct reports, and you'll take your turn on call like everyone else on the team. You will report to the leader of the DevOps team and work closely with Digital, Data Engineering, IT, Manufacturing Systems and Security.

Requirements

  • 10+ years in DevOps, site reliability, platform or infrastructure engineering, including several years setting technical direction for production platforms.
  • Deep AWS Expertise: Extensive hands-on experience designing and operating production AWS environments across many accounts, including multi-account governance, identity and access management, VPC networking and hybrid connectivity, DNS, EKS, containers, serverless and managed data services. An AWS Professional or Specialty certification is a plus.
  • Kubernetes in Production, in the Cloud and On-Premises: Proven experience running Kubernetes in production on EKS and on self-managed clusters built on VMs or bare metal, including cluster lifecycle, CNI networking, ingress, storage, RBAC and upgrades. Experience with Rancher or a comparable multi-cluster management platform is a strong plus.
  • Python and Infrastructure as Code: Strong experience with infrastructure as code; AWS CDK is strongly preferred. Our platform code is written in Python, so strong Python skills are required. You can read, understand and make targeted changes to TypeScript, which some of our application teams use.
  • CI/CD and GitOps: Experience building delivery pipelines with GitHub Actions or similar tools, and deploying to Kubernetes with GitOps tooling (Argo CD or Flux), Helm and Kustomize.
  • Reliability Engineering: A track record of establishing SLOs, observability (Prometheus, Grafana, OpenTelemetry, CloudWatch or similar), incident management and on-call practices, and of participating in on-call yourself.
  • Hybrid Networking: A solid grasp of networking across cloud and on-premises environments: routing, VPN, DNS, load balancing, certificates and firewalls, and how each of them fails.
  • Technical Leadership Without Authority: Ability to drive architectural decisions across teams you don't manage, write clear design documents, and bring skeptical stakeholders along.
  • Startup Orientation: Comfort in a fast-moving, resource-constrained environment. You know how to ship, how to cut scope without cutting corners, and how to improve a system while it's running.
  • Bachelor's degree in Computer Science, Engineering or a related field, or equivalent practical experience.

Nice To Haves

  • Exposure to manufacturing, industrial or OT environments and the constraints of running software near the plant floor.
  • Experience running GPU, machine learning or AI and agentic workloads on Kubernetes.
  • Familiarity with enterprise virtualization platforms such as VMware.
  • Policy as code (OPA Gatekeeper or Kyverno), service mesh, or GitHub Enterprise administration.

Responsibilities

  • Architect the Hybrid Kubernetes Platform: Define and build Slate's Kubernetes ecosystem across Amazon EKS and self-managed clusters running on virtualized infrastructure provided by our IT team. Own the reference architecture for cluster provisioning, multi-cluster management (such as Rancher), networking, identity, secrets, storage and upgrades.
  • Set the Direction for Our AWS Foundation: Evolve our multi-account AWS environment, including account governance and guardrails, identity and access management, and the networking and DNS that connect the cloud to our sites. Decide where new workloads live and how they connect.
  • Build the Reliability Practice: Establish our on-call rotation, SLOs and error budgets, alerting standards, incident response and blameless postmortems. Stand up the observability (metrics, logs, traces and synthetic checks) that tells us something is wrong before our users do.
  • Enable the Teams We Serve: Partner with our Digital teams (customer-facing website and backend services), Data Engineering, Manufacturing Systems and enterprise application teams. Give them well-supported paths: CI/CD templates, deployment patterns, reusable infrastructure libraries and self-service environments.
  • Plan for What's Next: Shape how the platform supports manufacturing and plant-floor workloads, data pipelines, and emerging AI and agentic workloads, whether each one runs on-premises or in the cloud.
  • Build In Security and Cost Discipline: Make least-privilege access, policy guardrails, secrets management and software supply chain controls the default. Keep cloud spend visible and deliberate through tagging, right-sizing and commitment planning.
  • Ship and Mentor: Write production infrastructure and code every week. Lead design reviews, mentor engineers on the team and across the organization, and write the documentation and runbooks that let others move without waiting on you.

Benefits

  • medical
  • dental
  • vision
  • life insurance
  • disability insurance
  • vacation
  • 401k
  • equity program
  • discretionary annual incentive program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service