Senior Software Engineer, Infrastructure

VSCOSan Francisco, CA
Hybrid

About The Position

We’re looking for a Senior Software Engineer, Infrastructure to own the platform that VSCO product and data teams ship on. You’ll join a small infra team that treats AWS, EKS, and GitOps as the default path for new systems, and you’ll spend real time pairing with other teams so - search, Workspace, and data - land in the right Terraform, Helm, and Flux instead of one-off snowflakes. This is a hands-on senior seat. You will design Terraform modules, cut production traffic to EKS, secure traffic with Cloudflare, look for optimizations in our AWS infrastructure, and leave on-call better than you found them. You like production ownership, you can sequence a cutover, and you can teach another team why their service belongs in shared-infra versus an app repo. The day-to-day Design, build, and operate the AWS and EKS platform using Terraform, Helm/Kustomize, and Flux GitOps across infra, shared-infra, and app-owned repos (Workspace, search, data). Own production cutovers end to end: ingress, DNS, workers and crons, telemetry monitors and dashboards, and a rollback plan you would actually use. Partner with product and data teams on the right home for new infra, including namespaces, secrets (SOPS, IRSA), preview environments, and not reinventing namespace or secret wiring that already lives in shared-infra. Build and maintain CI/CD on GitHub Actions, including self-hosted runners, image builds, and deploy pipelines that other teams can copy. That means supporting the build and test pipelines, container images, and runners for our language ecosystems (Go, Python, PHP, Node, Java, Ruby, and similar) so those teams can compile, test, and ship, rather than writing their application code yourself. Run core data stores and caches as platform services: MySQL, Mongo, Valkey/Redis, OpenSearch. Cloudflare and CloudFront: DNS and zone moves, CDN behavior, WAF, custom domains. Observability and autoscaling with monitors, dashboards, external-metrics HPA, with a bias toward fixing the system so the next person doesn’t page. Participate in the infra on-call rotation. Debug Linux, Kubernetes, and network issues, then turn the incident into Terraform, alerts, or docs. Review infra PRs and design docs. Help other engineers make the tradeoff, not just merge the change. Scope of the role Not an application-backend seat that happens to know Kubernetes. The center of gravity is platform: Terraform, EKS, DNS, GitOps, CI, and helping other teams land on it. This is a senior IC seat. You own projects end to end, review PRs, sequence cutovers, and partner with other teams. You are not expected to set org-wide architecture or operate primarily as a tech lead of other infra engineers. You will still open the PR, apply it to dev/preprod/prod in the right order, and watch the cutover.

Requirements

  • 5+ years in infrastructure, platform, or site reliability engineering, including time as a hands-on owner of production systems (not only tickets and runbooks).
  • Strong production experience with Kubernetes/EKS, Linux, and infrastructure-as-code (Terraform). You have shipped and operated this, not only used a chart someone else wrote.
  • AWS in production: EKS, IAM/IRSA, networking (VPC, NAT, DNS), and at least one of RDS/MySQL, Elasticache/Valkey, or OpenSearch.
  • GitOps and CI/CD in anger: Flux and/or Helm, plus GitHub Actions (self-hosted runners a plus).
  • Familiarity with the languages our CI builds (Go, Python, PHP, Node, Java, Ruby, and similar). You do not need to be a product engineer in these languages. You do need to be comfortable supporting the pipelines, images, and runners that compile and test them.
  • Experience leading technical architecture discussions, explaining tradeoffs, and driving a decision inside a small team.
  • Experience troubleshooting complex production issues, including network and DNS, and turning them into durable fixes.
  • Excellent communication: you can teach another team how to consume the platform, and you can write the cutover plan.

Nice To Haves

  • Cloudflare DNS, CDN, and zone migration experience (alongside or instead of CloudFront).
  • OpenSearch, Valkey/Redis, MySQL, Mongo, Kafka.
  • Datadog, Groundcover, Honeycomb (monitors, dashboards, bonus if you understand OTEL).
  • GitHub Actions Runner Controller, preview environments, or similar paved-road CI.
  • Flux, Helm, Kustomize, and a point of view on which repo should own a given piece of infra.
  • Using AI tools like Cursor to help force multiply your work
  • High-scale consumer or creator products.
  • A connection to photography, visual storytelling, or creative communities.
  • BS in Computer Science or related technical discipline, or equivalent practical experience.

Responsibilities

  • Design, build, and operate the AWS and EKS platform using Terraform, Helm/Kustomize, and Flux GitOps across infra, shared-infra, and app-owned repos (Workspace, search, data).
  • Own production cutovers end to end: ingress, DNS, workers and crons, telemetry monitors and dashboards, and a rollback plan you would actually use.
  • Partner with product and data teams on the right home for new infra, including namespaces, secrets (SOPS, IRSA), preview environments, and not reinventing namespace or secret wiring that already lives in shared-infra.
  • Build and maintain CI/CD on GitHub Actions, including self-hosted runners, image builds, and deploy pipelines that other teams can copy.
  • Support the build and test pipelines, container images, and runners for our language ecosystems (Go, Python, PHP, Node, Java, Ruby, and similar) so those teams can compile, test, and ship.
  • Run core data stores and caches as platform services: MySQL, Mongo, Valkey/Redis, OpenSearch.
  • Manage Cloudflare and CloudFront: DNS and zone moves, CDN behavior, WAF, custom domains.
  • Implement observability and autoscaling with monitors, dashboards, external-metrics HPA, with a bias toward fixing the system so the next person doesn’t page.
  • Participate in the infra on-call rotation.
  • Debug Linux, Kubernetes, and network issues, then turn the incident into Terraform, alerts, or docs.
  • Review infra PRs and design docs.
  • Help other engineers make the tradeoff, not just merge the change.

Benefits

  • Competitive salary & equity
  • Medical, dental, and vision insurance for employees and families
  • Flexible Time Off
  • Company-paid parental, medical and caregiver leave
  • Mental health resources
  • Tech reimbursements
  • Flexible time off
  • 401K retirement plan
  • Insurance (medical, dental, vision, life/AD&D, short and long term disability)
  • 11 paid holidays
  • Paid sick time as required by state and local law
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service