VP Engineering - Infrastructure & SRE

Sezzle
$400,000 - $600,000Remote

About The Position

Sezzle is seeking an exceptional VP Engineering - Infrastructure & SRE to own the strategy, reliability, security posture, and evolution of the platform that powers Sezzle. This role involves scaling a high-growth fintech while raising operational rigor, resilience, and compliance to the standards of demanding financial institutions. It is both a leadership and technical role, requiring hands-on engagement with cloud infrastructure, Kubernetes, databases, networking, observability, and site reliability engineering. The position involves setting technical vision, managing budgets, and ensuring platform availability, disaster recovery, and business continuity. The VP will also lead the organization into an AI-boosted SRE era, driving the adoption of AI agents and tooling for operational efficiency. This role reports to senior engineering leadership and collaborates with Security, Compliance, Engineering, and Finance teams.

Requirements

  • 15+ years of combined experience across infrastructure, platform, site reliability, software development, or related engineering disciplines, with substantial depth in infrastructure.
  • 5+ years of experience leading engineering teams.
  • Deep, hands-on expertise with AWS, including designing and operating production architectures at scale.
  • Deep, hands-on expertise with Kubernetes in production, including cluster lifecycle management and workload architecture.
  • Deep expertise with relational databases at scale, specifically RDS/Aurora (MySQL and/or Postgres), including high availability, replication, failover, performance tuning, and backup/recovery.
  • Proven ownership of disaster recovery and business continuity for a production platform, including defining targets and running failover tests.
  • Demonstrated AI-forward leadership, actively using AI tooling in engineering or operations and leading team adoption of AI-assisted workflows.
  • Track record of operating a 24/7, high-availability platform with mature incident command and postmortem practices.
  • Willingness to manage and participate in an on-call rotation and demonstrated ability to lead recovery from a full production outage.
  • Credible hands-on engineer, comfortable in a terminal, reading dashboards, and reviewing designs.
  • Experience owning significant cloud budgets and driving cost efficiency.
  • Strong grounding in infrastructure-as-code (Terraform or equivalent) and modern CI/CD practices.
  • Demonstrated ability to hire, develop, and retain strong infrastructure and SRE talent.
  • Bachelor's degree in Computer Science or a similar technical field.

Nice To Haves

  • Direct experience supporting PCI-DSS and SOC 2 programs from the infrastructure side.
  • Experience in fintech, payments, or banking, especially in environments with heightened regulatory expectations.
  • Experience deploying AIOps or LLM-based tooling in production operations.
  • Experience with multi-region and active-active architectures, chaos engineering, and formal operational resilience programs.
  • Proficiency with modern observability stacks (Prometheus, Grafana, Loki, Tempo, or commercial equivalents).
  • Familiarity with service mesh, zero-trust networking, secrets management, and workload identity patterns.
  • Experience with CI/CD pipelines, progressive delivery, and platform engineering / internal developer platform approaches.
  • Experience presenting to boards, auditors, or examiners.

Responsibilities

  • Own the infrastructure vision, strategy, and multi-year roadmap, focusing on scaling the platform while strengthening resilience, controls, and audit-readiness.
  • Lead, grow, and mentor the infrastructure, platform, and SRE organization, including hiring, career development, and fostering a culture of operational excellence.
  • Own end-to-end reliability by defining and enforcing SLOs, maturing incident management, and ensuring platform availability.
  • Manage and participate in the on-call rotation, serving as senior incident commander for high-severity events.
  • Direct AWS strategy, including account architecture, IAM, network design, multi-AZ/multi-region posture, service selection, and cost management (FinOps).
  • Own the Kubernetes platform as a product, including cluster architecture, upgrade strategy, workload isolation, and developer experience.
  • Own the database tier, focusing on Aurora RDS (MySQL and Postgres) availability, performance, capacity, and migration safety.
  • Design, implement, and test disaster recovery and business continuity plans, including regular game days and failover exercises.
  • Champion the AI-boosted SRE transformation by evaluating and deploying AI tooling for operational tasks.
  • Partner with Security and Compliance to manage infrastructure's role in PCI-DSS and SOC 2.
  • Drive infrastructure-as-code and platform automation.
  • Own vendor and technology strategy for the infrastructure domain, including build-vs-buy decisions and vendor risk management.
  • Communicate infrastructure risk, investment, and posture to executives, the board, and auditors in business terms.

Benefits

  • Unlimited PTO, volunteer hours and sabbatical
  • Life, STD/LTD, medical, dental and vision insurance
  • Highly discounted LifeTime gym membership
  • 401k with match
  • Collaborative fun co-workers
  • The opportunity to join the fastest growing FinTech alongside a team of motivated and driven individuals
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service