Manager, Site Reliability Engineering

DriveWealth•Chicago, IL
•$150,000 - $170,000•Hybrid

About The Position

As the Manager of Site Reliability Engineering, you'll lead a team of SRE Automation Engineers while remaining a hands-on technical authority for our Brokerage-as-a-Service platform. This isn't a purely people-management seat, you're expected to bring the same principal-level SRE depth to automation design and engineering as an individual contributor, while also building the team, setting technical direction, and developing your engineers' careers. This role is centered on reducing manual toil through engineering, applying Google's SRE principles: SLOs, error budgets, blameless postmortems, and systematic toil reduction, adapted to a regulated brokerage environment. You'll carry two responsibilities at once: driving the automation agenda, building and orchestrating workflows in Rundeck and Airflow to eliminate repetitive work, and growing your team of SRE Automation Engineers into a high-functioning automation practice. You'll guide the design of internal SRE platforms, automate complex workflows, and ensure our Kubernetes-based and colo ecosystems can handle the demands of global financial markets, while owning the people side of the team: mentorship, performance, and growth, and the day-to-day management of the team's Jira board. While this role includes participation in on-call rotations supporting our 24/7 global operations, your primary mission is to build systems that make manual intervention obsolete, and a team capable of sustaining that mission.

Requirements

  • Prior experience managing or leading SRE/DevOps engineers, ideally in a fintech or highly regulated environment. Able to flex between hands-on principal-level engineering and coaching and developing a team.
  • Working knowledge of Google's SRE practices—SLIs/SLOs, error budgets, toil reduction, and blameless postmortems—and experience adapting them to a regulated environment.
  • Proficient in Linux administration with a deep understanding of the TCP/IP stack, OSI model, DNS, and network troubleshooting.
  • Experience working in highly regulated financial environments or with FIX/API connectivity.
  • Hands-on experience managing production-grade clusters, including RBAC, autoscaling, Helm, and multi-cluster patterns.
  • Strong grasp of AWS core services, security, and high-availability patterns. Proficiency with boto3 and AWS CLI for automation.
  • Experience building secure, automated delivery pipelines and operating GitOps workflows (ArgoCD).
  • Strong scripting and development skills in Python or Golang, along with Bash and Ansible.
  • Experience with Grafana/Similar tools, Prometheus, Understanding of logs shipping, management and metric first alerting.
  • Experience with secrets management, vulnerability scanning, and securing the software supply chain.
  • Familiarity with using LLMs, Public MCPs, or Bedrock Agent Core to enhance SRE workflows.
  • Hands-on experience with Rundeck and Airflow for job orchestration and automation, plus experience managing Kafka, MQ, or SQS.

Responsibilities

  • Manage, mentor, and grow a team of SRE Automation Engineers—setting technical direction, running 1:1s, owning performance management and career development, and managing the team's Jira board to prioritize and track sprint work.
  • Lead the design and development of internal tooling and automation—including Rundeck and Airflow-based orchestration—to eliminate repetitive manual toil and improve developer velocity, staying hands-on with the most complex, highest-leverage automation work yourself.
  • Adapt Google's SRE principles to our environment—defining SLIs, SLOs, and error budgets, and using them to guide engineering and operational priorities.
  • Set architectural standards for modular, reusable IaC using Terraform and oversee GitOps workflows via ArgoCD.
  • Review software architecture and Kubernetes metrics to ensure high availability, capacity planning, and cost-optimization across AWS regions, and hold the team accountable to those standards.
  • Lead incident response for critical events, drive complex root-cause analysis (RCA), and champion a blameless post-mortem culture across the organization.
  • Partner with engineering leadership to align SRE priorities with business goals, and foster adoption of new tools, security standards, and reliability best practices across teams.

Benefits

  • competitive compensation
  • equity
  • 401(k) match
  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Disability insurance
  • Paid Parental Leave
  • wellness reimbursement
  • company-provided phone
  • personal development allowance
  • generous Paid Time Off (PTO)
  • observed holidays
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service