Senior SRE

Banyan Software
•$145,000 - $170,000•Remote

About The Position

We are seeking a highly experienced and hands-on SRE to own the operational excellence of the modernized SaaS applications produced by the Banyan AI Factory. This is not a role focused on building the factory itself; instead, you will run the reliability of the modernized applications the factory delivers to our Operating Companies (OpCos). You will join a team that provides 24x7 coverage with rotating on-call responsibilities, serving as Tier 1 Site Reliability Engineering (SRE) for our OpCos’ distributed applications. Day to day this will include: automated deployments, cloud service integration, application performance and availability monitoring/observability, and security incident response across our two target clouds — Amazon Web Services (AWS) and Microsoft Azure. The ideal candidate has a track record of keeping secure, highly available production systems running at scale.

Requirements

  • 5–7 years of progressive experience in Software Engineering, and/or Site Reliability Engineering, with a focus on operating distributed systems.
  • Deep expertise in Python, Javascript, or Go. Building automation and integrations between tools. This may be with AI assistance, but you must have a deep understanding of the code and scripting principals such as: authentication, parallelization, triggering, APIs, data transformation, etc.
  • Deep expertise in container technologies (Docker/Kubernetes) supporting highly scalable and resilient distributed systems.
  • Have experience working with modules at scale. This is a requirement for the role.
  • Hands-on experience operating production workloads on Amazon Web Services (AWS) (e.g., EC2, Lambda, EKS, S3, RDS) and / or Microsoft Azure (e.g., Container Apps, AKS, Container Storage).
  • Deep history of hands-on work with CI/CD platforms (GitHub Actions, GitLab CI) and embedding DevSecOps practices directly into operational workflows.
  • Experience with application level logging, troubleshooting, and tracing tools, with a proven track record operating highly available production systems.
  • Experience with AI-assisted engineering tools such as Claude Code or similar
  • Familiarity with APM tooling and practices (e.g., Datadog, New Relic, Dynatrace, or similar) to instrument, profile, and optimize application performance in production.
  • Demonstrated experience participating in on-call rotations, responding to production and security incidents, and executing disaster recovery procedures.
  • Exceptional communication, presentation, and collaboration skills, with a proven ability to coordinate across teams.
  • Bachelor’s degree in Computer Science or a related technical field.

Nice To Haves

  • Familiarity with advanced cloud security tools like Wiz, Prisma Cloud, and Checkov.

Responsibilities

  • Operate as part of a team providing round-the-clock coverage of OpCo containerized applications, participating in a rotating on-call schedule to ensure continuous availability and rapid response.
  • Serve as Tier 1 SRE for the modernized applications, managing day-to-day cloud integrations across our two target clouds — AWS and Azure — to keep production systems healthy, performant, and secure.
  • Implement and maintain robust application observability tooling (monitoring, logging, tracing) to track performance and availability, proactively detect degradation, and drive down mean-time-to-detect and mean-time-to-resolve.
  • Develop, maintain, test, and execute disaster recovery and business continuity procedures. Ensure the timely recovery and restoration of services following geographic disruptions, cyber incidents, infrastructure failures, or other disaster events.
  • Respond to security incidents and operational events affecting OpCo SaaS platforms, executing established runbooks, coordinating remediation.
  • Use Infrastructure-as-Code (Terraform) and CI/CD pipelines (e.g., GitHub Actions, GitLab CI) to manage, deploy, and automate the operational environments of modernized applications, reducing toil and improving consistency.
  • Build scale in our DevSecOps practice by designing, building, and operating AI agents that automate SRE tasks and incident response, reducing toil and accelerating detection, triage, and remediation.
  • Serve as a technical escalation point for operational challenges, applying strong analytical skills to resolve infrastructure, network, and automation issues across distributed, multi-tenant SaaS environments while navigating technical ambiguity.

Benefits

  • annual bonus
  • equity (when applicable)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service