About The Position

Caseware is a leading global audit and accounting software company with over 30 years of experience and a significant user base across 130 countries. This is a senior engineering role focused on enhancing production resilience, security, operational excellence, and the developer experience. The position involves designing, building, and evolving foundational systems, tooling, and operational practices to enable engineering teams to deliver secure, reliable, and scalable software. Key responsibilities include establishing reliability standards, defining SLOs, improving observability, automating operations, and strengthening incident management and post-incident learning. The role requires close collaboration with Engineering, Security, Platform, and Product teams to architect scalable distributed systems, optimize Kubernetes and AWS infrastructure, and build automated delivery pipelines for rapid and safe software releases. The goal is to reduce operational toil, improve system performance, increase platform reliability, and support business growth.

Requirements

  • 8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or related cloud-native engineering roles.
  • Deep expertise in AWS services, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security best practices.
  • Advanced experience operating and scaling production Kubernetes environments.
  • Strong hands-on experience with Istio service mesh, including traffic management, security, observability, and resiliency.
  • Proven expertise with Infrastructure as Code (IaC), preferably using AWS CDK.
  • Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms.
  • Strong troubleshooting, performance optimization, and incident management experience in distributed systems.
  • Excellent communication, collaboration, and technical leadership skills.
  • Experience designing and operating monitoring, logging, tracing, and alerting solutions for cloud-native platforms.
  • Strong knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling.
  • Experience defining and operationalizing SLIs, SLOs, alerting strategies, runbooks, and reliability metrics.
  • Proven ability to leverage observability data to improve service reliability, reduce incident impact, and optimize operational performance.
  • Strong proficiency in TypeScript and Node.js for platform engineering, automation, and operational tooling.
  • Experience building and maintaining scalable backend services, APIs, and event-driven systems.
  • Deep understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking.
  • Experience implementing zero-trust architectures, mTLS, and service-to-service security controls.
  • Commitment to high-quality engineering practices, including automated testing, code reviews, and observability-driven development.
  • Strong understanding of resilience engineering, including autoscaling, disruption management, failure testing, and safe deployment strategies.

Nice To Haves

  • Experience with progressive delivery practices such as canary, blue/green, and feature-flag-based deployments.
  • Experience working in regulated, compliance-driven, or security-sensitive SaaS environments.
  • Familiarity with FinOps principles and cost optimization strategies for cloud platforms.
  • Experience building internal developer platforms and self-service engineering tooling.
  • Cloud-native certifications such as CKA, CKAD, CKS, KCSA, or KCNA.
  • Kubestronaut certification or equivalent advanced Kubernetes expertise is highly regarded.

Responsibilities

  • Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes.
  • Design, implement, and continuously improve deployment, release, and rollback strategies across complex distributed systems.
  • Establish secure-by-default CI/CD pipelines with robust automation, governance, and policy-driven controls.
  • Enhance platform observability through metrics, logs, tracing, and actionable alerting to improve system visibility and operational efficiency.
  • Define, implement, and mature Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability standards across the organization.
  • Lead response efforts for high-severity incidents, ensuring timely resolution, effective communication, and meaningful post-incident reviews that drive continuous improvement.
  • Partner closely with engineering teams to strengthen platform standards, improve service resilience, optimize runtime performance, and embed reliability best practices.
  • Mentor and guide engineers on cloud-native technologies, site reliability engineering principles, and operational excellence practices, fostering a culture of continuous learning and accountability.

Benefits

  • competitive salary
  • comprehensive benefits
  • health insurance
  • retirement plans
  • discretionary bonus
  • commission
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service