Site Reliability Engineer

OnBoard
Remote

About The Position

The Cloud Operations Engineer III is a senior member of the cloud operations team, responsible for the reliability, observability, performance, and operational security of our multi-product SaaS platform. This role owns our Datadog observability practice — instrumentation standards, dashboards, SLOs, monitors, and alert routing — and leads the migration off our legacy monitoring stack. It is an engineering role, not a ticket-queue role: the expectation is that recurring operational work gets replaced with code. The Cloud Operations Engineer III participates in on-call, incident response and is measured on fewer customer-impacting incidents, faster detection and recovery, and less manual work year over year. The ideal candidate is a proactive problem-solver who thrives in dynamic, evolving environments and works effectively across departments to address complex challenges. They have experience partnering with cross-functional teams to understand and document requirements, then translating those needs into meaningful dashboards that improve service visibility (Observability) and support informed decision-making. They are passionate about automation, process improvement, and eliminating unnecessary manual effort. They confidently propose better approaches when opportunities for improvement arise.

Requirements

  • Bachelor's degree in Computer Science, Information Technology, or a related field, or equivalent practical experience.
  • 5-7 years of professional experience in cloud operations, site reliability, platform, or DevOps engineering for production SaaS systems.
  • Demonstrated hands-on depth with a modern observability platform — Datadog strongly preferred — including APM and distributed tracing, log pipelines and indexing controls, dashboards, monitors, and SLOs.
  • Strong scripting and automation ability in PowerShell, with the judgment to write tooling that other engineers can safely operate.
  • Strong knowledge of containers, container orchestration, and the Kubernetes ecosystem, including autoscaling, cluster upgrades, and diagnosing pod-level failures.
  • Production experience with Azure — Kubernetes Service, Azure SQL, Cosmos DB, Redis, Service Bus, Key Vault, and Entra ID — or equivalent depth in another major cloud.
  • Experience with infrastructure-as-code and CI/CD pipeline authoring (Bicep or Terraform; Helm or Kustomize; Azure DevOps preferred).
  • Proven incident response experience in a customer-facing production environment, including on-call participation and leading post-incident reviews.
  • Experience operating multi-region, multi-tenant systems.
  • Strong knowledge of platform security and operational best practices: secret and key rotation, least-privilege access, and vulnerability remediation.
  • Excellent problem-solving and analytical abilities, with strong written communication for runbooks, incident updates, and technical proposals.
  • Strong communication, and teamwork skills, including the ability to work effectively with legacy systems and their constraints.

Nice To Haves

  • experience migrating from a legacy monitoring stack to a consolidated observability platform
  • relevant Azure, Kubernetes, or Datadog certifications

Responsibilities

  • Own the Datadog platform across all products and environments, including agent lifecycle, instrumentation standards, unified service tagging, and per-cluster configuration.
  • Instrument services for APM and distributed tracing, log collection, and synthetic monitoring; partner with engineering teams to close instrumentation gaps in both legacy and modern codebases.
  • Build and maintain the dashboard, monitor, and SLO catalog; define SLIs and error budgets for critical user journeys and use them to drive prioritization with engineering and product.
  • Design high-signal alerting: reduce noise and duplicate alerts, tune thresholds, and ensure every alert has an owner and a runbook.
  • Develop and maintain automation in PowerShell, Python, and Bash for provisioning, configuration, diagnostics, remediation, and reporting.
  • Extend our infrastructure-as-code estate — Bicep modules, Kubernetes manifests, Helm releases, and Azure DevOps pipeline templates — so environments and regions are reproducible and drift-free.
  • Convert manual runbooks into automated or self-service workflows: cluster upgrades, secret and certificate rotation, tenant provisioning, data retention purges, and access provisioning.
  • Implement and maintain platform security controls and audit-ready operational evidence: managed identities, secret and key rotation, least-privilege access, and image and dependency scanning.
  • Author and maintain runbooks, on-call guides, and architecture documentation, and provide technical leadership and mentorship to junior engineers on observability, automation, and incident response.

Benefits

  • Fully remote work with company provided equipment (laptop, software, etc.)
  • Employment with a growing, casual, fun, philanthropic minded company
  • US Based Employees
  • Comprehensive, high-quality medical/prescription drug plan options, as well as dental and vision plan offerings.
  • An employer contribution to your Health Savings Account (HSA) if you participate in a High Deductible Healthcare Plan.
  • Medical Flexible Spending Accounts available.
  • Dependent Care Flexible Spending Accounts available.
  • Basic life insurance in the amount of $50,000 or 1 X’s your salary (whichever is higher).
  • Short and long-term disability and Accidental Death and Dismemberment benefits at no cost to you.
  • 401K Retirement Savings Plan with automatic enrollment at the first of the month following 60 days of employment at 5% to help you secure your financial freedom. We offer a generous company match that starts on the first of the month following 60 days of employment. The company match is dollar for dollar on the first 3% of your pay that you contribute and $0.50 on the dollar on the next 2%, for a total match of 4%.
  • Paid Time Off (PTO)/Holiday
  • CAN Based Employees Employer paid Life and Accidental Death Insurance
  • Contribution to Health Care Spending Account
  • Dependent Life Insurance
  • Optional Life Insurance
  • LTD Insurance
  • Drug and Paramedical Coverage
  • Dental Insurance
  • Vision Insurance
  • EAP
  • AUS Based employees Superannuation rate of 12%
  • Monthly stipend of $400 AUD to purchase private medical insurance
  • UK Based Employees (via EPG) Pension - Aegon
  • Passageways/OnBoard contributes 8% of the employee's basic salary
  • Employees can contribute up to 100% of salary subject to max limits
  • Enrolled from Day 1 of employment
  • Private Medical Insurance
  • Life Assurance
  • Income Protection
  • Critical Illness
  • Employee Assistance Programme
  • Serious Illness Benefit
  • Help@Hand
  • Cashplan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service