Senior Site Reliability Engineer – APAC Region

Nucleus Security,
Remote

About The Position

Are you looking for more in life than just building another web app? Does upending cyber security resonate with you? We're a growth stage cyber security startup that is paving the way forward for how vulnerability management is run in large enterprise organizations. For our customers, vulnerability management has always been a game of catch up, with limited asset coverage and manual processes. Nucleus’ core goal is to build a fast and scalable platform that solves these problems and many more so that vulnerability management isn't just possible, it's easy. We're looking for a passionate Senior Site Reliability Engineer to join our growing team of engineers APAC region.

Requirements

  • 8+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, Infrastructure Engineering, or related field.
  • Strong hands-on experience with cloud platforms, including AWS, GCP, Azure, and/or OpenShift (OCP).
  • Deep experience with Kubernetes, containers, and production container orchestration.
  • Experience building and maintaining highly available, scalable, and secure production infrastructure.
  • Strong experience with Infrastructure as Code, preferably Terraform, and configuration/automation tools such as Ansible.
  • Strong scripting and automation skills using Python, Bash, or similar languages.
  • Experience building and maintaining CI/CD pipelines using GitHub, GitLab, Bitbucket, or similar platforms.
  • Strong experience with observability and monitoring platforms such as Prometheus, Grafana, Loki, CloudWatch, or equivalent tools.
  • Experience with incident response, root-cause analysis, production troubleshooting, and reliability engineering practices.
  • Experience using AI-assisted engineering tools to improve infrastructure automation, troubleshooting, documentation, code generation, or operational workflows.
  • Ability to identify opportunities where AI and automation can reduce operational toil, improve signal detection, and accelerate incident investigation.
  • Strong understanding of Linux, networking, security, cloud architecture, and distributed systems.
  • Ability to provide technical leadership, mentor engineers, and help drive a culture of automation, reliability, and continuous improvement.
  • Must be located in the APAC region and have citizenship in the country you currently reside.
  • Minimum 8 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.
  • Strong hands-on experience with cloud providers like AWS, GCP, Azure, and OpenShift (OCP).
  • Strong experience with Kubernetes, Infrastructure as Code (terraform, tofu, CloudFormation), and automation (Python, bash, PowerShell).
  • Proven experience supporting highly available production systems, including observability, incident response, troubleshooting, and reliability improvements.

Nice To Haves

  • Experience integrating LLMs or AI-enabled tools into engineering or operational workflows.
  • Familiarity with AI-assisted log analysis, anomaly detection, incident summarization, or root-cause investigation.
  • Experience building internal automation or tooling that combines APIs, scripting, infrastructure data, and AI models.
  • Understanding of how to use AI safely in production engineering environments, including data security, access controls, validation, and human review.

Responsibilities

  • Maintain Reliable, Secure, and AI-Assisted Production Operations: Keep production systems highly available, secure, patched, and performant. Use AI-assisted tooling to accelerate troubleshooting, identify risks, analyze incidents, and improve operational response.
  • Build and Maintain Kubernetes, Cloud, and DevOps Infrastructure: Own and improve Kubernetes clusters, containerized workloads, Infrastructure as Code, CI/CD pipelines, and cloud infrastructure. Leverage AI-assisted development and automation tools to improve delivery speed, configuration quality, and operational consistency.
  • Build Observability and Automation That Reduces Toil: Improve monitoring, alerting, logging, dashboards, and automated remediation to identify issues earlier and reduce repetitive operational work. Apply AI and intelligent automation to correlate signals, surface anomalies, assist with root-cause analysis, and automate common SRE workflows.

Benefits

  • 100% company-paid health, dental, vision, life, and short-term disability insurance options
  • Generous 401k contribution (not a match)
  • Flexible PTO + 10 company holidays
  • Equity in a high-growth, VC-backed startup
  • Fantastic company culture
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service