Site Reliability Engineer

Ancestry
$106,029 - $118,503Hybrid

About The Position

You will help maintain high site availability and ensure the best customer experience through complex technical problem resolution, creating and improving procedures, facilitating communication, participating in active incidents, developing in-house tooling, and effectively leveraging AI agentic workflows. Your role is important to the success of Ancestry and requires both depth and breadth in many cloud operations technologies including an understanding of observability practices, incident management, and software development. You will participate in team meetings and team projects as assigned. You will split your time between active incident triage and mitigation, and software development and automation.

Requirements

  • 4+ years in a continual, high-availability enterprise production environment
  • 4+ years in managing and monitoring AWS technologies and services with tools like New Relic or AWS Application Signals
  • 2+ years in an enterprise level NOC or Command Center environment
  • 2+ years in developing AI agentic workflows and knowledge of best practices in prompt engineering, context management, and AI SDLC
  • 4+ years in scripting, automation, and software development with knowledge of various languages (e.g. Python, Java,, Bash, Powershell)
  • Demonstrated logical thought processes through new technologies, systems, concepts and procedures, and the ability to use reports and data to improve operational results
  • Document issues leveraging defined processes
  • An understanding of TCP/IP LAN/WAN networking technologies and troubleshooting techniques, virtualization technologies
  • Knowledge of hardware or software-based firewalls, load balancers, intrusion detection systems and proxy servers
  • Knowledge of relational and NoSQL database technology query languages (MSSQL, MySQL, Cassandra and CouchDB, etc)
  • Knowledge of CI/CD tools like Harness, Chef, or Puppet

Responsibilities

  • Provide frontline support, including the ability to provide technical solutions, for a complex system including a variety of cloud infrastructure technologies, databases, and network systems
  • Analyze data from monitoring tools and take appropriate actions based on patterns, outliers, or anomalies
  • Partner with development teams to find ways to reduce the time to issue detection and resolution
  • Manage bridges with on-call rotational support groups and executives during outage events or as deemed necessary
  • Report on Incident cases through our automation tools and ticket tracking system
  • Research, proof, and author technical documentation
  • Create and maintain AI agentic workflow context
  • Assist in automation and tool development to reduce toil
  • Participate in team meetings and team projects - both independently and as assigned
  • Be available to cover other shifts (i.e.: days, weekends, and holidays)

Benefits

  • health, dental and vision
  • bonus
  • equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service