Site Reliability Engineer II

KFCLouisville, KY

About The Position

SRE II is responsible for leading SRE initiatives such as meeting SLI/SLOs for critical applications, deployment of monitoring and logging tools, creation of automation with the goal of reduction of toil within the environment. SRE II leads KFC SRE in delivering a blameless RCA to improve KFC’s systems. They are to work cross-functionally with product owners to agree about SLI/SLOs for each critical application. SRE II is there assist in the design and building of cloud-based infrastructure observability to improve performance, cost, and reliability. SRE II is also the resident expert for the SRE team when it comes to deployment of IaC for a containerized environment.

Requirements

  • 5+ or more years of IT experience
  • 1+ years in a professional technical role with multi-cloud experience preferred: Azure AWS, GCP
  • 1+ years of experience with bash and/or Linux command-line utilities
  • 1+ years of experience in Python, Go, REST APIs, GraphQL.
  • 2+ years experience with industry-standard observability tools such as DataDog, Elastic, Dynatrace, etc.
  • Expertise in CI/CD best practices and methodologies: GitLab
  • Experience with DevOps based Infrastructure as Code toolkits such as Terraform, Ansible, etc.
  • Working experience with Incident and Problem tracking systems (i.e. ServiceNOW, Jira, etc. )
  • Excellent collaboration skills with multiple engineering functions, business leaders, vendors, exhibiting excellent teamwork and strong verbal and written communication skills along with strong troubleshooting and analytical skills
  • Understanding of current trends of large-scale infrastructure environments
  • Proven ability to work autonomously focused on long-term results

Nice To Haves

  • Bachelor’s Degree preferred

Responsibilities

  • Automation – Lead the SRE team’s goal of elimination of TOIL through automation of processes, creation of tools, etc.
  • Infrastructure as Code – Work cross collaboratively with Platform Engineering, Development, and Product teams to design, deploy, and maintain cloud native applications and infrastructure.
  • Observability – Actively monitoring KFC environments utilizing tools such as Data Dog, etc. to proactively identify issues within the KFC production environment. Rotational on-call system to support the infrastructure 24/7/365.
  • Complex Distributed System Troubleshooting – Provide SRE support to KFC platforms and stakeholders through investigation, analysis, leading technical resolution, and post-mortem actions. Deliver root cause analysis and corrective actions after critical incidents within the SLA listed.
  • Design - Collaborate with Product Owners to determine SLI/SLOs for projects within scope. Design and develop a dashboard to hold SLI/SLO standards
  • Deploy - Configure monitoring and logging tools used for generating alerts about the health of our systems and applications

Benefits

  • medical
  • dental
  • vision
  • legal
  • accidental death and dismemberment
  • FSA/HSA (depending on enrolled medical plan)
  • short-term disability
  • long-term disability
  • life insurance
  • 401(k) plan
  • 4 weeks of vacation
  • paid sick leave
  • 10 paid holidays
  • a floating day off
  • half day Fridays year-round
  • 2 paid days for volunteer time each calendar year
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service