Lead Software Engineer

State FarmDunwoody, GA
Hybrid

About The Position

The Digital Experience (DE) Resiliency & Availability team is seeking two experienced Lead Software Engineers to improve the reliability, resiliency, and engineering effectiveness of critical customer-facing digital experiences. These roles work horizontally across DE, partnering with engineering teams and technology leaders to solve complex reliability challenges, strengthen engineering practices, and reduce customer impact from disruptions. This is an opportunity to apply deep technical expertise beyond a single product or application and influence how resilient software is built and operated across Digital Experience. Both roles are hands-on technical leadership positions that work across engineering teams to solve complex problems and improve how applications are built and operated. The positions have two complementary focus areas: Engineering Excellence, focused on helping teams improve engineering practices and operational maturity; and Reliability Improvement, focused on using engineering telemetry and hands-on problem solving to improve customer-visible availability and performance. Candidates may be considered for either role based on their experience and interests.

Requirements

  • Significant hands-on software engineering experience designing, developing, testing, deploying, troubleshooting, and supporting production applications using modern programming languages such as Java, Python, or similar languages.
  • Experience designing and operating resilient distributed systems, including practical application of patterns for fault tolerance, dependency failures, scalability, recovery, and graceful degradation.
  • Hands-on AWS experience developing and operating cloud-native applications and services.
  • Experience with containers and orchestration platforms such as Kubernetes or OpenShift; Red Hat OpenShift Service on AWS (ROSA) experience is highly preferred.
  • Strong working knowledge of Site Reliability Engineering (SRE) principles and demonstrated experience applying resiliency, observability, operational readiness, and incident-management practices to production applications.
  • Experience with observability and distributed tracing, including using telemetry to troubleshoot complex application and dependency issues and establishing effective monitoring and alerting; Dynatrace experience is preferred, with comparable experience using similar platforms also considered.
  • Experience with modern software delivery practices including CI/CD, automated testing, Git-based workflows, deployment automation, security/vulnerability management, and Infrastructure as Code (IaC) using OpenTofu, Terraform, or similar technologies.
  • Experience conducting root-cause analysis, resolving complex production issues, and translating incident findings into sustainable engineering improvements.
  • Experience using AI-assisted engineering capabilities to improve software development, testing, analysis, troubleshooting, automation, or operational workflows.
  • Demonstrated technical leadership, communication, mentoring, and consulting skills, including the ability to assess technical risk and influence engineering teams without direct authority.

Nice To Haves

  • Red Hat OpenShift Service on AWS (ROSA) experience is highly preferred.
  • Dynatrace experience is preferred, with comparable experience using similar platforms also considered.

Responsibilities

  • Partner with engineering teams to design, build, troubleshoot, and improve resilient, highly available production systems.
  • Provide hands-on technical leadership and consulting across product teams, using engineering expertise to diagnose complex problems and guide implementation of sustainable solutions.
  • Apply Site Reliability Engineering (SRE) principles to improve availability, performance, observability, incident readiness, and operational maturity.
  • Use observability and distributed tracing to identify failure patterns, troubleshoot dependencies, measure customer impact, and drive data-informed reliability improvements.
  • Assess engineering practices and coach teams on automated testing, deployment readiness, monitoring and alerting, production support, and other engineering best practices.
  • Participate in production incident analysis and post-incident reviews, identifying root causes and driving corrective actions that prevent recurrence.
  • Build tooling, automation, dashboards, and AI-assisted engineering capabilities that identify risk earlier and improve engineering effectiveness.
  • Support resiliency exercises, performance testing, GameDays, chaos testing, and other practices that validate how applications behave under failure or degraded conditions.
  • Collaborate with engineers, architects, SRE and platform teams, and technology leaders across organizational boundaries while mentoring engineers and influencing technical direction.

Benefits

  • Compensation is based on our standard 38:45-hour work week
  • Potential starting salary range: $120,000 - $160,000
  • Starting salary will be based on skills, background, and experience
  • High end of the range limited to applicants with significant relevant experience
  • Potential yearly incentive pay up to 15% of base salary
  • Annual raise and bonus
  • Robust health and wellbeing programs
  • State Farm pays most of your healthcare premium
  • Multiple healthcare plan options, including a high deductible plan
  • All medical plans provide 100% coverage for in-network preventative care
  • Access to vision, dental, telemedicine, 24/7 mental health professionals
  • Industry leading training programs
  • Top-notch tuition assistance programs
  • Employee resource groups
  • Mentoring
  • Fertility/IVF/adoption assistance
  • College coaching
  • National discount programs
  • Interactive monthly financial workshops
  • Free financial coaching
  • Generous time off policies
  • Opportunity to initially earn up to 20 days annually plus parental leave
  • Paid holidays
  • Celebration day
  • Life leave (40 hours/year)
  • Bereavement leave
  • Community service/education support days
  • Matching Gift Program
  • Good Neighbor Grant Program
  • Employee Assistance Fund
  • Free financial advisors
  • 401(k) plan with company contributions of up to 7% of your salary
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service