Lead System Engineering

AT&T•Plano, TX
•$158,200 - $237,400•Onsite

About The Position

The Lead System Engineer will be responsible for proactive functionality launched by reviewing user stories and code changes. This role supports the scrum team to configure alerting, monitoring, dashboarding and impact analysis tools to optimize on-call processes and procedures. Key responsibilities include documenting knowledge, conducting post-incident reviews, and driving Mean Time To Repair (MTTR) reduction. The Lead System Engineer will coordinate and guide cloud migration of microservice-based architecture on various cloud environments, and build and implement cloud services for high availability, performance, monitoring, and incident response. This role will enable and provide infrastructure support for DevOps teams, including on-prem and cloud administration. Additionally, the Lead System Engineer will implement and enhance automation frameworks for the delivery of microservices-based applications using Java, J2EE, Jenkins, Maven, Linux, and K8s, on both on-prem and in the cloud. The role involves working with application developers to collect requirements for future releases, implementing monitoring and alerting, creating dashboards for specific metrics, setting thresholds, and triggering alerts. It also includes performing root cause analysis, brainstorming incident resolutions, and providing corrective and preventative measures. The Lead System Engineer will continuously analyze system performance in production, troubleshoot issues, and proactively identify areas for optimization. This role will also work with the team to gather requirements, research, evaluate, design, plan, deploy, and support the ELK stack on Linux, building highly-resilient, high-performance, scalable, and flexible systems. Experience with Azure Cloud, Linux systems administration, and scripting automation is crucial. The role requires designing, analyzing, and troubleshooting large-scale distributed systems, debugging production issues, and developing tools such as Azure Resource Manager (ARM), Terraform, and Ansible. Familiarity with Git, Visual Studio Team Services (VSTS), CI/CD pipelines, Microsoft Azure Public Cloud, PowerShell, Python, RESTful and WebSocket APIs, TCP/IP stack, internet routing, load balancing, Log Analytics, Dynatrace Prometheus, Nagios, Kafka, Docker, Kubernetes, Serverless, Function, and Lambda is essential.

Requirements

  • Bachelor’s degree, or foreign equivalent degree, in Applied Computer Science, Electronics Engineering, or Computer Engineering
  • 5 years of progressive, postbaccalaureate experience in the job offered, or 5 years of progressive, postbaccalaureate experience in a related occupation working with Azure Cloud, Linux systems administration and scripting automation
  • Designing, analyzing and troubleshooting large-scale distributing systems
  • Debugging production issues across services and levels of the stack
  • Developing tools Azure Resource Manager (ARM), Terraform, and Ansible
  • Working with Git and other source control systems
  • Working with Visual Studio Team Services (VSTS)
  • Using tools to create and manage Continuous Integration (CI) and Continuous Delivery (CD) pipelines
  • Working with Microsoft Azure Public Cloud, PowerShell, Python systems automation, RESTful and WebSocket APIs
  • Working on Transmission Control Protocol (TCP) and Internet Protocol (IP) stack, internet routing and load balancing
  • Using technologies including Log Analytics, Dynatrace Prometheus, Nagios, and Kafka
  • Implementing, designing, deploying Docker, Kubernetes, Serverless, Function and Lambda

Responsibilities

  • View site reliability engineers for CLOUD(Azure) to proactive functionality launched by reviewing user stories and code changes.
  • Support scrum team to configure alerting, monitoring, dashboarding and impact analysis tools to optimize on-call processes and procedures.
  • Document knowledge, conducting post-incident reviews, and drive lower Mean Time To Repair (MTTR) reduction.
  • Coordinate and guide cloud migration of microservice based architecture on cloud various environments.
  • Build and implement cloud service for the high availability, performance, monitoring, and incident response.
  • Enable and provide infrastructure support for DevOps team including on-prem and cloud administration.
  • Implement and enhance automation framework for delivery of microservices based arch applications using Java, J2EE, Jenkins, Maven, linux, and K8s, on both on-prem and in cloud.
  • Work with application developers on a day-to-day basis to collect requirements for next release.
  • Implement monitoring and alerting creating dashboards for specific metrics, set thresholds, and trigger alerts based on those thresholds interpret the alerts and automatically heal system.
  • Perform root cause analysis brainstorming session on incident resolutions provide corrective and preventative measures to perform, avoid and mitigate future incidents working with DevOps teams.
  • Demonstrate ability to work across teams to continuously analyze system performance in production, troubleshoot consumer and engineering reported issues, and proactively identify areas in need of optimization.
  • Work with team to gather requirements, research, evaluate, design, plan, deploy, and support the ELK stack on Linux.
  • Build highly-resilient, high-performance, scalable, and flexible systems.
  • Work with Azure Cloud, Linux systems administration and scripting automation.
  • Design, analyze and troubleshoot large-scale distributing systems.
  • Debug production issues across services and levels of the stack.
  • Develop tools Azure Resource Manager (ARM), Terraform, and Ansible.
  • Work with Git and other source control systems.
  • Work with Visual Studio Team Services (VSTS).
  • Use tools to create and manage Continuous Integration (CI) and Continuous Delivery (CD) pipelines.
  • Work with Microsoft Azure Public Cloud, PowerShell, Python systems automation, RESTful and WebSocket APIs.
  • Work on Transmission Control Protocol (TCP) and Internet Protocol (IP) stack, internet routing and load balancing.
  • Use technologies including Log Analytics, Dynatrace Prometheus, Nagios, and Kafka.
  • Implement, design, deploy Docker, Kubernetes, Serverless, Function and Lambda.

Benefits

  • Medical/Dental/Vision coverage
  • 401(k) plan
  • Tuition reimbursement program
  • Paid Time Off and Holidays (based on date of hire, at least 23 days of vacation each year and 9 company-designated holidays)
  • Paid Parental Leave
  • Paid Caregiver Leave
  • Additional sick leave beyond what state and local law require may be available but is unprotected
  • Adoption Reimbursement
  • Disability Benefits (short term and long term)
  • Life and Accidental Death Insurance
  • Supplemental benefit programs: critical illness/accident hospital indemnity/group legal
  • Employee Assistance Programs (EAP)
  • Extensive employee wellness programs
  • Employee discounts up to 50% off on eligible AT&T mobility plans and accessories, AT&T internet (and fiber where available) and AT&T phone
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service