Major Incident Response Production Support

FidelityWestlake, TX
Onsite

About The Position

We are seeking a technical problem solver with a strong background in production support to join our Major Incident Management team. This role combines Major Incident Management responsibilities with hands-on technical expertise across cloud, infrastructure, and application environments. The ideal candidate is passionate about troubleshooting, thrives in high-pressure situations, and is eager to grow their technical and coordination skills in a dynamic environment.

Requirements

  • Bachelors or equivalent with 2+ years of experience or Masters with 0+ years of experience.
  • A minimum of 2 + years of hybrid experience in Production Support, Development or SRE Experience.
  • Hands-On experience developing or supporting highly distributed multi-tiered systems at scale.
  • Ability to triage while leading an incident call, perform root cause analysis, and be decisive under pressure.
  • Solid understanding of Cloud Computing and DevOps concepts including CI/CD Pipelines.
  • Perform second-level support using visibility tools such as Datadog and Splunk and expert hands on experience with one or more (Datadog, Splunk, Grafana).
  • Understanding of ITIL processes (Incident, Problem, Change Management).
  • Effective business communication and influencing skills.
  • Cloud Platforms: AWS, Azure
  • Operating Systems: Unix, Linux, Windows Server
  • Scripting and Development: Shell Script, .NET
  • Middleware and Integration: Tomcat, Apache
  • Monitoring Tools: Splunk, Datadog, Grafana
  • Databases: Oracle DB, MS-SQL, Sybase
  • ITSM Tools: JIRA, ServiceNow

Nice To Haves

  • Relevant certifications (ITIL, AWS, Azure) are a plus.
  • Retail trading , Asset Trading, or Financial Services Technology experience is a plus.

Responsibilities

  • Lead incident calls, perform root cause analysis, and make decisive actions under pressure.
  • Provide second-level support using visibility tools such as Datadog and Splunk.
  • Support a wide range of environments including application, distributed systems, mainframe, network, cloud, end-user computing, and security.
  • Act as the first line of defense for production incidents, providing detection, coordination, escalation, communication, and mitigation across the enterprise.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service