Production Operations Lead

Group 1001Zionsville, IN

About The Position

The Production Operations Lead is responsible for the day-to-day operational leadership of a Prod Ops squad supporting the stability, availability, and reliability of Group 1001's production systems. This role serves as the primary escalation point within the team, coordinating incident response across multiple engineering teams, facilitating problem management to root cause resolution, and ensuring change management processes are followed consistently. The Production Operations Lead guides and mentors Production Operations Analysts, ensuring effective triage, ticketing, escalation, and stakeholder communication during service-impacting events. The Production Operations Lead is accountable for operational outcomes: SLA adherence, notification cadence, shift handoff quality, and process compliance. This role bridges the gap between technical engineering teams and business stakeholders — translating incident impact into clear executive communication while keeping technical responders focused on resolution. An effective Production Operations Lead will demonstrate the ability to manage high-severity incident bridges, enforce escalation timelines, drive post-incident root cause analysis to completion, and build trust with cross-functional teams through consistent process execution.

Requirements

  • Bachelor's degree from an accredited college or university in a related technical discipline, or the equivalent combination of education, technical certifications or training, or work experience
  • 5+ years of experience in IT operations, NOC, or production support environments
  • 2+ years in a lead, senior, or supervisory capacity within an operations team
  • Demonstrated experience coordinating incident response across multiple teams, including bridge facilitation and stakeholder communication
  • Strong understanding of ITIL processes — specifically Incident, Problem, and Change Management
  • Experience with ITSM/ticketing platforms (ServiceNow, Jira Service Management, or equivalent)
  • Experience with monitoring and observability tools (Grafana, Datadog, CloudWatch, PagerDuty, or similar)
  • Excellent written and verbal communication skills — able to produce concise executive-level status updates and facilitate focused technical bridges simultaneously
  • Proven ability to work effectively under pressure, particularly during high-severity incidents requiring rapid decision-making and clear communication
  • Strong organizational skills with the ability to manage multiple concurrent incidents, problems, and operational tasks
  • Ability to mentor and develop team members while maintaining operational accountability

Nice To Haves

  • ITIL v4 Foundation certification (or higher)
  • Experience in financial services, insurance, or other regulated industries
  • Familiarity with SLO/SLI frameworks, error budget policies, and reliability engineering principles
  • Experience with Jira workflow configuration and operational reporting
  • Background in cloud infrastructure environments (AWS, Azure, GCP)
  • Experience onboarding new teams or platforms into operational support models
  • Familiarity with SOC 2 compliance requirements as they relate to incident and change management

Responsibilities

  • Serve as Incident Manager for P0/P1/P2 production incidents — spin up bridges, coordinate technical responders, enforce escalation timelines, and deliver stakeholder notifications per defined cadence
  • Assign, prioritize, and oversee BAU (day-to-day) operational work across the squad — including access requests, runbook updates, monitoring tuning, and operational maintenance
  • Guide and mentor Production Operations Analysts (Level I and II), ensuring effective triage, documentation, escalation, and communication during incidents
  • Facilitate Problem Review Board (PRB) meetings within 48 hours of incident resolution and drive root cause analysis to completion
  • Coordinate After Action Reviews (AARs) within one week of significant incidents, ensuring corrective actions are assigned with owners and due dates
  • Identify recurring incident patterns and proactively create problem tickets when thresholds are met (3+ incidents, same system, within 30 days)
  • Represent Prod Ops at the Change Advisory Board (CAB) — review changes for risk, validate implementation plans, and provide operability sign-off for changes affecting monitoring, alerting, or escalation paths
  • Enforce customer-facing communication requirements (2 business day stakeholder notice after approval) and year-end freeze policies
  • Ensure shift handoff quality — incoming shifts receive full context on active incidents, ongoing changes, and outstanding action items
  • Maintain squad-level KPIs: incident response times, notification compliance, change success rate, and problem RCA completion within SLA
  • Serve as the primary Prod Ops liaison for assigned platform or business unit, building relationships with engineering leads, app owners, and business stakeholders
  • Lead monthly proactive alert reviews with SRE to identify systemic issues before they escalate to major incidents
  • Validate operational readiness for new services onboarding to production — ensuring monitoring, alerting, runbooks, and escalation paths exist before go-live
  • Participate in post-incident reviews to improve response processes and update Method & Procedure documentation

Benefits

  • Comprehensive health, dental, and vision insurance plan options
  • Basic and Supplemental Life Insurance
  • Short and Long-Term Disability
  • Employee Assistance Program
  • Wellness programs
  • 401K plan, with matching contributions by the Company
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service