Platform Operations Engineer

NRGHouston, TX

About The Position

The Platform Operations Engineer is responsible for coordinating and improving the operational effectiveness of the Home Services technology platform. This role partners across Engineering, Product, Architecture, Quality Assurance, and Platform teams to improve platform reliability, operational visibility, and service performance through coordination, reporting, operational excellence, and continuous improvement. The Platform Operations Engineer supports production operations, operational reporting, environment coordination, monitoring and observability, and AI-enabled operational capabilities to ensure engineering teams have the visibility, processes, and operational support necessary to deliver reliable technology solutions.

Requirements

  • Bachelor's degree in Information Systems, Computer Science, Engineering, or a related field, or an equivalent combination of education and relevant work experience.
  • Five (5) or more years of experience in Platform Engineering, Application Support, IT Operations, DevOps, Cloud Operations, Site Reliability Engineering, or a related technical discipline.
  • Experience supporting mission-critical production applications within an enterprise environment.
  • Experience coordinating production support activities across multiple Engineering and business teams.
  • Experience developing, monitoring, and reporting operational KPIs, dashboards, and metrics.
  • Experience with application monitoring, logging, observability, and alerting platforms.
  • Experience coordinating environment management activities, including planning, release readiness, deployments, refreshes, and cross-team scheduling.
  • Experience supporting cloud platforms and enterprise applications, preferably Microsoft Azure and/or AWS.
  • Experience working with IT Service Management (ITSM) processes, including Incident, Problem, Change, and Release Management.
  • Strong analytical, troubleshooting, organizational, and problem-solving skills with the ability to identify trends and recommend operational improvements.
  • Excellent verbal and written communication skills with the ability to collaborate effectively across technical and business teams.
  • Experience with enterprise operational tools such as Azure DevOps, GitHub, ServiceNow, Azure Monitor, Application Insights, Datadog, Splunk, or similar platforms.
  • Experience applying AI, automation, or scripting to improve operational efficiency, monitoring, reporting, or support processes.

Responsibilities

  • Coordinate production support activities across Engineering, Product, and Platform teams to facilitate timely incident resolution and effective communication.
  • Support incident management processes by coordinating investigations, root cause analysis activities, corrective actions, and post-incident follow-up with the appropriate engineering teams.
  • Coordinate operational readiness activities for releases, maintenance events, and platform changes.
  • Track recurring operational issues and coordinate continuous improvement initiatives with responsible teams.
  • Develop, maintain, and communicate operational dashboards, KPIs, and service performance metrics.
  • Analyze operational data to identify trends, risks, and opportunities for operational improvement.
  • Coordinate recurring operational reviews and provide visibility into platform health, service levels, and operational performance.
  • Support the definition, measurement, and reporting of operational objectives and service quality indicators.
  • Partner with engineering teams to ensure appropriate instrumentation, logging, monitoring, and alerting are implemented across platform services.
  • Identify gaps in observability and coordinate improvements with engineering teams.
  • Support adoption of monitoring standards and operational reporting practices that improve proactive issue detection and platform visibility.
  • Promote consistent telemetry and operational reporting across platform services.
  • Coordinate environment planning, scheduling, availability, and utilization across multiple environments.
  • Facilitate environment requests, refreshes, deployments, and conflict resolution activities.
  • Maintain visibility into environment readiness, dependencies, and operational risks.
  • Communicate environment status, planned activities, and potential impacts to stakeholders.
  • Identify opportunities to leverage AI and automation to improve production support, operational efficiency, and platform reliability.
  • Partner with engineering teams to implement AI-assisted operational workflows, monitoring, and support processes.
  • Support adoption of AI-enabled operational tools that improve issue detection, operational insights, knowledge management, and engineering productivity.
  • Evaluate emerging AI capabilities and recommend practical applications that enhance platform operations.
  • Support development and maintenance of operational documentation, runbooks, standard operating procedures, and knowledge resources.
  • Identify opportunities to improve operational processes through automation and standardization.
  • Coordinate operational improvement initiatives that enhance platform reliability and support efficiency.
  • Support operational governance by providing reporting, metrics, and operational insights.
  • Partner with Platform Architects, Solution Engineers, Engineering teams, Quality Assurance, and Delivery teams to improve operational effectiveness.
  • Coordinate cross-functional activities, dependencies, and communications impacting platform operations.
  • Provide visibility into operational risks, dependencies, and platform readiness to support informed decision-making.
  • Foster collaboration and continuous improvement to enhance operational maturity and service performance.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service