AI Reliability Operations Engineer

Lenovo•Chicago, IL
•$80,000 - $90,000•Hybrid

About The Position

We are looking for an AI Reliability Operations Engineer to support the operational health of Qira's production and non-production systems. Qira is Lenovo’s cross-device Personal AI that works across phones, PCs, and other Lenovo and Motorola products. This role spans system monitoring, alert response, incident response, and observability across the full AI stack, including model performance, inference pipelines, and cloud services. You will also have visibility into SDLC operations across staging and pre-production environments, helping ensure that releases and configuration changes land cleanly. This is a foundational role in keeping Qira stable and available for users around the world.

Requirements

  • Direct experience in incident response or an SRE-adjacent role, not just monitoring or support.
  • Experience with observability tools such as Grafana, Datadog, or cloud-native dashboards.
  • Experience with alerting tools such as PagerDuty or OpsGenie, and ticketing systems such as Jira or ServiceNow.
  • Experience Troubleshooting: can isolate where a problem actually lives (which service, which layer) without being handed the answer.
  • Experience in Azure, including core cloud concepts and how services in an Azure environment are monitored.

Nice To Haves

  • Experience in a technical operations, SRE, or production support environment.
  • Exposure to AI or ML systems, including awareness of how model quality and data pipelines are monitored.
  • Basic scripting ability or comfort reading and adapting existing scripts and runbook commands.
  • Experience working across time zones in a globally distributed team.
  • Working knowledge of SRE concepts: P50/P95/P99 latency, MTTA/MTTM/MTTR (MTTx), the four golden signals (latency, traffic, errors, saturation), and how to apply them to triage.
  • Clear, precise written communication in English, including accurate incident updates under pressure.
  • Ability to work assigned on-call coverage

Responsibilities

  • Perform incident response: contain issues and work cross-functionally with dev teams to fully resolve them.
  • Monitor system health via Grafana dashboards, catching issues early and verifying resolution.
  • Serve as point of contact for change requests (CRs), triaging bug tickets from internal testers to the correct dev group.
  • Keep incident and CR records clear and accurate in ticketing systems.
  • Monitor production and non-production systems using observability dashboards, alerting tools, and AI-specific signals, including model performance, inference latency, and data pipeline health.
  • Watch proactively for early warning signals across Qira's cloud services, device integrations, and AI components, not just respond to alerts after they fire.
  • Review alert thresholds, update runbooks, and flag procedural gaps to the shift lead or SRE. This role is expected to improve the process, not just execute it.
  • Build and improve tooling and scripts to reduce manual, repetitive work across incident and CR handling.
  • Maintain documentation standards: keep runbooks, CR records, and process docs accurate and usable by anyone on the team.
  • Observe and report on SDLC operations across staging and pre-production environments, flagging anomalies and supporting engineering teams during releases and configuration changes.
  • Verify system health before and after deployments and configuration changes, and assist engineering with deployment checks.

Benefits

  • Bonus and/or commission
  • Various benefits can be found on www.lenovobenefits.com
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service