About The Position

The Technical Support team at OpenAI is responsible for ensuring that developers and enterprises can reliably build mission-critical solutions using OpenAI models. This role involves providing technical guidance, resolving complex issues, and supporting customers in maximizing value and adoption from deploying our highly-capable models. The team collaborates closely with Technical Success, Product, Engineering, and other departments to deliver the best possible customer experience at scale, with an automation-first mindset and leveraging AI to scale support operations. This Senior Support Engineer position will collaborate directly with strategic enterprise accounts and product teams, helping solve difficult customer problems. The role is considered the last line of defense before the core Engineering team and will provide technical guidance for the most challenging issues. The Senior Support Engineer will design and run operational processes to monitor top strategic customers and a 24x7 response team, working closely with Infrastructure and Engineering teams. This role is crucial for the success of innovative, disruptive, and high-scale AI solutions built with the OpenAI API platform, characterized by low volume and high difficulty. The position is based remotely in Toronto, Canada.

Requirements

  • Bachelor’s degree in Computer Science or a related field.
  • A strong software engineering foundation.
  • 8+ years of experience in technical operations roles such as SRE/NOC, designing monitoring systems and resolving production issues in fast-paced and mission-critical environments.
  • A strong track record of troubleshooting complex technical problems at the systems level.
  • Deep familiarity with modern monitoring, alerting, and observability practices.
  • Hands‑on experience setting up or managing metrics, logging, and tracing for distributed systems (e.g., understanding of SLIs/SLOs, alert tuning, dashboard creation).
  • Proven experience leading incident response for high‑severity outages or service disruptions.
  • Able to perform real‑time incident coordination, root cause analysis, and drive follow‑ups (post‑mortems, action items) to prevent recurrence.
  • Knowledge of industry best practices for incident management and fault diagnosis.
  • Strong skills in scripting or software engineering (e.g., Python or similar) to automate repetitive tasks and integrate tools.
  • Solid understanding of cloud infrastructure and distributed systems fundamentals.
  • Comfortable working with cloud services, load balancers, databases, and containerized applications.
  • Effective at working cross‑functionally in a high‑trust environment.
  • Strong communication skills to explain technical issues and resolutions to both engineering and non‑technical stakeholders.
  • Able to coordinate efforts across teams and comfortable providing updates in the midst of an ongoing incident.

Responsibilities

  • Be among the foremost technical and troubleshooting experts for our API platform at OpenAI.
  • Proactively identify and implement opportunities to scale support operations by leveraging automation and advancements in AI technologies.
  • Configure and use advanced monitoring and alerting workflows to proactively detect customer impacting issues in real time.
  • In partnership with engineering, contribute to reliability reviews and preparedness for new features, launches, or strategic customer requirement updates.
  • Design and refine incident response processes and documentation across strategic customers, engineering and support teams.
  • Analyze operational metrics and incident RCAs to identify areas for improvement.
  • Provide support coverage during holidays and weekends based on business needs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service