Staff - Artificial Intelligence Operations

Bread Financial•Chadds Ford Township, PA
•$139,800 - $279,600•Hybrid

About The Position

The Staff Artificial Intelligence Operations Engineer will implement, optimize, and support Artificial Intelligence (AI) and Machine Learning (ML)-enabled operational capabilities that improve the reliability, performance, and efficiency of Technology Operations. This role will be a technical contributor focused on leveraging AI Operations and observability platforms to support event correlation, anomaly detection, alert noise reduction, incident response automation, and operational analytics across the enterprise. The Engineer will partner closely with the Technology Fusion Center, Site Reliability Engineering (SRE), Incident Management, Change Enablement, and Service Management teams to implement intelligent operational capabilities and improve service reliability. This is a hands-on technical role requiring expertise in observability platforms, automation, monitoring technologies, and operational processes, while providing guidance and knowledge sharing across technical teams.

Requirements

  • Bachelor’s Degree in Computer Science or related field of study.
  • 1 cloud certification- AWS cloud practitioner, Microsoft Azure Fundamentals or Google Cloud- Cloud Digital Leader
  • 8-10 years of experience in Technology Operations, Site Reliability Engineering (SRE), Observability Engineering, Cloud Operations, AI Operations, or related technical disciplines.

Nice To Haves

  • Experience implementing Generative AI and Agentic AI solutions within highly regulated enterprise environments, with consideration for governance, risk, security, and compliance requirements.
  • Demonstrated experience building, deploying, and managing AI agents and agent-based workflows in production environments.
  • Experience with and strong understanding of ITIL frameworks, service management processes, and operational best practices.
  • Hands-on experience designing, implementing, and scaling Site Reliability Engineering (SRE) tools, practices, and operational processes.
  • Proven ability to assess existing business and technology processes, identify opportunities for operational efficiencies, and drive automation initiatives that improve scalability and effectiveness.

Responsibilities

  • Partner with Fusion Center, SRE, Incident Management, Change Enablement, and application teams to deploy and optimize AI Operations capabilities that improve monitoring effectiveness, reduce alert fatigue, and support incident response activities.
  • Configure, administer, and optimize AI Operations capabilities within observability and monitoring platforms (e.g., Dynatrace, BigPanda, ServiceNow, Port, or similar tools). Support integrations with monitoring, logging, event management, and ITSM platforms.
  • Identify, evaluate, and implement operational use cases that leverage AI/ML capabilities to improve root cause analysis, anomaly detection, predictive alerting, and operational efficiency.
  • Develop and maintain automated remediation workflows, runbook automation, and event-driven operational processes that reduce manual effort and improve service restoration times.
  • Provide technical guidance to peers, contribute to operational best practices, document solutions, and support adoption of AI Operations capabilities across technology teams.
  • Evaluate emerging observability and AI Operations capabilities, participate in proof-of-concept activities, and provide recommendations for platform enhancements and process improvements.

Benefits

  • medical
  • prescription drug
  • dental
  • vision
  • life insurance
  • short and long-term disability
  • 401(k)
  • paid holidays
  • Flexible Time Off (FTO)
  • Paid Sick and Safe Time (“PSST”)
  • company stock purchase
  • annual incentive bonus
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service