About The Position

The team operates at the cutting edge of AI, functioning within the broader Customer Experience Engineering organization as a specialized Customer Reliability Engineering (XRE) group to support Cisco’s customer experience (CX) platform. Our primary mission is to ensure seamless, highly reliable production systems by fiercely resolving complex customer reliability issues, managing critical escalations, and ensuring AI-driven solutions are secure and performant. We are a collaborative, agile group of MLOps experts, reliability engineers, and software developers who thrive on troubleshooting and stabilizing complex technical ecosystems. Working closely with design, data science, and product management, our efforts directly protect Cisco's CX product roadmap and ensure long-term customer success. What’s most exciting is the opportunity to shape the reliability of a rapidly evolving AI landscape, directly impacting the customer experience by tackling complex challenges in hybrid cloud environments, LLM orchestration, and inference optimization every day.

Requirements

  • Bachelors degree and 8+ years of related experience, or Masters and 6+ years of related experience.
  • Experience within Software Engineering and DevOps job families, specifically focusing on building scalable infrastructure.
  • Experience operationalizing Large Language Models such as OpenAI, Anthropic, Llama or similar and utilizing LLM frameworks such as LangChain or LangSmith.
  • Experience in Python and/or Java/J2EE.
  • Experience developing robust APIs using popular frameworks such as FastAPI, Flask, Spring Boot or similar.
  • Experience with CI/CD concepts for machine learning and familiarity with model tracking and MLOps frameworks such as ClearML, MLflow, Weights & Biases or similar and containerization utilizing Docker and Kubernetes.

Nice To Haves

  • Expertise in Agentic AI with hands-on experience in AI Agent development, including designing, orchestrating, and deploying autonomous systems and multi-agent workflows.
  • Broad perspective and ability to articulate architectural design choices regarding GPU compute scaling, model latency, and AI routing, including knowledge of AI API gateways and proxying solutions like Apache APISIX or LiteLLM.
  • Strong verbal and written communication skills to effectively negotiate delivery trade-offs (e.g., latency vs. model accuracy vs. cost) with cross-functional partners.
  • Experience contributing to threat modeling, specifically identifying and mitigating risks regarding prompt injection or AI data leakage.
  • Proven leadership ability to facilitate knowledge-sharing sessions, lead postmortems, and mentor junior team members on technical designs.

Responsibilities

  • Drive the operationalization of complex autonomous agent architectures into secure, scalable, and high-performing production environments.
  • Define and establish robust MLOps practices, foundational ML pipelines, and model serving architectures to bring innovative Agentic AI concepts out of the lab and into customer-facing products.
  • Design and implement robust telemetry monitoring tools to track token usage, control compute costs, monitor agent reasoning paths, and ensure optimal model performance and security.
  • Proactively resolve complex infrastructure challenges by managing containerized models on GPU clusters and addressing model inference issues.
  • Guide architectural choices based on deep market knowledge and mentor peers, ensuring our Agentic AI platforms remain reliable, scalable, and ahead of the curve.

Benefits

  • medical, dental and vision insurance
  • a 401(k) plan with a Cisco matching contribution
  • paid parental leave
  • short and long-term disability coverage
  • basic life insurance
  • grants of Cisco restricted stock units
  • 10 paid holidays per full calendar year, plus 1 floating holiday for non-exempt employees
  • 1 paid day off for employee’s birthday, paid year-end holiday shutdown, and 4 paid days off for personal wellness determined by Cisco
  • 16 days of paid vacation time per full calendar year, accrued at rate of 4.92 hours per pay period for full-time employees (non-exempt)
  • flexible vacation time off program (exempt)
  • 80 hours of sick time off provided on hire date and each January 1st thereafter
  • up to 80 hours of unused sick time carried forward from one calendar year to the next
  • Additional paid time away may be requested to deal with critical or emergency issues for family members
  • Optional 10 paid days per full calendar year to volunteer
  • annual bonuses
  • performance-based incentive pay
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service