Lead SRE

BMC Software
•$152,925 - $254,875•Remote

About The Position

This role focuses on building, operating, and continuously improving reliable, scalable, secure, and observable AI/ML and Generative AI platforms in production. The responsibilities span the full lifecycle from deployment and monitoring through incident management, optimization, and continuous improvement. At this level, the individual will execute defined work with guidance and grow through review and mentorship, contributing reliable work to a project. Scope includes completing defined operations/observability tasks with guidance, learning AI failure modes, and on-call basics.

Requirements

  • SRE fundamentals — SLOs/SLIs, monitoring, alerting, incident management (depth scales with level).
  • Linux, networking, distributed systems, troubleshooting, performance analysis.
  • CI/CD and DevSecOps practices; containers/orchestration (Docker, Kubernetes, OpenShift); cloud platforms.
  • Generative AI/LLM literacy; RAG/agents/tooling concepts as level requires.
  • MLOps/LLMOps tooling and AI observability as level requires.
  • Scripting/automation (Python, Bash); security-first, calm incident handling.
  • Linux/scripting fundamentals; exposure to monitoring or cloud (project/internship OK).
  • Curiosity about how LLM/agent systems fail in production.
  • Calm, careful work under guidance; readiness for eventual on-call with support.
  • Past experience: Contributed to monitoring, runbooks, or operational fixes within an existing workflow; basic exposure to cloud/containers.
  • Delivery evidence: Reliable execution of defined ops tasks; clear notes; careful changes under guidance.
  • Shared expectation: Completes defined tasks with guidance and checks work carefully.

Nice To Haves

  • Agent frameworks and AgentOps for multi-agent systems.
  • LLM cost-optimization and inference-serving stacks (vLLM, Ollama).
  • Enterprise on-prem or hybrid deployments; Agile/Atlassian.

Responsibilities

  • Build and operate production-grade infrastructure and operational frameworks for LLM, GenAI, RAG, and Agentic AI applications.
  • Define and defend SLOs/SLIs for agent and model behaviour — latency, availability, quality, and cost.
  • Monitor production AI systems for performance, drift, hallucination rates, and quality regressions; act before customers are affected.
  • Lead incident response for AI systems: detection, triage, mitigation, and blameless post-mortems.
  • Build runbooks, automation, and self-healing for common AI operational issues; participate in on-call as required.
  • Operate model/agent lifecycle in production: versioning, rollout, rollback, and safe promotion with evaluation gates.
  • Manage LLM cost, capacity, and scaling; manage model serving/inference infrastructure.
  • Maintain observability/tracing for agents, tools, and prompts (Langfuse, OpenTelemetry, OpenSearch).
  • Operate AI/ML workloads on IBM Z where required; harden the AMI Platform for reliability, scalability, and security.
  • Partner with AI Engineers, Data Scientists, DevOps, and Security to embed reliability from design onward.

Benefits

  • Variable plan
  • Country specific benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service