LLM Ops Engineer

LiteraDenver, CO
$105,000 - $130,000Hybrid

About The Position

At Litera, AI is becoming a critical enabler of how we build products, improve customer experiences, and drive innovation. As an LLM Ops Engineer, you will create the secure, scalable, and reliable foundation that allows our engineering teams to leverage AI confidently and efficiently across the business. Your work will ensure that AI capabilities are available, governed, cost-effective, and ready to support production applications at scale. This role is instrumental in accelerating AI adoption while maintaining the performance, security, and resilience required for enterprise software.

Requirements

  • 3+ years of experience in DevOps, Platform Engineering, MLOps, or a related field, including hands-on experience operating LLMs in production environments.
  • Experience deploying, managing, and scaling models across multiple AI providers such as OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, or Google Vertex AI.
  • Strong expertise in building highly available, secure infrastructure, including load balancing, failover strategies, secrets management, and access controls.
  • Experience with API management, gateway technologies, and production-grade AI service operations.
  • Strong Python programming skills with experience developing and supporting scalable systems.
  • Proven ability to solve complex technical challenges and thrive in a fast-paced, evolving environment while collaborating across teams.

Nice To Haves

  • Experience fine-tuning or training large language models for domain-specific applications.
  • Familiarity with ML orchestration tools and frameworks such as Kubeflow, MLflow, or Apache Airflow.
  • Experience with LLM evaluation frameworks, retrieval-augmented generation (RAG), vector databases, or inference optimization techniques.
  • Knowledge of infrastructure-as-code, Kubernetes, compliance frameworks, or large-scale AI cost optimization strategies.

Responsibilities

  • Build and operate a scalable AI platform that enables engineering teams to seamlessly access and deploy models across multiple providers and environments.
  • Ensure high availability and resiliency of AI services through intelligent routing, failover strategies, and production-grade infrastructure.
  • Establish secure and compliant AI operations by protecting model access, safeguarding sensitive data, and enforcing governance standards.
  • Create a consistent developer experience through unified APIs, self-service capabilities, tooling, and best practices that accelerate AI adoption.
  • Optimize AI platform performance, reliability, and cost efficiency through proactive monitoring, analytics, and provider strategy management.
  • Lead the evolution of Litera’s AI operations capabilities by evaluating emerging technologies and recommending scalable solutions.
  • Deliver observability and operational excellence through dashboards, alerting, quality monitoring, and service-level metrics.
  • Support the safe deployment of AI solutions by implementing testing frameworks, quality controls, and production readiness standards.

Benefits

  • medical coverage
  • dental coverage
  • vision coverage
  • 401(k) with company match
  • incentive and recognition programs
  • company bonus plan
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service