About The Position

Join us in building the next generation of AI infrastructure that will power innovation across the customer organization. We’re seeking a senior software engineer to support our AI infrastructure team. In this role, you’ll help build and maintain the foundation for customer AI capabilities while supporting a broader ecosystem of AI-enabled applications. Your focus will be ensuring access to the highest available quality LLMs to users throughout the inference software stack. Requirements shift quickly as mission needs evolve and new technologies emerge, so you'll keep sharpening your skills and learning to turn loosely defined problems into working solutions. Lead the evaluation, configuration, and deployment of inference models for our users. Translate high-level stakeholder goals into engineering outcomes. Mentor engineers and grow development best practices and standards for the team. Develop in-house services and techniques to guarantee continual high-quality inference service for our customer. Engage with other teams in our organization to establish solid infrastructure for our services and integrate LLM-powered tools for user needs. Collaborate with teammates on surge efforts to support short-term, high-priority inference needs from our customer.

Requirements

  • Experience with Python and/or other modern programming languages.
  • Experience with Argo CD and/or other CI/CD frameworks.
  • Experience with Kubernetes/Helm.
  • Experience with AWS or other cloud service providers.
  • Strong communication skills and ability to mentor other engineers.
  • Excellent leadership and stakeholder management skills.
  • Ability to balance hands-on engineering with leadership and coordination responsibilities.
  • 12 years of experience, or 4 additional years in place of a B.S. in a technical discipline.

Nice To Haves

  • Experience with vLLM, LiteLLM, or similar inference-serving frameworks.
  • Experience with other LLM hosting frameworks and practices.
  • Experience supporting production software using Site Reliability Engineering (SRE) best practices.
  • Experience with Elastic, Grafana/Prometheus, or other observability frameworks and practices.
  • Experience with Docker and containerization.
  • Experience in traffic shaping and quality-of-service engineering.
  • Knowledge of and interest in hosting AI capabilities.

Responsibilities

  • Lead the evaluation, configuration, and deployment of inference models for our users.
  • Translate high-level stakeholder goals into engineering outcomes.
  • Mentor engineers and grow development best practices and standards for the team.
  • Develop in-house services and techniques to guarantee continual high-quality inference service for our customer.
  • Engage with other teams in our organization to establish solid infrastructure for our services and integrate LLM-powered tools for user needs.
  • Collaborate with teammates on surge efforts to support short-term, high-priority inference needs from our customer.

Benefits

  • 24 days PTO accrued annually
  • 11 federal holidays
  • 401k is 100% vested on your start date
  • Company makes a direct contribution worth 10% of your salary
  • Company covers 100% of healthcare costs for employees
  • Company covers 50% toward dependents' healthcare costs
  • Educational assistance towards college classes
  • Costs associated with job-related training and certifications are covered
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service