Machine Learning Ops Engineer

CareforthUNAVAILABLE, Minnesota
Remote

About The Position

The ML Ops Engineer is a critical specialist within the Product & Technology Organization, responsible for the intersection of machine learning, software engineering, and platform operations. You will design and maintain the infrastructure required to scale ML models across clinical risk intelligence, caregiver risk scoring, composite risk trajectory, NLP signal capture, LLM-powered enablement tools, and insights reporting — from research through reliable production. The ideal candidate has deep experience with distributed systems, containerization, model lifecycle governance, and automated ML pipelines, and thrives collaborating with Data Scientists and Data Engineers to ensure models are deployable, monitorable, HIPAA-compliant, and continuously improving.

Requirements

  • Bachelor's or Master's Degree in Computer Science, Software Engineering, or a related technical field.
  • 7+ years of professional experience in DevOps, Data Engineering, or ML Engineering, with at least 4 years focused on Machine Learning operations.
  • Proven track record of owning and operating production ML systems including model serving, monitoring, and lifecycle governance.
  • Experience in healthcare or regulated data environments; familiarity with HIPAA technical safeguards required.
  • Experience mentoring engineers and contributing to platform standards and technical roadmaps.
  • Expert-level containerization and orchestration: Docker, Kubernetes/EKS; Infrastructure-as-Code with Terraform or AWS CloudFormation.
  • Deep experience with ML lifecycle tooling: MLflow, Databricks (Delta Lake, Unity Catalog, Workflows, Spark, Vector Search), and AWS SageMaker.
  • Strong AWS proficiency: S3, Lambda, API Gateway, SageMaker, Bedrock, Redshift, Athena, DynamoDB, Glue, Step Functions, Kinesis, Transcribe, Comprehend Medical, Secrets Manager.
  • Experience with stream processing (Kafka/Kinesis), event-driven pipeline design, and feature store architecture.
  • Familiarity with LLM API integration, prompt versioning, RAG infrastructure, and LLM quality and cost governance.
  • Strong Python proficiency; Bash scripting; Go or Java a plus.
  • Clear communicator able to translate operational constraints into actionable guidance for Data Scientists and product teams.
  • Collaborative, detail-oriented, and committed to reproducibility, HIPAA audit-readiness, and operational excellence.
  • Self-directed and intellectually curious; proactive in evaluating and adopting emerging MLOps tooling.

Responsibilities

  • Design and implement automated ML pipelines for model training, evaluation, and deployment using MLflow, Databricks Workflows, and AWS SageMaker Pipelines; own model registry governance including versioning, promotion, and retirement.
  • Build and manage scalable model serving infrastructure (REST APIs, WebSocket APIs) using Docker and Kubernetes/EKS for real-time and batch scoring across all risk and enablement model domains.
  • Architect and operate a sub-model orchestration layer, score computation service, score history, and audit logging compliant with HIPAA requirements.
  • Design and maintain feature store architecture, temporal feature computation, and data versioning to ensure training-serving consistency across all ML domains.
  • Implement and govern LLM API integrations (Claude/GPT via Bedrock) including prompt versioning, rate limiting, cost tracking, response logging, and output guardrails.
  • Build production monitoring and alerting for data drift, model decay, scoring latency, and pipeline failures; implement A/B testing infrastructure for controlled model rollouts.
  • Automate ML infrastructure provisioning using Terraform or AWS CloudFormation; manage secrets, access controls, PHI redaction, and HIPAA-compliant data handling across all services.
  • Help define and lead the enterprise MLOps platform strategy, drive adoption of core tooling, and mentor junior engineers on operational standards and best practices.
  • Establish CI/CD and Continuous Training (CT) workflows to enable rapid, safe, and auditable ML iteration across all project domains.
  • Perform other duties and special projects as assigned.

Benefits

  • Flexible schedules
  • Remote-first culture
  • Nationally recognized wellness program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service