MLOps Engineer

Fractal Analytics•New York, NY
•$100,000 - $125,000

About The Position

We are hiring a senior MLOps Engineer on a consulting basis to help operationalize a portfolio of machine learning solutions in purchase and underwriting. You will work alongside our AI/MLOps team and partner with our Data Science and Data Engineering teams to deliver inference services, data pipelines, feature engineering frameworks, and model lifecycle management capabilities. This is a hands-on engineering role. You should expect to spend most of your time writing production code, designing scalable services, building reusable ML platform capabilities, and ensuring machine learning solutions are production-ready across development, test, and production environments.

Requirements

  • Deep hands-on Python expertise for both data engineering and backend application development.
  • Strong experience developing production-grade backend services using FastAPI.
  • Hands-on expertise with the Databricks platform, including: MLflow, Delta Lake, Databricks Workflows, Databricks Asset Bundles (DABs), Spark-based distributed processing.
  • Proven experience designing and implementing end-to-end MLflow-based training and inference architectures across multiple environments.
  • Experience with AWS services including SQS, EKS, and Aurora PostgreSQL.
  • Experience implementing event-driven and asynchronous architectures using Kafka and/or SQS.
  • Strong understanding of the end-to-end ML lifecycle, including feature engineering, training, deployment, monitoring, and retraining.
  • Experience creating architecture diagrams, technical design documentation, implementation plans, and operational runbooks.
  • Strong Docker and Kubernetes fundamentals.
  • Experience with CI/CD pipelines using GitHub Actions, Jenkins, or similar tools.
  • Hands-on experience using both GitHub Copilot and Claude Code as part of day-to-day software engineering workflows.
  • Excellent written and verbal communication skills with the ability to collaborate effectively with technical and business stakeholders.

Nice To Haves

  • Prior experience working in Group Insurance, Life Insurance, or Underwriting domains.
  • Experience operationalizing GenAI or LLM-based applications and services.

Responsibilities

  • Design and build FastAPI services that expose models to downstream applications, including request/response contracts, authentication and authorization, input validation, error handling, structured logging, tracing, and metrics.
  • Implement queue-based asynchronous serving patterns for higher-latency and higher-throughput workloads using technologies such as SQS and Kafka.
  • Containerize services with Docker and deploy them onto Kubernetes/EKS environments with production-grade observability and scalability.
  • Design and implement model inference patterns across both batch and real-time serving use cases.
  • Design and implement end-to-end MLflow-based patterns for model training, experiment tracking, model registry, deployment, and inference.
  • Build and manage reproducible ML workflows, ensuring seamless model promotion across development, testing, and production environments.
  • Manage model artifacts, versions, lineage, and deployment governance using MLflow and Databricks-native capabilities.
  • Establish standardized MLOps frameworks and reusable implementation patterns that can be adopted across multiple ML use cases.
  • Build reproducible training and inference pipelines on Databricks and PySpark, from raw sources through curated feature and training datasets.
  • Design and develop pipelines leveraging Databricks capabilities including MLflow, Delta Lake, Databricks Workflows, and Databricks Asset Bundles (DABs).
  • Own data preprocessing, transformation, and feature engineering code and evolve existing solutions into scalable, reusable production frameworks.
  • Ensure consistent feature engineering and data processing logic across training, batch inference, and real-time serving paths.
  • Participate in solution design discussions with Data Science, Data Engineering, Platform Engineering, business, and IT stakeholders.
  • Create technical architecture diagrams, solution blueprints, design documents, and implementation runbooks.
  • Evaluate tradeoffs between scalability, maintainability, performance, and operational complexity while defining target-state architectures.
  • Provide technical guidance and recommendations on MLOps best practices, deployment approaches, and platform adoption.
  • Implement monitoring for model performance, prediction drift, data quality, service health, and pipeline reliability.
  • Diagnose production incidents across pipelines and services, identify root causes, and drive durable fixes.
  • Establish operational standards for deployment, monitoring, alerting, and supportability.
  • Apply strong software engineering fundamentals including testing, code reviews, CI/CD, versioning, dependency management, and infrastructure automation.
  • Build and maintain shared libraries, starter templates, reusable frameworks, and engineering standards.
  • Leverage GitHub Copilot and Claude Code to improve developer productivity, code quality, and delivery velocity.
  • Document solutions clearly enough for internal teams to own and operate after the engagement concludes.

Benefits

  • health insurance
  • dental insurance
  • vision insurance
  • life insurance
  • disability plans
  • 401(k) Plan
  • paid holidays
  • Parental Leave
  • free time PTO policy
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service