MLOps Engineer

Fractal AnalyticsNew York, NY
$120,000 - $140,000

About The Position

We are hiring a senior MLOps Engineer on a consulting basis to help operationalize a portfolio of machine learning solutions in purchase and underwriting. You will work alongside with our AI / ML Ops team and partner with our Data Science and Data Engineering teams to deliver the inference layer, data and feature pipelines. This is a hands-on engineering role. You should expect to spend most of your time writing production code — designing services, hardening pipelines, and to get ML production solutions developed, deployed, monitored, and consistent across batch and real-time paths.

Requirements

  • Deep hands-on Python in using it for both data engineering and application development, and comfort across SQL, PySpark, and shell scripting.
  • Production experience building services with FastAPI (or a comparable Python web framework), including auth, validation, error handling, and observability.
  • Experience building queue-based asynchronous processing systems — familiarity with at least one of Kafka, RabbitMQ, SQS, Redis Streams, Celery, or equivalent — and the operational concerns that come with them (retries, idempotency, back-pressure, dead-letter queues).
  • Strong Docker and general containerization skills; comfortable with Kubernetes concepts even if a platform team runs the cluster.
  • Hands-on Databricks experience including working knowledge of MLFlow and fluency with distributed compute in Spark.
  • Working experience with common ML libraries (scikit-learn, XGBoost, PyTorch or similar) — enough to be a competent partner to data scientists, not necessarily to build novel models.
  • Strong grasp of the end-to-end ML lifecycle and a track record of building or migrating feature engineering code with an explicit focus on training / batch / real-time parity.
  • Comfort reading and refactoring batch ML or data pipeline code — understanding intent and edge cases before rewriting.
  • CI/CD (Jenkins, GitHub Actions, or equivalent), version control workflows, and orchestration (Airflow, Prefect, or equivalent).
  • Excellent written and verbal communication; able to drive alignment with data scientists, platform engineers, and business stakeholders without a manager brokering every conversation.

Nice To Haves

  • Prior Experience working in Group Insurance Domain or Life Insurance Underwriting Domain.
  • Experience operationalizing LLM-based systems — inference serving, evaluation, cost and latency controls.

Responsibilities

  • Model serving — synchronous and asynchronous: Design and build FastAPI services that expose models to downstream applications, including request/response contracts, authentication and authorization, input validation, error semantics, and structured logging, tracing, and metrics. Implement queue-based asynchronous serving for higher-latency or higher-throughput workloads — producers and consumers, worker concurrency, retries and back-off, dead-letter handling, back-pressure, idempotency, and end-to-end traceability of a request across the pipeline. Containerize services with Docker and deploy them so that scaling, rollout, and rollback are boring.
  • Data preprocessing, feature engineering, and pipelines: Own the data preprocessing, transformation, and feature engineering code that sits between raw sources and the model — refactoring notebook or script-style logic into modular, tested, and reusable components. Work with existing code from prior batch solutions: read it carefully, understand the business logic and edge cases baked in, and evolve it into the target-state pipelines rather than throwing it away. Build reproducible training and batch inference pipelines on Databricks and PySpark, from raw sources through curated feature and training datasets. Manage model artifacts, versions, and promotion across environments so that what runs in production is always known and reproducible.
  • Data and feature parity across the ML lifecycle: Guarantee that the feature values a model sees at training time match what it sees at batch scoring and real-time serving — same definitions, same transformations, same edge-case handling. Design feature engineering code so that a single implementation (or a rigorously validated pair) serves both offline (Spark/batch) and online (low-latency Python) paths, avoiding the classic “training/serving skew” failure mode. Establish parity checks and reconciliation between training data, batch outputs, and real-time predictions as a first-class part of the pipeline — not an afterthought. Bring a solid working understanding of the end-to-end ML lifecycle — from data acquisition, preprocessing, and feature engineering through training, evaluation, deployment, monitoring, and retraining — and use that lens to make design trade-offs across batch and real-time solutions.
  • Reliability and observability: Implement monitoring for model performance, prediction drift, data quality, and pipeline health, with actionable alerts routed to the right owners. Diagnose production incidents in pipelines and services, identify root causes, and drive fixes through to closure — including the durable fix, not just the mitigation.
  • Engineering practices: Apply strong software engineering fundamentals — testing, code review, CI/CD, semantic versioning, and dependency hygiene — to ML code that has historically not had them. Build and maintain shared libraries, utilities, and repository patterns that other ML use cases can adopt. Document what you build clearly enough that internal teams can own it after the engagement ends.
  • Collaborate and document: Work closely with Data Science, Data Engineering, business partners, and IT teams to align on requirements, handoffs, and production readiness. Produce clear documentation of pipelines, frameworks, and operational runbooks so ownership can transition smoothly to internal teams.

Benefits

  • health insurance
  • dental insurance
  • vision insurance
  • life insurance
  • disability plan
  • 401k
  • paid holidays
  • Parental Leave
  • free time PTO policy
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service