Senior ML Ops Engineer | $165K-$175K + Hybrid + Equity | AI Powered Outage Intelligence SaaS Startup

PhillyTech.Co•King of Prussia, PA
•$165,000 - $175,000•Hybrid

About The Position

This is a 3-day in-office hybrid role in King of Prussia, PA. The company is a fast-growing B2B SaaS outage intelligence company helping major enterprises reduce downtime through advanced automation and real-time intelligence. Their technology helps organizations understand outages faster, automate operational workflows, reduce unnecessary costs, and accelerate repair times. Their platform supports critical infrastructure operations where reliability matters. As the company expands its customer base and develops new products, they are investing further in the machine learning systems behind their intelligence platform. This is not a role where you inherit a finished ML platform and simply maintain it. They are looking for someone who has previously helped build an ML stack from the ground up, understands what production ML infrastructure looks like as it scales, and wants meaningful ownership over the systems supporting machine learning in production.

Requirements

  • 4+ years of professional software engineering or data engineering experience.
  • 2+ years building and operating machine learning systems in production.
  • Hands on experience building an ML platform from the ground up or significantly scaling an existing ML platform.
  • Experience continuing to own and operate ML infrastructure as it matured.
  • Experience supporting multiple production models with complex training workloads.
  • Experience building and maintaining production data pipelines at scale, including managing their operating costs.
  • Experience working with multiple external data sources that behave differently and may arrive inconsistently.
  • Working knowledge of machine learning modeling and evaluation, with enough depth to review and challenge the work of ML engineers.
  • Entrepreneurial mindset and interest in working within a fast moving startup environment.

Nice To Haves

  • Advanced proficiency with Python 3 in mature production environments + Kubernetes + PostgreSQL.
  • Experiment tracking, model registry, and model serving frameworks.
  • Workflow orchestration.
  • Infrastructure as code tooling.
  • Geospatial or time series systems.
  • Large real time data flows.
  • Experience working for a data intelligence company.
  • Bachelor's degree in Computer Science or a related field.

Responsibilities

  • Own and extend the ML platform end to end, including training orchestration, experiment tracking, model registry, deployment, and production monitoring.
  • Build and operate data pipelines supporting model training and online inference.
  • Design processes for backfills, replays, and recovery when upstream data feeds fail.
  • Ensure training data accurately reflects what was known at the point in time it represents.
  • Build and manage reliable processes for moving trained models into production.
  • Ensure model training runs and results are reproducible.
  • Monitor deployed model performance over time, including models where outcomes are confirmed later.
  • Manage training and inference costs as data volume and the number of production models grow.
  • Contribute to technical architecture design and reviews.

Benefits

  • Hybrid work model, onsite in King of Prussia 3 days per week
  • Equity in a fast-scaling SaaS company
  • Fully paid medical, dental, and vision options
  • Life and AD&D insurance
  • PTO
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service