Software Engineer, ML Platform

nyra health
Hybrid

About The Position

As a Software Engineer on the ML Platform, you will build the systems behind every nyra labs experiment and model release. You will work across data processing, distributed training, experiment management, evaluation, inference, and release infrastructure. Your goal is to give a small research team the leverage to run ambitious experiments quickly, reproducibly, and reliably. This is not a conventional backend role. You will work directly with researchers, understand how models are developed, and turn recurring research bottlenecks into dependable platform capabilities.

Requirements

  • Excellent Python skills and experience building maintainable production systems.
  • Familiarity with PyTorch training, model evaluation, GPU workloads, and the ML development lifecycle.
  • Experience with cloud infrastructure, containers, orchestration, job scheduling, or distributed computing.
  • Experience building reliable pipelines and working with large, versioned datasets.
  • You care about observability, debuggability, failure recovery, and clear system boundaries.
  • You build reusable capabilities instead of solving the same problem repeatedly.
  • You understand that research workflows change quickly and infrastructure must support exploration.
  • You use coding agents, automation, and custom tooling to increase your own leverage and that of the team.

Nice To Haves

  • You know what needs a platform and what needs a small script.
  • You look for improvements that make the entire team faster.
  • You treat reproducibility and data integrity as core product requirements.
  • You can identify bottlenecks and own the solution end to end.
  • You enjoy working closely with researchers and translating experimental needs into durable systems.

Responsibilities

  • Build reliable pipelines for ingesting, validating, transforming, versioning, and accessing large speech datasets.
  • Improve distributed training, orchestration, checkpointing, resource scheduling, and failure recovery.
  • Create tooling for configuration, tracking, comparison, reproducibility, and artifact management.
  • Make it easy to run benchmarks, inspect regressions, compare releases, and understand model behavior.
  • Optimize models for efficient cloud and on-device use where relevant.
  • Automate model packaging, documentation, validation, and open-source publishing.
  • Build internal tools that remove friction from the daily work of researchers and engineers.
  • Establish observability, access controls, and operational practices appropriate for sensitive clinical data.

Benefits

  • Attractive compensation
  • Phantom Stock Options
  • company benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service