Member of the Technical Staff - Systems ML Engineer

TransfyrCambridge, MA
Onsite

About The Position

Transfyr is building physical AI for science. We are developing systems that capture real scientific work and turn it into a high-fidelity, machine-readable record. This infrastructure helps teams learn from failures, transfer know-how, train scientists, and provide grounded data for models and robots. We are tackling complex problems at the intersection of science, perception, machine learning, and robotics. We have significant traction, are backed by a $25M seed round, and collaborate with leading AI labs. Our advisors include prominent figures in AI and science. We are looking for ambitious and pragmatic individuals to join our team.

Requirements

  • Demonstrated expertise in ML systems engineering, including optimizing and deploying large-scale models in production.
  • Experience debugging and fixing performance and stability issues in deployed systems.
  • Experience building infrastructure for reproducible, monitored ML deployments.
  • Experience optimizing inference throughput and resource utilization across cloud and edge.
  • Deep knowledge of distributed training and serving frameworks (e.g., PyTorch/JAX distributed strategies, gradient accumulation, mixed precision training, checkpoint/recovery systems).
  • Strong cloud administration skills (e.g., AWS services, infrastructure as code like Terraform, Kubernetes orchestration, cost optimization, security best practices, compliance requirements).
  • Experience deploying and optimizing ML systems on edge or on-prem infrastructure.
  • Understanding of the ML stack from hardware to deployment and serving.
  • Skilled at debugging complex failures across the stack (GPU/NCCL issues, data loading bottlenecks, memory leaks, performance or convergence problems).
  • Deep experience optimizing algorithms for cloud and edge environments, including computer vision and other ML algorithms.
  • GPU-level work like CUDA and kernel tuning.
  • High agency: Identify needs, build scaffolding, and push work forward.
  • Biased toward action: Prototype quickly, test assumptions, and iterate based on failure.
  • Successful in ambiguity: Make progress with incomplete labels, delayed feedback, and evolving success criteria.
  • Thoughtful: Make deliberate tradeoffs between model complexity, robustness, and operational cost.
  • Clear, direct communicator.
  • Intense: Care deeply about the mission and work hard.

Nice To Haves

  • Experience writing custom GPU kernels when off-the-shelf ops aren't fast enough.
  • Experience managing cloud and edge infrastructure that holds up in real lab environments.
  • A passion for and experience in science.
  • A passion for and experience with AI.
  • Demonstrated experience working in fast-moving/ambiguous environments (like startups!).

Responsibilities

  • Profile and optimize training and inference workloads.
  • Manage cloud and edge infrastructure.
  • Optimize cloud spend.
  • Ensure security and compliance.
  • Work closely with the perception and research teams to optimize model performance.
  • Profile and optimize performance using tools like Nsight and PyTorch Profiler.
  • Implement optimizations such as kernel fusion, sharding, and tiling.
  • Improve the efficiency of distributed training pipelines using PyTorch Distributed.
  • Design and maintain high-performance GPU kernels in Triton or CUDA.
  • Design and optimize data loading and inference pipelines.
  • Manage deployment across cloud infrastructure and edge devices in lab environments.
  • Debug and resolve performance bottlenecks, resource issues, and failures.
  • Partner with the research team on training efficiency.
  • Collaborate with perception engineers on data pipelines.
  • Build monitoring, versioning, and rollback into deployments.

Benefits

  • Competitive compensation (cash + equity)
  • Full benefits (low/no-cost health insurance options, HSA, 401K with matching, lunch subsidy, etc.)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service