Sr. Staff Machine Learning Systems Engineer

Hims & Hers,
$240,000 - $265,000Remote

About The Position

Hims & Hers is seeking a Senior Staff engineer to own the development and implementation of trustworthy AI/ML systems within a regulated healthcare environment. This role involves managing data pipelines for model feeding and evaluation, as well as the infrastructure for evaluating model effectiveness and safety. The position requires setting technical direction for building, versioning, and trusting data and judgments used in AI product evaluation. The scope includes raw data ingestion, feature/dataset pipelines, evaluation methodology, statistical rigor, and reporting surfaces for clinical and product teams. The role is designed for individuals who thrive on ambiguous, cross-team challenges, bringing clarity, charting paths, and driving projects from inception to production with organizational buy-in.

Requirements

  • 10+ years of experience in ML infrastructure, data engineering, or evaluation/testing systems, with a track record of impact beyond a single team or project.
  • Hands-on depth in evaluation systems: designing and calibrating LLM judges/scorers, building statistically sound regression-testing methodology (e.g., paired significance testing with proper correction for multiple comparisons), measuring agreement against human labels, and designing adversarial/red-team evaluation approaches.
  • Hands-on depth in data pipeline engineering: dataset versioning, feature and benchmark pipelines, labeling and calibration workflows, and high-throughput ingestion and transformation systems.
  • A history of building things that became the standard approach for others.
  • Experience leading multi-team projects to completion, including navigating and resolving genuine technical disagreement.
  • A track record of mentoring other engineers, including senior ones, and visibly raising the bar for the teams around you.
  • Excellent communication skills, comfortable adapting ideas for different audiences and building support before launch.
  • Strong Python skills.
  • Sufficient statistical fluency to design and defend a testing framework for production decisions.

Nice To Haves

  • Experience with Databricks, MLflow, Unity Catalog, or similar data/eval platforms.
  • Experience building reporting tools for people without direct engineering access (e.g., automated Slack digests, spreadsheet reports for non-technical teams).
  • Prior experience in a regulated industry (healthcare, fintech, life sciences).
  • A track record of company-wide talks or write-ups that changed how other teams approached a problem.

Responsibilities

  • Own the evaluation process as a whole, not just a segment.
  • Set the technical direction for evaluation systems, including metric, judge, and scorer design, statistical methodology for regression decisions, and infrastructure for tracking failures.
  • Design and scale data pipelines for ingestion, transformation, dataset versioning, and labeling/calibration workflows to support evaluation and data science.
  • Proactively address challenges in scaling and complexity of AI evaluation.
  • Lead multi-team and multi-quarter projects.
  • Define evaluation processes for new AI services from scratch.
  • Drive initiatives to replace manual review processes with automated, statistically sound gates.
  • Own the approach to adversarial and red-team evaluation as a risk-reduction program, designing test suites and failure taxonomies.
  • Work through complex, cross-team technical disagreements and drive alignment among engineering, product, and AI leaders.
  • Transform ambiguous, cross-team problems into reusable solutions.
  • Originate new approaches and methodologies that become reusable standards.
  • Take vague, cross-team pain points from rough ideas to fully-specified, shipped systems.
  • Lead major platform improvements, including re-architecting core systems and modernizing operations.
  • Build working relationships across ML engineering, data science, platform engineering, clinical, legal, and product teams.
  • Become a go-to expert on evaluation methodology and data pipeline design, sharing knowledge through internal talks and documentation.
  • Mentor other engineers, including experienced ones, and raise the technical and statistical bar of teams.

Benefits

  • Competitive salary & equity compensation for full-time roles
  • Unlimited PTO
  • Company holidays
  • Quarterly mental health days
  • Comprehensive health benefits including medical, dental & vision
  • Parental leave
  • Employee Stock Purchase Program (ESPP)
  • 401k benefits with employer matching contribution
  • Offsite team retreats
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service