Data Engineer

Probably GeneticSan Francisco, CA
Hybrid

About The Position

Probably Genetic is changing the lives of patients living with severe, complex diseases. Our data platform is used by drug developers and patient advocacy groups to develop and launch treatments for these patients. Our technology discovers undiagnosed patients online, analyzes their disease state using machine learning and at-home testing, and enables compliant communication with patients. In doing so, we help patients access diagnoses, clinical trials, and treatments as early as possible. We are a tight-knit group of hard-working, ambitious problem solvers united by a mission greater than ourselves. We do well by doing right by patients. Our annually recurring revenue is growing >6x year over year, we’re profitable, and our roadmap is packed with innovations in bioinformatics, machine learning, and drug development. We are building an all-star team to help us bring our vision to life, and we want you to be a part of it. Probably Genetic has raised multiple rounds of funding from Silicon Valley’s best investors, including Threshold, Khosla, and Y Combinator, giving us the ability to pay competitive salaries, offer great benefits, and provide meaningful equity. We’re dedicated to ensuring your journey with us is unforgettable, with incredible team retreats to places like Barbados, the Alps, Mexico, Costa Rica, and Portugal, just to name a few. About the role We are looking for a Data Engineer to build data pipelines and visualizations for key business metrics, machine learning (ML) models, and growth metrics to enable our data-driven organization.

Requirements

  • 5+ years of relevant experience in data engineering
  • Experience using Python and SQL for ETL pipelines and data analysis
  • Experience in data reporting and dashboarding using variety of BI tools
  • Track-record of solving extremely difficult, ambiguous problems
  • A good person. We work with some of the most marginalized populations on the planet and empathy is key
  • Patient-focused and motivated to have a lasting, positive impact on humanity
  • Comfortable in a fast-paced, often ambiguous environment with rapid change
  • Action-oriented and excited to build a company from the ground up

Nice To Haves

  • Experience in contributing to a Python Django backend
  • Expertise with healthcare data, bioinformatics, ML infrastructure, or growth marketing

Responsibilities

  • Collaborating with stakeholders in Business Development, Growth, Product, ML and Finance to design data visualizations and pipelines
  • Developing ETL pipelines from a PostreSQL database using Python Django and AWS services such as S3, Glue, and Athena
  • Integrating with external data sources such as Meta and Google Ad platforms, product analytics platforms, and financial platforms to extract valuable data and insights
  • Developing visualizations in a business intelligence (BI) tool of your choice to create reports and dashboards
  • Collaborating directly with the CEO and executive team on investor updates using data from BI reports and dashboards
  • Identifying missing data points which should be collected and working with the engineering team to collect all valuable data
  • Writing documentation to onboard team members to BI tooling and ensuring that datasets are available for ad-hoc analysis

Benefits

  • 30 days of vacation a year
  • Hybrid, flexible work
  • A “work from anywhere” policy, up to 4 weeks a year
  • Competitive equity grants
  • All-expenses paid quarterly team retreats
  • Benefits including medical and dental
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service