Cmptl & Data Sci Rsh Spec 4 Rp

University of California, San FranciscoSan Francisco, CA
Hybrid

About The Position

This role involves a hybrid of computational, data, and cyberinfrastructure (CI) fields including bioinformatics, geological information services (GIS), data analytics, and computational chemistry. The position applies computational, computer science, data science, and cyber infrastructure (CI) research and development principles, combined with relevant domain science knowledge, to conduct research and technology integration. Key responsibilities include the research, design, development, analysis, operation, and support of high-performance computing (HPC) and data science research, software, tools, and hardware resources. The role also involves developing data algorithms and performing computations, statistical analyses, and interpretation and reporting of research findings. This specialty is for positions primarily focused on research, utilizing computational and data science technology as a tool. The Preclinical Design and Clinical Translation of Regimens for Tuberculosis (PReDiCTR-TB) Consortium is a model-informed drug development (MIDD) platform that integrates computational science, translational pharmacology, and global clinical insight to accelerate the design and delivery of TB regimens. PReDiCTR-TB functions as a strategic intelligence engine, linking preclinical evidence, synthetic experiments, mechanistic models, and clinical data into a predictive framework. It utilizes quantitative systems pharmacology (QSP), AI-driven analytics, and probabilistic decision modeling to prioritize regimens with the highest probability of clinical success. The consortium employs a regimen-first, translation-driven approach, optimizing combinations, dosing strategies, and treatment durations through iterative simulation-validation cycles grounded in human-relevant biology. This framework aims to reduce reliance on empirical experimentation and increase translational fidelity. The consortium delivers predictive insights to de-risk development, optimized trial designs, and faster go/no-go decisions. The market for MIDD and AI in Clinical Trials is projected to grow significantly, driven by data and model-driven efficiencies in patient selection, drug repurposing, and real-time trial monitoring. The UCSF Savic Lab is establishing a modern, interconnected, and secure data & model infrastructure to support advanced analytics, multi-institution translational science, and accelerated drug development decision support. The Savic Lab is a global leader in model-informed drug development for infectious diseases, focusing on TB, HIV, malaria, pediatric infectious diseases, translational PK/PD, and systems pharmacology. Through its leadership in PReDiCTR-TB, the lab collaborates internationally to integrate computational science, mechanistic modeling, and clinical translation into actionable strategies for global health.

Requirements

  • 5+ years of hands-on experience in data engineering, distributed systems, or enterprise platform architecture.
  • A proven track record of architecting, deploying, and maintaining production-grade distributed data architectures.
  • Direct experience handling large-scale data ingestion, multi-tenant databases, and event-driven data flows.
  • Exceptional communication and presentation skills, with a demonstrated ability to explain complex technical concepts to non-technical executive stakeholders.
  • Bachelor's degree in Computer / Computational / Data Science, or Domain Sciences with computer / computational / data specialization or equivalent experience.

Nice To Haves

  • Data mesh environments, data fabric layers, and semantic web modeling.
  • Apache Kafka, Pulsar, AWS Kinesis, Apache Spark, Flink, or Apache Beam.
  • Distributed object storage (S3-compatible API environments) and NoSQL systems (Cassandra, DynamoDB, MongoDB).
  • Familiarity with NIH-preferred schemas and ontologies (e.g., BIDS for imaging, OMOP common data models, LOINC, SNOMED CT, or FHIR transfer protocols).
  • Data cataloging, automated data lineage tools, and strict data access control frameworks (HIPAA / NIST compliance).
  • Master's degree in Computer / Computational / Data Science, or Domain Sciences with computer / computational / data specialization preferred.

Responsibilities

  • Applies advanced HPC / data / CI research and development concepts to plan, design, develop, modify, debug, deploy and evaluate highly complex HPC (software and / or hardware) or data science or computational science or CI software and technologies or combination thereof.
  • Analyzes existing highly complex software, scientific codes, data science / analytics codes / algorithms, and HPC related hardware or works to formulate logic for new and highly complex systems and devises new algorithms.
  • Performs highly complex analysis and tests / debugs highly complex software and hardware.
  • Applies highly complex programming principles.
  • Initiates large and complex research projects with multi-institutional scope in HPC / data / CI areas. May involve collaboration with domain science experts.
  • Architect the Ecosystem: Design and implement distributed data platforms, data fabrics, and data mesh capabilities that safely unify complex datasets.
  • Specifies, develops, implements and executes highly complex software and hardware research and development plans.
  • Performs or directs highly complex HPC, computational and data modeling, performance and integration testing.
  • Works with research communities to develop, implement, and optimize computational and data analysis / analytics software / tools / algorithms / research codes with broad applicability.
  • Build Resilient Pipelines: Oversee the development of scalable data pipelines optimized for massive batch processing of pre-clinical and clinical data.
  • Initiates and contributes to HPC / data science / CI research proposals with partners internal and external to the institution, in collaboration with other researchers and PIs. May lead a proposal of small to moderate size as a PI.
  • Enable Cross-Domain Sharing: Create secure, cross-institution data frameworks that align with NIH Data Management and Sharing (DMS) policies and FAIR principles.
  • Drive Bio-Informatics Decisions: Set the architectural direction for formatting and structuring data that directly feeds pharmacometrics, biostatistics, and AI/ML pipelines.
  • Understands and applies advanced research and development practices, community standards and department policies and procedures.
  • May serve as technical lead for multiple research and development projects of moderate to broad scope.
  • Lead and Translate: Convert high-level clinical research data goals into concrete technical roadmaps.
  • Collaborate across software engineering, machine learning, cyber security, clinical pharmacology teams, and consortium participants.
  • Brief Stakeholders: Author technical documentation, lead architectural reviews, and deliver executive briefings to internal leadership and external consortium partners.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service