AI Predictive Biology Postdoctoral Fellow (KBase Project)

Lawrence Berkeley National LaboratoryBerkeley, CA
$99,192 - $110,808Onsite

About The Position

Berkeley Lab’s (LBNL) Environmental Genomics and Systems Biology (EGSB) Division has an opening for a Postdoctoral Fellow to join the US Department of Energy’s (DOE) Systems Biology Knowledgebase (KBase) team! In this exciting role, you will leverage the extensive functional data available through KBase, including genome-wide fitness data, curated phenotypes, metabolic reconstructions, and analysis results integrated with their associated genomes and accumulated over more than a decade of research. Genome and protein foundation models are advancing quickly, yet almost none of that progress has been translated into reliable prediction of biological function. Current measures such as model perplexity and structure recovery do not directly assess functional prediction. A key challenge is the absence of large-scale, systematically generated functional measurements needed to adapt, evaluate, and validate these models. You will have the opportunity to work at the intersection of AI, machine learning, and computational biology to adapt and augment existing open models for predictive biology. Approaches may include fine-tuning, parameter-efficient tuning, probing, and retrieval- or knowledge-conditioned inference.

Requirements

  • A recent Ph.D. (within the last 1-2 years) in Computational Biology, Bioinformatics, Machine Learning, Computer Science, Statistics, or a related field.
  • A strong background in artificial intelligence, machine learning, computational biology, and/or large-scale biological data and measurements to develop and evaluate innovative approaches for predictive biology.
  • Demonstrated experience adapting large pretrained deep learning models to downstream scientific applications, including fine-tuning, parameter-efficient methods, model probing, or retrieval augmentation.
  • Experience applying deep learning methods to biological sequence or other omics data.
  • Demonstrated rigor in model evaluation and experimental design, including selection of appropriate baselines, defensible splits, and appropriate statistical analysis and interpretation of results.
  • Strong programming skills in Python and proficiency with PyTorch or an equivalent framework.
  • Strong organizational skills including experience maintaining detailed and accurate records of results and analyzed data.
  • Excellent verbal and presentation skills including experience preparing research reports, manuscripts, and scientific publications for group meetings, conferences, and scientific journals.
  • Demonstrated interpersonal communication skills including experience conducting independent, data-driven research and collaborating with an interdisciplinary research team.
  • You must have less than 3 years of paid postdoctoral experience.

Nice To Haves

  • Direct experience working with protein or genomic language models.
  • Experience with genome-wide fitness data (e.g., RB-TnSeq), functional genomics, or comparative genomics at scale.
  • Experience with model calibration, uncertainty quantification, or active learning.
  • Experience with genome-scale metabolic modeling or pathway analysis.
  • Demonstrated experience with training and calibrating other large-scale complex models.
  • Demonstrated ability to pretrain large-scale models from scratch, including distributed multi-GPU training.
  • Demonstrated experience developing and openly releasing benchmarks, datasets, models, or leaderboards.

Responsibilities

  • Adapt and augment open genome and protein foundation models for biological function prediction through fine-tuning, parameter-efficient tuning, model probing, and retrieval- or knowledge-conditioned inference.
  • Augment models with KBase functional signal - including genome-wide fitness measurements, curated phenotypes, metabolic reconstructions, and other mechanistic information as training, conditioning, or retrieval context.
  • Develop and release benchmarks that tie model performance to real biological tasks, including gene function, fitness, phenotype, and pathway completion, with defined splits, baselines, and evaluation protocols.
  • Evaluate model calibration and uncertainty and characterize failure modes against taxonomic distance, annotation quality, and data sparsity.
  • Deploy validated models and adapters as documented KBase capabilities that can be used directly by researchers, and incorporate results and predictions into ongoing model evaluation data.
  • Publish research findings, methods, benchmarks, and datasets, and openly release evaluation code.

Benefits

  • Exceptional health benefits.
  • Generous paid time off, sick time off, and holidays.

Stand Out From the Crowd

Upload your resume and get instant feedback on how well it matches this job.

Upload and Match Resume

What This Job Offers

Job Type

Full-time

Career Level

Entry Level

Education Level

Ph.D. or professional degree

© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service