Scientific Data Engineer

Lawrence Berkeley National LaboratoryBerkeley, CA
$117,132 - $197,676Hybrid

About The Position

Lawrence Berkeley National Laboratory is hiring a Scientific Data Engineer within the Scientific Data Division. The Computational Biosciences Group has an immediate opening for a software and data engineer in the area of multi-modal data modeling and analysis with applications to bioscience research. You will develop new methods and software tools that enable scientific knowledge discovery using modern data management and machine learning technologies and advance the state-of-the-art in data-intensive analysis. Your projects will focus on the domains of omics/structural biology data and neurophysiology data. Under limited instruction, you will be part of an experienced team conducting R&D in the areas of FAIR data science, AI, and modern methods for data understanding. You will be working as part of a multi-disciplinary team composed of computer scientists, data scientists, and bioinformaticians. Please note this is a scientific software/data engineering position– it is not a pure machine learning or AI research position, and it is not a pure data science or analytics position. Applications that do not demonstrate experience developing scientific software, scientific data models, or scientific data pipelines will not be competitive.

Requirements

  • Typically requires a minimum of 5 years of related experience with a Bachelor’s degree in computer science, data science, machine learning, bioinformatics, or equivalent; or 3 years and a Master’s degree; or equivalent work experience designing and developing software for data modeling or analysis; or a PhD in a relevant STEM field
  • Demonstrated experience developing software in a scientific or research context, such as in a research group, a scientific user facility, or on a scientific software project
  • Strong programming experience in Python.
  • Experience testing large code bases
  • Experience contributing to community-driven open source software
  • Demonstrated experience in one or more of the following areas: data management, scientific data analysis, machine learning
  • Works well in a collaborative team environment
  • Demonstrated capability with the Git version control and continuous integration systems, such as GitHub or GitLab
  • Ability to work effectively with domain scientists whose expertise is outside computing, and to translate their requirements into technical designs.
  • Excellent oral and written communication skills.
  • Demonstrated ability to work effectively as part of a cross-disciplinary team.

Nice To Haves

  • Working proficiency in C++ or Javascript is a plus.
  • Master's or PhD in Computer Science or related field, with 5 or more years of professional experience designing and developing scientific data modeling or analysis software
  • Experience working with modern scientific data formats and database systems, such as HDF5, Zarr, MongoDB, PostgreSQL, MySQL, and Redis
  • Experience with Neurodata Without Borders, LinkML, or similar software ecosystems
  • Experience working with large biological data, such as in the areas of neurophysiology, microbiology, genomics, or protein design
  • Experience designing or working with structured data models, schemas, ontologies, or data standards
  • Familiarity with FAIR data principles, persistent identifiers, provenance, and controlled vocabularies and ontologies
  • Experience preparing scientific datasets for use by machine learning pipelines or LLM-based agents
  • Experience working with cloud object storage, cloud computing, High-Performance Computing, data lakehouse architecture, or containerization.
  • Experience developing web-based graphical user interfaces (GUIs) or application programming interfaces (APIs) for scientific data analysis and management

Responsibilities

  • Design and develop user-friendly software packages for scientific data management and analysis
  • Work with domain experts to develop FAIR data models (i.e., models of the structure organization of the data) and management solutions for bioscience applications
  • Support machine learning and AI use of biological data by making it well-structured, documented, and efficiently accessible
  • Work closely with the community of developers of the Neurodata Without Borders and LinkML open source data ecosystems, as well as the Joint Genome Institute.
  • Maintain and manage open source software products, including managing development priorities, software releases, continuous integration, and testing
  • Design, implement and maintain high performance computing and cloud solutions for visualization and analysis of complex biological data
  • Develop machine learning and AI solutions for analysis of biological data in close collaboration with diverse teams of scientists
  • Train scientists and research software engineers in the use of the developed software products at workshops and conferences
  • Demonstrate good judgment in selecting methods and techniques for obtaining solutions.
  • Network with senior internal and external personnel in their own area of expertise.

Benefits

  • Exceptional health and retirement benefits, including pension or 401K-style plans
  • A culture where you’ll belong - we are invested in our teams!
  • In addition to accruing vacation and sick time, we also have a Winter Holiday Shutdown every year.
  • Parental bonding leave (for both mothers and fathers)
  • Pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service