Data Scientist

Rice University•United States,
•$105,000•Remote

About The Position

The Rice University Department of Computer Science and Department of Materials Science and NanoEngineering are seeking a Data Scientist to build and operate the data infrastructure for READINESS, a new $20 million NSF-funded project transforming four materials synthesis systems into an AI-driven, remotely accessible autonomous laboratory. The READINESS facility will generate substantial and diverse data, including growth recipes, reactor time-series, user interactions, optical, SEM/TEM and AFM images, spectra, and continuous robot telemetry. The Data Scientist will be responsible for developing and maintaining the central data lake that captures, indexes, secures, and serves this information across the project. The position combines data engineering and analysis, including deploying an S3-compatible object store, developing ingestion connectors for live instruments and robot controllers, and establishing SQL and vector-search capabilities under Rice single sign-on. The role will also compute embeddings over images and spectra, develop retrieval and similarity-search services, and build AI agents capable of answering researchers' questions directly from the data. Reporting to the faculty leads of the READINESS Data Infrastructure working group, the Data Scientist will collaborate closely with the synthesis platform teams, robotics and AI/digital-twin working groups, Rice IT and Information Security, and graduate and undergraduate students involved in the project. This is a hands-on computer science and data engineering position that requires substantial ownership of the project's central data infrastructure. A background in materials science, chemistry, or laboratory science is not required; domain-specific knowledge can be developed through collaboration with the project's scientific experts. Working within a small team, the Data Scientist will develop the foundational systems through which data generated across the READINESS facility is captured, organized, and made accessible.

Requirements

  • Bachelor’s degree in computer science, data science, engineering, or a related quantitative field.
  • Three or more (3+) years of related professional experience in data engineering, data science, software engineering, or research computing.
  • Strong programming skills and ability to write production-quality, tested, documented code.
  • Proficiency with SQL and relational data modeling, including scientific or operational database schemas.
  • Experience with object storage (S3, Ceph, or MinIO), columnar formats (Parquet), and distributed query engines (Trino, Presto, Spark, or similar).
  • Working knowledge of Linux administration, containers (Docker/Kubernetes), and at least one major cloud platform (AWS or Azure).
  • Familiarity with PyTorch, embedding models, vector databases, and retrieval-augmented or agentic LLM applications.
  • Understanding of authentication and access controls (SSO/SAML/OIDC, role-based access, audit logging) and secure research-data handling.
  • Ability to connect software to physical instruments and devices using network or file-based protocols; ROS/rosbag experience is a plus.
  • Ability to learn an unfamiliar scientific domain, translate collaborators’ requirements into working systems, manage competing requests, and communicate clearly.

Nice To Haves

  • Five or more years of professional experience building and operating data systems.
  • Experience building data pipelines for scientific instruments, laboratory automation, manufacturing, or IoT/telemetry.
  • Experience operating Ceph or comparable software-defined storage, or managing cloud object-storage tenancies at 100 TB+ scale.
  • Experience with GPU-accelerated similarity search (FAISS, Milvus, Qdrant) and serving open-weight LLMs.
  • Experience in academic research or national laboratories, including working with institutional IT and security offices.
  • Experience monitoring production data systems and implementing backup and recovery.

Responsibilities

  • Designs, deploys, and operates the READINESS central data lake, including the S3-compatible object store (self-hosted Ceph or cloud), bucket layout, versioning, and access policies.
  • Develops and maintains ingestion connectors that capture data unattended from synthesis tools, characterization instruments, and robotic systems, and land it in the store with structured, de-identified metadata and provenance links.
  • Designs and maintains the project's metadata and provenance schema and the Python client library that all project teams use to read and write data.
  • Deploys and administers the SQL query layer (Trino and Hive Metastore over Parquet) and programmatic APIs for researchers, the user portal, and digital-twin and AI teams.
  • Builds embedding pipelines and GPU-resident vector indexes over images, spectra, and logs, and exposes similarity-search and retrieval-augmented-generation services.
  • Develops and demonstrates AI agents that answer questions about experiments by retrieving from the data lake, in collaboration with the AI & Digital Twins working group.
  • Implements and maintains security controls, including Rice single sign-on with MFA, tiered role-based access, encryption, and immutable audit logging, and works with Rice Information Security and Research Security on design reviews and compliance.
  • Monitors system health, performs integrity checks and backups, and documents and tests recovery procedures.
  • Gathers requirements from synthesis platform, characterization, and robotics teams; attends platform team and working-group meetings; and sets node-wide standards for data formats, metadata, and access.
  • Writes documentation, user guides, and onboarding materials, and trains researchers and partner-site collaborators to use the data infrastructure.
  • Mentors graduate and undergraduate students working on data infrastructure projects.
  • Performs all other related duties as assigned

Benefits

  • Benefits-eligible position
  • Central standard time working hours
  • Fully remote work arrangement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service