Research Engineer - Model Evaluation & MLOps

SciforiumSan Francisco, CA
$155,000 - $200,000

About The Position

As a Research Engineer focused on Model Evaluation & MLOps, you will build the tools and infrastructure needed to evaluate, deploy, and operate multimodal foundation models reliably. You will rapidly enable Sciforium’s models and the latest open-weight models on GPUs, automate quality and performance benchmarking, and improve the MLOps workflows that connect research experiments to reliable releases.

Requirements

  • 2+ years of professional ML or software engineering experience, including work on production ML systems, ML platforms, or MLOps infrastructure.
  • Strong Python and software engineering skills, with experience building reliable production systems.
  • Hands-on experience with PyTorch, TensorFlow, or JAX and a good understanding of modern language or multimodal model architectures.
  • Experience with model evaluation or benchmarking and core model lifecycle workflows such as experiment tracking, versioning, deployment, or monitoring.
  • Experience running, benchmarking, and debugging models with one or more GPU inference runtimes, such as vLLM, SGLang, TensorRT-LLM, or equivalent, in containerized cloud or on-premises environments.
  • Ability to document systems clearly and collaborate across research, infrastructure, and product engineering teams.
  • MS or PhD in Computer Science, Computer Engineering, Machine Learning, or a related technical field, or equivalent practical experience.

Nice To Haves

  • Familiarity with Hugging Face Transformers or similar model libraries.
  • Experience enabling models on AMD GPUs and ROCm.
  • Contributions to open-source evaluation, model, or ML infrastructure projects.

Responsibilities

  • Rapidly integrate new internal and open-weight language and multimodal models into our GPU evaluation and inference environments.
  • Build automated benchmarks for model quality and systems performance, including latency, throughput, and memory usage.
  • Create standardized, reproducible comparisons across Sciforium models, external baselines, and runtime configurations.
  • Build and maintain experiment tracking, model registry, and versioning for models, datasets, and evaluation configurations.
  • Automate the path from research checkpoints to validated deployments through CI/CD and reproducible workflows.
  • Monitor model quality and systems performance, and diagnose failures or regressions across model and deployment pipelines.
  • Build reusable tools that help researchers launch evaluations, compare experiments, and reproduce results.
  • Profile end-to-end model workloads and collaborate with distributed systems, inference, and GPU kernel engineers on deeper performance issues.

Benefits

  • Medical, dental, and vision insurance
  • 401k plan
  • Daily lunch, snacks, and beverages
  • Flexible time off
  • Competitive salary and equity
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service