Senior Applied Scientist

OracleSeattle, WA
$114,600 - $234,600

About The Position

The OCI AI Evaluation Science team builds the evidence behind model-selection, product-readiness, and launch decisions. We evaluate frontier foundation models and AI systems across capabilities such as reasoning, coding and agentic coding, retrieval-augmented generation, AI agents, NL2SQL, multimodal understanding, multilingual performance, and responsible AI. As a Senior Applied Scientist on the team, you will independently own complex evaluation work from problem definition, benchmark development, and final recommendation to executive leadership. You will translate ambiguous product and customer questions into measurable hypotheses, select or create appropriate benchmarks, design experiments, build evaluation pipelines, validate data and metrics, analyze failure modes, and communicate conclusions to science, engineering, product, and leadership stakeholders. This is hands-on applied science. You will write high-quality code, work with large and imperfect datasets, develop and calibrate automated evaluators, and turn one-off analyses into reproducible evaluation protocols and reusable infrastructure. You will examine more than aggregate benchmark scores, considering factors such as statistical validity, data provenance, contamination, robustness, cost, latency, reliability, safety, and operational constraints. The work sits at the point where research results become product decisions. Success requires scientific rigor, strong engineering judgment, clear writing, and the ability to make progress when requirements, model access, data, or infrastructure are still evolving. You will collaborate closely with other scientists, software engineers, product teams, data and human-annotation teams, and external partners to deliver evaluation results that are technically defensible and useful in practice. You will develop novel benchmarks and evaluation methodologies that are publishable at top tier AI conferences.

Requirements

  • Scientific rigor
  • Strong engineering judgment
  • Clear writing
  • Ability to make progress when requirements, model access, data, or infrastructure are still evolving.

Responsibilities

  • Independently own complex evaluation work from problem definition, benchmark development, and final recommendation to executive leadership.
  • Translate ambiguous product and customer questions into measurable hypotheses.
  • Select or create appropriate benchmarks.
  • Design experiments.
  • Build evaluation pipelines.
  • Validate data and metrics.
  • Analyze failure modes.
  • Communicate conclusions to science, engineering, product, and leadership stakeholders.
  • Write high-quality code.
  • Work with large and imperfect datasets.
  • Develop and calibrate automated evaluators.
  • Turn one-off analyses into reproducible evaluation protocols and reusable infrastructure.
  • Examine factors such as statistical validity, data provenance, contamination, robustness, cost, latency, reliability, safety, and operational constraints.
  • Collaborate closely with other scientists, software engineers, product teams, data and human-annotation teams, and external partners.
  • Deliver evaluation results that are technically defensible and useful in practice.
  • Develop novel benchmarks and evaluation methodologies that are publishable at top tier AI conferences.

Benefits

  • Flexible medical
  • Life insurance
  • Retirement options
  • Volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service