Data Scientist - Evaluating Foundation Models and AI Agents for National Security

Pacific Northwest National Laboratory•Richland, WA
•Onsite

About The Position

The National Security Directorate (NSD) at PNNL drives science-based, mission-focused solutions to address complex, real-world threats. The AI and Data Analytics Division within NSD combines domain expertise with advanced hardware and software to deliver computational solutions for complex data and analytic challenges. They work in multidisciplinary teams, connecting foundational research to engineering and operations to innovate quickly and field results faster. Their strengths span the entire data analytics lifecycle, from acquisition and management to analysis and decision support.

Requirements

  • BS/BA and 2 years of data science experience, OR MS/MA, OR PhD.
  • Experience developing or evaluating LLMs, VLMs, pretrained transformer models, AI agents, or integrated AI systems.
  • Experience using Python and machine-learning frameworks.
  • Ability to obtain and maintain a federal security clearance.
  • U.S. Citizenship.
  • Must pass a Federal background investigation.
  • Must pass pre-employment drug testing and comply with PNNL Workplace Substance Abuse Program.
  • Must demonstrate non-use of illegal drugs, including marijuana, for 12 consecutive months preceding completion of the Questionnaire for National Security Positions (QNSP).

Nice To Haves

  • Experience building evaluation datasets, test harnesses, agentic harnesses, automated graders, behavioral tests, or analysis pipelines.
  • Experience analyzing model behavior, failure modes, robustness, calibration, uncertainty, or performance across meaningful data slices.
  • Familiarity with agent orchestration frameworks such as LangGraph or similar technologies involving routing, retrieval, tools, memory, state, and human oversight.
  • Experience with deterministic NLP, computer vision, retrieval, search, rules, or classical machine-learning methods that complement generative AI.
  • Experience in one or more related areas such as explainable AI, model probing, mechanistic interpretability, AI security, red teaming, uncertainty quantification, or privacy-preserving AI.
  • Experience contributing to applied research through technical reports, publications, prototypes, or open-source software.
  • Prior experience supporting government, national security, or other high-consequence applications.
  • Active TS/SCI clearance.

Responsibilities

  • Design and implement evaluation studies for AI/ML, with a focus on pretrained transformer-based foundation models (LLMs, VLMs) and the AI agents that use them.
  • Prepare evaluation datasets, prompts, mission scenarios, meaningful data slices, test cases, scoring rubrics, and deterministic baselines.
  • Develop evaluation environments, test harnesses, and agentic harnesses for exercising models, tools, retrieval systems, and end-to-end workflows.
  • Execute controlled experiments at the model, component, workflow, and system levels.
  • Analyze model behavior to identify capabilities, weaknesses, failure modes, and sensitivity to various factors.
  • Examine agent trajectories, intermediate decisions, tool calls, retrieved evidence, memory/state, error propagation, recovery behavior, and points of human intervention.
  • Evaluate performance using aggregate metrics, case-level analysis, behavioral categories, robustness testing, calibration, uncertainty characterization, and adversarial/off-nominal scenarios.
  • Evaluate hybrid systems combining generative AI with deterministic NLP, computer vision, retrieval, search, rules, or classical machine learning.
  • Develop and validate automated or model-based graders against human judgments and mission-relevant criteria.
  • Apply functional explainability, model probing, or selected interpretability methods to understand model and system behavior.
  • Document evaluation methods, assumptions, results, uncertainties, and limitations in reproducible code, visualizations, technical reports, and briefings.
  • Contribute technical material, data, figures, demonstrations, or preliminary results to proposals and sponsor engagements.

Benefits

  • Medical insurance
  • Dental insurance
  • Vision insurance
  • Robust telehealth care options
  • Several mental health benefits
  • Free wellness coaching
  • Health savings account
  • Flexible spending accounts
  • Basic life insurance
  • Disability insurance
  • Employee assistance program
  • Business travel insurance
  • Tuition assistance
  • Relocation
  • Backup childcare
  • Legal benefits
  • Supplemental parental bonding leave
  • Surrogacy and adoption assistance
  • Fertility support
  • Company-funded pension plan
  • 401(k) savings plan with company match
  • Up to 120 vacation hours per year
  • Ten paid holidays per year
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service