Senior Data Analyst

Rightsline Inc•El Segundo, CA
•$110,000 - $120,000•Remote

About The Position

Reporting to the SVP Engineering, the Senior Data Analyst – LLMOps will work with the Application Development team to own the observability, evaluation, and quality measurement of Rightsline’s LLM-powered features. This is a senior individual-contributor role in which you will instrument the LLM stack end to end — traces, prompts, tool calls, latency, token spend, and output quality — build the evaluation suites that tell us whether a change made the product better or worse, and turn that telemetry into decisions the engineering and product teams act on. You will work across a SaaS application, a wide range of third-party tools, an AWS-hosted infrastructure, and an increasingly agentic development workflow. The position is based in North America, has no direct reports, and partners closely with the wider engineering, product, and data teams. The Senior Data Analyst – LLMOps makes the behavior of Rightsline’s AI features measurable. Quality questions about prompts, retrieval, and agent workflows are too often answered with intuition and spot checks; you will replace that with tracing, offline and online evaluations, regression suites, and dashboards that show quality, cost, and latency trends over time. You will be working in a collaborative environment with a team who shares your passion for building something impactful, and success in the role looks like engineers and product managers reaching for your evaluation results and dashboards before they ship — and trusting what they see.

Requirements

  • 5+ years in data analytics, analytics engineering, or data science, including hands-on work with LLM-based systems in production
  • Strong SQL and Python (including pandas and notebook-based analysis), with the discipline to build repeatable pipelines rather than one-off queries
  • Hands-on experience with LLM observability and evaluation tooling such as LangSmith, Langfuse, LangChain/LangGraph, or equivalent
  • Practical understanding of LLM application architecture — prompting, retrieval-augmented generation, tool calling, and agent workflows — and how each of them fails
  • Experience designing evaluation methodology: golden datasets, scoring rubrics, LLM-as-judge grading, human annotation, and A/B testing on production traffic
  • Working knowledge of experiment design and statistics, enough to say confidently whether a difference is real
  • Ability to work independently, meet deadlines, and be accountable for your work, while operating within established standards
  • Excellent written and verbal communication skills, and a strong sense of ownership, urgency, and initiative — you can explain a quality regression to an engineer and to an executive in the same afternoon
  • BS in Computer Science, Statistics, Data Science, or equivalent knowledge

Nice To Haves

  • Experience with AWS services such as Lambda, SQS/SNS, API Gateway, Step Functions, and Amazon Bedrock
  • Familiarity with dbt, Airflow or a similar orchestrator, Git, and BI tools such as Looker, Power BI, or Tableau
  • Working knowledge of MS SQL, Redis or another NoSQL store, and vector stores used for retrieval
  • Exposure to model providers and platforms (Anthropic, OpenAI, Amazon Bedrock), guardrail and PII-redaction tooling, and prompt versioning and release workflows
  • Experience with coding agents or AI developer tools (Claude Code, GitHub Copilot, Cursor, or similar) in day-to-day analysis work

Responsibilities

  • Instrument Rightsline’s LLM and agent workflows end to end with tracing and structured logging, using tools such as LangSmith, Langfuse, or equivalent
  • Define and maintain the metrics that describe AI system health — latency, token and cost per request, error and fallback rates, tool-call success, and retrieval hit rates
  • Build and maintain dashboards and alerting so quality and cost regressions surface in hours rather than in customer tickets
  • Design and maintain offline evaluation suites — golden datasets, scoring rubrics, and LLM-as-judge grading — for prompts, retrieval pipelines, and agent flows
  • Run online evaluations and A/B experiments on production traffic, and report results with enough rigor to support a ship / no-ship decision
  • Track quality across model, prompt, and retrieval changes, and maintain the benchmark history that shows whether the product is actually improving
  • Partner with engineers on human review and annotation workflows, including labeling guidelines and inter-rater agreement
  • Build and maintain the pipelines that move trace, evaluation, and usage telemetry into the analytics stack using SQL and Python
  • Analyze how customers actually use AI features — adoption, drop-off, and failure patterns — and translate findings into prioritized recommendations
  • Produce recurring reporting on AI quality, cost, and usage for engineering, product, and leadership audiences
  • Monitor token spend and model usage, and identify where prompt, model, or caching changes cut cost without hurting quality
  • Investigate quality incidents — bad or unsafe outputs, hallucinations, tool failures, latency spikes — from trace to root cause
  • Follow and help extend Rightsline’s development best practices and standards in the pipelines, queries, and evaluation code you ship
  • Attend and contribute to team meetings, and present telemetry and evaluation findings in a way non-specialists can act on
  • Raise the team’s LLMOps literacy — help engineers instrument their own features, write their own evals, and read the dashboards

Benefits

  • Equal opportunity workplace
  • Dedicated to growing a diverse team
  • Inclusive environment where everyone feels empowered to bring their authentic selves to work
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service