About The Position

Apple’s Sales organization generates the revenue needed to fuel our ongoing development of products and services. This, in turn, enriches the lives of hundreds of millions of people around the world. We are, in many ways, the face of Apple to our largest customers. Apple's US Decision Intelligence (DI) team is looking for a talented individual who is passionate about crafting, implementing, and operating AI solutions that have a direct and measurable impact on Apple Sales and its customers. We’re seeking a visionary AI Evaluations Engineer to own the end-to-end evaluation pipeline for our AI products and agentic workflows. This role will focus on implementing and maintaining evaluation frameworks, instrumentation, and workflows that help us understand how well our AI systems perform, where they fail, and how they improve over time. You own the evaluation gate and the standards. This role will operate in both capacities, to augment existing AI roadmap, as well as innovate and trailblaze new frontier-technology projects, crafting AI experiences that reduce time to insight and catalyze decision making.

Requirements

  • 5+ years of experience in data and AI-related fields such as AI engineering, software development, ML engineering, data science, or QA roles.
  • Eagerness and ability to learn new skills and solve dynamic problems in an encouraging and expansive environment.
  • Strong Python skills.
  • Hands-on experience with AI evaluation techniques, such as Golden datasets, LLM-as-a-Judge, or rubric-based scoring.
  • Experience with different LLM ecosystems (OpenAI, Anthropic, Gemini, etc.), RAG pipelines, vector databases (e.g., Pinecone, FAISS, Milvus, PostgreSQL).
  • Proficiency in SQL and experience with at least one major data analytics platform, such as Hadoop, Spark, or Snowflake.
  • Experience with CI/CD or release validation workflows.
  • Experience working with data science teams on insights generation leveraging LLMs.
  • Strong time management skills with the ability to collaborate across multiple teams.
  • Able to balance competing priorities, long-term projects, and ad hoc requirements.
  • Ability to work in a fast-paced, dynamic, constantly evolving business environment.
  • Hands-on experience with Langfuse or similar tools for LLM observability.
  • Comfortable working with product/domain experts to translate fuzzy correctness criteria into measurable rubrics or metrics.
  • B.S. degree in Computer Science/Engineering, or equivalent work experience

Nice To Haves

  • Sound communication skills - expert at messaging domain and technical content, at a level appropriate for the audience.
  • Strong ability to gain trust with stakeholders and senior leadership.
  • Familiarity with embeddings, retrieval algorithms, agents, and data modeling for vector and graph databases.
  • Other complementary technologies for distributed systems architecture and asynchronous messaging, agent communication, and caching like RabbitMQ, Redis, and Valkey are preferred.
  • Experience working across global teams to ensure alignment of product development.
  • Applied knowledge of GenAI and RAG strategies, microservices, recommendation systems, and context engineering.
  • Working knowledge of agent evaluation concepts like trajectory vs. end-to-end vs. component-level evaluation, tool-call correctness.
  • Advanced degree (MS or Ph.D.) in Economics, Electrical Engineering, Statistics, Data Science, or a similar quantitative field is preferred.

Responsibilities

  • Own the end-to-end evaluation pipeline for AI products and agentic workflows.
  • Implement and maintain evaluation frameworks, instrumentation, and workflows.
  • Understand how well AI systems perform, where they fail, and how they improve over time.
  • Augment existing AI roadmap.
  • Innovate and trailblaze new frontier-technology projects.
  • Craft AI experiences that reduce time to insight and catalyze decision making.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service