Manager; AI Evaluation Engineering

Caterpillar Inc.Peoria, IL
$147,760 - $240,110

About The Position

Join the AI Engineering team of Cat Digital and take charge of leading a team dedicated to evaluating and validating our advanced generative AI solutions—including intelligent agents, digital assistants, and other innovative capabilities. You will shape the development of cutting-edge products that redefine how customers and dealers interact with technology through AI-driven experiences.

Requirements

  • Knowledge of major products and services and product and service groups; ability to apply knowledge of product and service appropriately to diverse situations.
  • Knowledge of software development tools and activities; ability to produce software products or systems in line with product requirements.
  • Knowledge of software development life cycle; ability to use a structured methodology for delivering and managing new or enhanced software products to the marketplace.
  • Knowledge of software quality assurance and testing; ability to apply appropriate processes, tools, and techniques for assuring a high level of quality in computer software products and systems.

Nice To Haves

  • Proven experience leading software quality, test automation, validation, and AI evaluation teams, with the ability to scale processes, tools, metrics, and engineering practices across multiple products and enterprise initiatives.
  • Demonstrated success delivering enterprise-scale software and GenAI solutions across hybrid cloud and embedded/edge environments, leveraging modern software engineering practices including CI/CD, automated testing, incident management, feature flags, and progressive deployment strategies (e.g., blue/green and canary deployments).
  • Deep expertise in GenAI architecture and system design, including LLMs, SLMs, multimodal, speech, and real-time models; prompt engineering; agentic systems; tool use and orchestration; RAG architectures; vector databases, embeddings, and chunking strategies; fine-tuning techniques such as LoRA; and emerging standards including MCP and A2A.
  • Hands-on experience with AI development and evaluation ecosystems, including Azure, AWS, GCP, Azure AI Foundry, SageMaker, Bedrock, Snowflake Cortex, LangChain, LangGraph, Langfuse, Arize, LangSmith, Humanloop, Ragas, DeepEval, Phoenix, or similar technologies.
  • Deep understanding of AI evaluation methodologies and the challenges of non-deterministic systems, including metric design, RAG quality assessment, output reliability, safety evaluation, A/B testing, human-in-the-loop validation, and balancing deterministic and probabilistic testing approaches.
  • Expertise in modern AI observability and monitoring practices, including tracing, telemetry, evaluation data collection, root-cause analysis, and the use of AI technologies to enhance testing through automated test generation, synthetic data creation, scenario expansion, and LLM-assisted evaluation.

Responsibilities

  • Providing strong technical support and clear direction to ensure the team is aligned with company goals and capable of evaluating advanced AI projects.
  • Overseeing the performance of both individual team members and the team as a whole, fostering a culture of learning by identifying and addressing training and development needs.
  • Taking ownership of the quality of AI engineering products, making sure all solutions are robust, reliable, and meet high standards.
  • Establishing and supervising the implementation of engineering best practices to maintain consistency and excellence in development processes.

Benefits

  • Medical, dental, and vision benefits
  • Paid time off plan (Vacation, Holidays, Volunteer, etc.)
  • 401(k) savings plans
  • Health Savings Account (HSA)
  • Flexible Spending Accounts (FSAs)
  • Health Lifestyle Programs
  • Employee Assistance Program
  • Voluntary Benefits and Employee Discounts
  • Career Development
  • Incentive bonus
  • Disability benefits
  • Life Insurance
  • Parental leave
  • Adoption benefits
  • Tuition Reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service