Principal Machine Learning Engineer

Palo Alto Networks•Office - USA - CA - Headquarters, CA
•$163,200 - $264,000•Onsite

About The Position

The AI Canvas team is building our next-generation data exploration experience for cybersecurity. AI Canvas is designed to transform how security teams explore, understand, and act on complex security data by combining natural language interaction, real-time visualization, dashboards, and intelligent agents in a unified experience. We are building beyond traditional dashboards and static user interfaces. AI Canvas is based on a modern agent- and skill-driven architecture, bringing together AI, data exploration, visualization, and collaboration. We are looking for engineers who are comfortable operating in ambiguity, excited by hard technical problems, and motivated to build reliable AI systems that security teams can trust. As a Principal Machine Learning Engineer on the AI Canvas team, you will take significant technical ownership of the AI layer powering the platform. You will help define the architecture, build production-grade AI systems, and shape how intelligence is integrated throughout the product. You will work across large language models (LLMs), retrieval-augmented generation (RAG), AI agents and assistants, agent harnesses, natural-language-to-query generation, evaluation systems, guardrails, and cloud-based AI infrastructure. This is a highly technical and hands-on role. You will work closely with ML, backend, UI, product, and design teams to solve challenging AI problems in cybersecurity, where reliability, accuracy, scalability, latency, observability, and trust are critical.

Requirements

  • Bachelor’s degree in Computer Science, Machine Learning, Engineering, or a related technical field, or equivalent practical experience.
  • 7+ years of software engineering, machine learning engineering, or related industry experience, including significant experience building production systems.
  • Strong experience designing and building machine learning or AI-powered applications at scale.
  • Hands-on experience building applications using large language models and generative AI technologies.
  • Experience with one or more areas such as RAG, AI agents, AI assistants, tool-calling systems, agent orchestration, or LLM-based workflows.
  • Strong proficiency in Python and experience developing production-grade backend or ML services.
  • Experience designing evaluation systems for AI/ML applications, including offline evaluation, regression testing, quality measurement, and production monitoring.
  • Strong understanding of modern ML and AI concepts, including embeddings, retrieval, ranking, prompt engineering, model inference, and experimentation.
  • Experience building scalable systems on public cloud platforms such as GCP, AWS, or Azure.
  • Strong software engineering fundamentals, including system design, distributed systems, APIs, testing, CI/CD, and observability.
  • Ability to lead complex technical initiatives across multiple teams while remaining deeply hands-on.
  • Strong communication skills and the ability to translate ambiguous product problems into clear technical architectures and execution plans.

Nice To Haves

  • Master’s or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field.
  • Deep experience building and operating LLM-powered products or agentic AI platforms in production.
  • Experience with AI development frameworks and tooling for model orchestration, agent systems, evaluation, tracing, or observability.
  • Experience building natural-language-to-SQL or natural-language-to-query systems, semantic layers, or AI-powered data exploration products.
  • Experience designing LLM evaluation harnesses, synthetic test generation, LLM-as-a-judge approaches, or human-in-the-loop evaluation workflows.
  • Experience with vector databases, search and retrieval infrastructure, embeddings, and large-scale knowledge systems.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Experience designing highly available, low-latency AI inference and backend services.
  • Familiarity with AI security, prompt injection defenses, data privacy, access control, and guardrails for enterprise AI applications.
  • Experience in cybersecurity, security analytics, observability, or large-scale data platforms.
  • Contributions to open-source AI, ML, agent, or infrastructure projects.

Responsibilities

  • Provide technical leadership for the architecture and development of the AI capabilities powering AI Canvas, from early design through production deployment and continuous improvement.
  • Design and build scalable, production-grade systems using LLMs, RAG, AI agents, assistants, tools, skills, and agent harnesses.
  • Architect intelligent workflows that enable users to explore complex security data through natural language, including natural-language-to-query generation, follow-up interactions, clarification, investigation, and troubleshooting.
  • Define and evolve the architecture for agentic AI systems, including orchestration, context management, tool invocation, memory, reasoning workflows, and multi-step task execution.
  • Build robust evaluation frameworks and harnesses for measuring AI quality, including correctness, relevance, reliability, regression detection, and end-to-end product behavior.
  • Establish evaluation methodologies using automated metrics, LLM-based evaluators, human evaluation, and representative production datasets.
  • Design and implement guardrails and safety mechanisms to improve reliability, reduce hallucinations, enforce system constraints, and ensure responsible behavior of AI-powered features.
  • Drive improvements in model and system quality through prompt engineering, retrieval strategies, model selection, fine-tuning where appropriate, and systematic experimentation.
  • Build scalable RAG and knowledge-retrieval systems capable of grounding AI responses in large, complex, and evolving security datasets.
  • Partner closely with backend and platform engineers to build reliable APIs, services, and infrastructure supporting AI workloads at production scale.
  • Establish engineering best practices around observability, debugging, tracing, versioning, reproducibility, testing, and monitoring of LLM and agent-based systems.
  • Evaluate emerging models, frameworks, and AI infrastructure and determine when and how they should be incorporated into the product.
  • Balance rapid experimentation with the engineering rigor required to operate mission-critical AI systems in production.
  • Mentor engineers, influence technical direction across teams, and raise the engineering bar for applied AI development.

Benefits

  • A description of our employee benefits may be found here.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service