Director AI Engineering (Remote Eligible)

Capital One•San Jose, CA
•$244,700 - $335,100•Remote

About The Position

At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent — along with our deep experience in machine learning — position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Generative AI Training team builds the platform Capital One's scientists and engineers use to train, fine-tune, and experiment with foundation models at scale. We run the GPU clusters and distributed training infrastructure behind the company's generative AI: fair-share scheduling across teams, large-scale fine-tuning and reinforcement learning, resilience for long-running jobs, and the utilization controls that keep that expensive hardware working efficiently. We also provide the self-service environments where teams deploy, serve, and evaluate models during experimentation, built on tools like KServe and vLLM. The platform we build is a central leverage point for AI across Capital One, shaping how fast and how affordable the rest of the company can build with foundation models.

Requirements

  • Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 6 years of experience developing AI and ML algorithms or technologies
  • At least 3 years of people leadership experience

Nice To Haves

  • 5+ years of experience managing and leading an engineering team
  • 7+ years of experience building and operating large-scale ML or GPU training infrastructure on cloud platforms (e.g., AWS, Google Cloud, Azure, or equivalent private cloud)
  • Hands-on experience with distributed training at scale: multi-node, multi-GPU jobs and parallelism strategies (e.g., PyTorch FSDP, DeepSpeed, Megatron), with proficiency in Python, Go, C++, or CUDA
  • Experience with the ML orchestration and scheduling stack (e.g., Kubernetes, Kubeflow, Kueue, Slurm, Ray, KServe, vLLM)
  • Experience operating large GPU fleets with a focus on reliability, fault tolerance, utilization, and cost efficiency
  • Experience right-sizing GPU clusters, instance types, interconnect, and quotas to training and experimentation workload requirements (e.g., model size, parallelism strategy, throughput targets)
  • Passion for staying abreast of the latest AI and ML-systems research, and judiciously applying novel training and optimization techniques
  • Excellent communication and presentation skills, with the ability to articulate complex AI and infrastructure concepts to peers
  • Experience building and leading a multi-team AI organization delivering multiple enterprise capabilities concurrently
  • Proven ability to expand and execute long-term AI platform strategies aligned to enterprise priorities and regulatory frameworks
  • Experience establishing cross-functional operating rhythms and review cadences (OKRs, AI governance councils, quarterly reviews)

Responsibilities

  • Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One.
  • Oversee the design, development, testing, deployment, and operation of the platform's core systems: distributed training and fine-tuning, reinforcement learning workflows, fair-share GPU scheduling, job resilience and fault tolerance, GPU utilization and efficiency, and self-service environments for model experimentation and evaluation.
  • Make high judgment build-vs-buy decisions across a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more.
  • Invent and introduce state-of-the-art techniques to improve the scalability, cost, throughput, and reliability of large-scale distributed training and fine-tuning.
  • Own GPU capacity planning and cost governance: right-size clusters, instance types, and quotas to the needs of training and experimentation workloads across teams.
  • Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One.
  • Attract and retain top talent in the AI industry and nurture personal and professional development for your team. Foster a culture of learning and staying abreast of the state-of-the-art in AI.
  • Translate the enterprise AI strategy into portfolio-level execution plans across multiple product areas, balancing innovation with delivery discipline
  • Scale AI engineering practices across teams through shared infrastructure, reusable components, and unified observability and governance frameworks
  • Establish enterprise standard for Responsible AI, including fairness metrics, model evaluation protocols, documentation requirements, and audit readiness
  • Partner with research, compliance, and enterprise risk teams to ensure deployed systems meet emerging ethical and regulatory standards

Benefits

  • comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being
  • performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service