Sr. ML Engineer

Visa•Austin, TX
•$123,400 - $191,100•Hybrid

About The Position

Visa AI Studio is Visa's AI operating system: a single platform for building, deploying, and operating predictive models, foundation models, and AI agents at global scale. It gives every team at Visa a common, self service way to train and experiment, build with generative AI, develop and run production AI agents, manage features and memory, and operate everything with governance and observability built in. Visa AI Studio is the foundation for how AI gets built across the company. We are looking for an Machine Learning Engineer to join the AI Engineering Platform team within Visa AI Studio, working on the systems that let every team at Visa train, deploy, and operate models and agents themselves. This is a modern AI engineering role: it sits deliberately at the intersection of core AI/ML knowledge and systems ands software engineering. You need to understand how models actually work, including architectures, training dynamics, embeddings, retrieval, evaluation, and the behavior and failure modes of agents that plan and call tools. And you need to be equally fluent in the systems side: distributed computing, large scale training and batch systems, API and SDK design, observability, and infrastructure that holds up at scale. You should also be AI native in how you build: comfortable pairing with coding agents, LLM powered tooling, and automated evaluation to design and ship the platform itself faster.

Requirements

  • 2+ years of relevant work experience and a Bachelors degree, in Computer Science, Engineering, or a related technical field, OR 5+ years of relevant work experience
  • Professional experience in software engineering, ML engineering, or platform/infrastructure engineering.
  • Good understanding of core machine learning and deep learning concepts: model training, evaluation, transformer/foundation model architectures, and embeddings.
  • Strong systems and software engineering fundamentals: distributed systems, API/SDK design, and cloud infrastructure (e.g., Kubernetes and a major cloud provider such as AWS, Azure, or GCP).
  • Proficiency in at least one language commonly used in AI/platform engineering (e.g., Python, Java, Go).
  • Hands on fluency with modern AI assisted development, using coding agents, copilots, and LLM based tooling to accelerate day to day engineering work.

Nice To Haves

  • 4 or more years of relevant work experience.
  • Experience building with or on top of foundation models / large language models: prompting, fine tuning, retrieval augmented generation (RAG), and evaluation frameworks.
  • Experience designing or operating AI agent systems: tool calling, multi agent orchestration, memory systems, or agent frameworks.
  • Experience building internal developer platforms, SDKs, or self service tooling used by other engineering or data science teams.
  • Experience with feature stores, embedding/vector systems, or knowledge graph and memory systems used by ML or agentic applications.
  • Experience with training and batch efficiency techniques such as distributed scheduling, quantization, caching, or compute right sizing at scale.
  • Experience operating production systems at high scale, such as large scale distributed training jobs or high volume batch processing pipelines.
  • Experience implementing AI governance, responsible AI, or model risk/compliance controls within a platform.

Responsibilities

  • Be involved in designing, build, and operate core components of the AI Engineering Platform: training and experimentation infrastructure, the Batch Platform, agent development frameworks, and the unified AI Studio developer experience.
  • Design, build, and operate the Batch Platform: large scale scheduled and on demand batch scoring, offline model execution, and high throughput data pipelines for training and feature generation.
  • Build self service APIs, SDKs, and tooling that let engineers and data scientists across Visa train, fine tune, deploy, evaluate, and monitor models and agents without platform team intervention.
  • Design and extend agent runtimes and orchestration primitives, including tool calling, memory, planning, and multiagent coordination, for agents that operate safely and predictably in production.
  • Apply core ML and deep learning knowledge (model architectures, embeddings, fine tuning, evaluation methodology) to platform design decisions, not just infrastructure decisions
  • Maintain infrastructure that supports large scale distributed training, high throughput batch processing, and efficient GPU and compute cluster utilization.
  • Be capable of cost and efficiency optimizations for training and batch workloads, such as distributed scheduling, caching, quantization, and compute right sizing.
  • Embed governance, responsible AI, audit, and monitoring capabilities directly into platform components, including drift, hallucination, and anomaly detection for agentic systems.
  • Partner with the AI Context Platform, AI Runtime Services, and AI Governance Platform teams to deliver a coherent, end to end platform experience.
  • Stay current with AI research and the broader platform/tooling ecosystem, and bring in techniques and patterns before they become standard practice.
  • Use AI coding agents and LLM powered developer tools as part of your own workflow to design, build, and ship AI Studio capabilities faster, treating AI assisted engineering as a core skill, not a side habit.

Benefits

  • Medical
  • Dental
  • Vision
  • 401(k)
  • FSA/HSA
  • Life Insurance
  • Paid Time Off
  • Wellness Program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service