Senior Lead Software Engineer- AI/ML Platform

JPMorgan Chase & Co.Jersey City, NJ
$152,000 - $260,000

About The Position

Be an integral part of an agile team that's constantly pushing the envelope to enhance, build, and deliver top-notch technology products. As a Senior Lead Software Engineer at JPMorgan Chase within Corporate - AIML Data Platforms team, you will design, build, and operate the foundational cloud infrastructure that enables data scientists and machine learning engineers to develop, train, and deploy intelligent solutions across the firm. In this role you will serve as a technical leader, driving platform reliability, scalability, and automation while collaborating with cross-functional teams to solve complex infrastructure challenges. Your work will directly accelerate the firm’s AI/ML capabilities—enabling faster experimentation and production-grade deployments that create measurable business impact.

Requirements

  • Formal training or certification on software engineering concepts and 5+ years applied experience
  • Experience delivering secure, production-quality code in Python or Java.
  • Strong foundations in distributed systems, microservices, and platform architecture/design principles.
  • Proven ability to architect and operate cloud-native infrastructure on AWS (compute, networking, storage, security) and other major clouds.
  • Demonstrated expertise with infrastructure-as-code tooling, specifically Terraform, in large-scale cloud environments.
  • Hands-on experience with Docker and Kubernetes, including AWS EKS operations.
  • Experience building or supporting production AI/ML platforms (training, deployment, and model serving/inference), including GPU infrastructure/tooling.
  • Strong DevOps/platform engineering practices: CI/CD, release automation, automated testing, and observability (monitoring/logging/tracing).
  • Experience with SQL/NoSQL databases and data integration; strong Linux, scripting, and networking fundamentals.
  • Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, code review, test acceleration, troubleshooting) with the ability to set team expectations for validating AI outputs for correctness, performance, and security
  • Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching senior engineers/leads on compliant usage patterns and controls.

Nice To Haves

  • Proficiency in Go or Python for automation, tooling development, or platform service implementation.
  • Experience with MLOps frameworks and tools such as Kubeflow, MLflow, or similar AI/ML lifecycle management platforms.
  • Working knowledge of ML frameworks (PyTorch, TensorFlow, Hugging Face, scikit-learn) for model integration and operationalization.
  • Exposure to multi-cloud or hybrid cloud architectures and platform portability strategies.

Responsibilities

  • Builds and maintains reusable AI/ML platform infrastructure and shared services to support development, deployment, and operations at scale.
  • Architects, deploys, and operates secure cloud and container-based environments for training and inference, including GPU-intensive workloads.
  • Design and implement platform tooling, automation, and infrastructure-as-code solutions to streamline model deployment, environment provisioning, release management, and operational support.
  • Develops and maintains production-grade services, APIs, SDK integrations, and workflows that support model training, serving, evaluation pipelines, and AI application lifecycle management.
  • Partners with data science, ML engineering, and application teams to translate model and compute requirements into platform standards and deployment patterns.
  • Optimizes platform reliability, scalability, latency, and cost through orchestration, scheduling, and hardware acceleration.
  • Establishes operational best practices including monitoring, logging, observability, access controls, incident response, and production troubleshooting.
  • Supports enterprise LLM operationalization, including fine-tuning workflows, inference optimization, and evaluation; contribute to documentation and engineering standards.
  • Drives adoption and governance of approved AI-assisted engineering practices across teams to improve code quality, delivery speed, and operational outcomes (e.g., AI-assisted code review/refactoring, test acceleration, release readiness, incident/root-cause analysis), while establishing measurable validation standards (secure coding, peer review, automated testing) and promoting reuse of proven patterns and automation within the SDLC/TLM toolchain.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including approved AI-assisted development and automation capabilities, to improve the value realized by automation at scale.

Benefits

  • comprehensive health care coverage
  • on-site health and wellness centers
  • a retirement savings plan
  • backup childcare
  • tuition reimbursement
  • mental health support
  • financial coaching
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service