About The Position

We have an exciting and rewarding opportunity for you to take your software engineering career to the next level. As a Software Engineer III at JPMorganChase within the AI/ML data platform team you serve as a seasoned member of an agile team to build and operate scalable, reliable ML training systems and pipelines on AWS and other cloud platforms. You will productionize training workloads (often GPU-based), improve performance and cost efficiency, and enable repeatable, well-governed training across environments.

Requirements

  • Formal training or certification on software engineering concepts and 3+ years applied experience
  • Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code.
  • Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management).
  • Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows).
  • Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow).
  • Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology.
  • Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis).
  • Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling).
  • Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2)
  • Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security.
  • Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices.

Nice To Haves

  • Experience running training workloads across multiple cloud platforms and managing portability, performance, and governance across environments.
  • Familiarity with cloud-native networking/storage patterns for high-throughput training and artifact management.
  • Experience optimizing training input pipelines (sharding, prefetching, caching, format choices such as Parquet/WebDataset) and working with large datasets.
  • Familiarity with distributed compute frameworks (Spark, Ray, Dask) for feature/dataset generation.
  • Familiarity with workflow orchestration tools (Airflow-like systems, Argo Workflows-like patterns) and model registry concepts.
  • Experience optimizing training cost/performance (right-sizing, scheduling policies, interruptible capacity strategies where applicable budge guardrails, quota planning).
  • Strong observability practice for training systems: metrics/logs/traces, GPU telemetry, dashboards, and alert tuning.

Responsibilities

  • Design, build, and maintain end-to-end ML training platform.
  • Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility.
  • Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting.
  • Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls.
  • Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks.
  • Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls).
  • Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow.
  • Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness.
  • Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service