AWS Platform Engineer

Indicium AINew York, NY
$135,000 - $190,000

About The Position

Indicium AI is seeking a Senior AWS Platform / DevOps Engineer to design, automate, and scale cloud infrastructure with a core focus on enabling next-generation AI/ML applications. The role involves leveraging expertise in AWS, Kubernetes (EKS), Infrastructure as Code (Terraform), and CI/CD automation to build a resilient platform foundation. Additionally, the position requires expanding into MLOps, deploying and orchestrating AI workloads, vector stores, and model inference pipelines using services like Amazon Bedrock and SageMaker. This is an ideal opportunity for a high-performing DevOps or Platform Engineer, particularly those with a cloud consulting or professional services background, who wish to build production-grade AI infrastructure at scale.

Requirements

  • 4+ years of hands-on experience in DevOps, Platform, or Infrastructure Engineering using AWS, Terraform, and Kubernetes (EKS) in production environments.
  • Deep understanding of Docker, Helm, Kubernetes architecture, and continuous delivery tools (ArgoCD, Flux, GitHub Actions, or GitLab CI).
  • Practical exposure or strong motivation to build infrastructure for AI/ML workloads (e.g., model serving, SageMaker, Amazon Bedrock, vector databases, GPU node groups).
  • Experience working in cloud consulting, system integration, or fast-paced delivery environments where adapting to new tools and client requirements is second nature.
  • Proficiency in scripting/programming with Python, Bash, or Go.
  • Ability to articulate architectural trade-offs, collaborate across software and data teams, and document infrastructure patterns.

Nice To Haves

  • AWS Certified Solutions Architect or DevOps Engineer - Professional.
  • Hands-on experience with GPU node scaling via Karpenter on EKS.
  • Familiarity with MLOps frameworks like MLflow, Kubeflow, or Ray.
  • Experience implementing cost optimization (FinOps) strategies for heavy compute/AWS workloads.

Responsibilities

  • Design, build, and maintain modular Infrastructure as Code (IaC) using Terraform and AWS best practices across multi-account environments.
  • Manage production AWS EKS clusters, Helm charts, and ingress controllers; configure dynamic autoscaling (Karpenter/HPA) to handle specialized CPU/GPU compute workloads.
  • Establish automated GitOps delivery pipelines (ArgoCD or Flux) to ensure zero-downtime releases for microservices and AI application services.
  • Build and optimize cloud infrastructure supporting Machine Learning workflows, including model serving, vector databases (e.g., OpenSearch, pgvector), and API integrations with services like Amazon Bedrock and SageMaker.
  • Implement fine-grained IAM policies, network security (VPCs, security groups), secrets management, and observability stacks (Prometheus, Grafana, CloudWatch).
  • Partner with software developers and data/AI teams to streamline internal developer workflows, reduce friction, and standardize deployment patterns.

Benefits

  • Competitive pay with performance bonuses
  • Comprehensive health coverage
  • 401k Employer Match
  • Generous PTO
  • Flexible holidays
  • Parental leave
  • Paid company shutdown the last week of December
  • Generous learning budget
  • Dedicated research time
  • Unencumbered access to state-of-the-art AI tools
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service