Principal DevOps Engineer

KēSTA I.T.Salt Lake, UT

About The Position

KēSTA I.T. is actively seeking a Principal DevOps Engineer for an immediate full-time opportunity with an industry-leading client. This role offers leadership, responsibility, and the chance to make a significant impact within a thriving and stable organization. We are seeking a technical leader with deep AWS and AI infrastructure expertise to join an enterprise technology organization. This role plays a key part in architecting and driving cloud and AI infrastructure strategy across AWS and private cloud environments. The Principal DevOps Engineer will establish platform standards that enable Bedrock-backed agents, agentic workflows using AgentCore, and high-throughput LLM serving to operate reliably, securely, and cost-effectively at enterprise scale. This position combines hands-on technical depth with architectural leadership, platform strategy, and mentorship. The ideal candidate has extensive experience building enterprise-scale AWS and Kubernetes environments and is comfortable establishing the infrastructure and operational standards required to support emerging AI workloads.

Requirements

  • 8+ years of experience in DevOps, platform engineering, cloud infrastructure, or a related discipline
  • Extensive experience administering Amazon EKS or Kubernetes within enterprise-scale environments
  • Advanced scripting experience using Bash and Python
  • Extensive experience designing and managing CI/CD pipelines and GitOps workflows using GitHub Actions, ArgoCD, or similar technologies
  • Deep experience with AWS services, including EC2, VPC, S3, EKS, IAM, CloudFront, ALB/NLB, and RDS
  • Strong experience designing Infrastructure as Code and cloud-native architectures using Terraform and AWS CDK
  • Extensive Linux systems administration experience
  • Proven expertise with AWS cloud and AI infrastructure, including Amazon Bedrock, AgentCore, and cloud-native architecture patterns
  • Experience deploying and operating agentic AI systems within AWS, including AgentCore and Bedrock Agents
  • Strong experience with Dynatrace and enterprise observability strategies, including monitoring AI workloads
  • Experience with centralized logging and monitoring technologies such as CloudWatch and Dynatrace
  • Deep understanding of cloud networking, including AWS VPC, Transit Gateway, security groups, load balancing, and routing
  • Strong understanding of cloud-native principles, Infrastructure as Code, and GitOps methodologies
  • Knowledge of CIS, NIST, and SOX compliance frameworks and associated audit processes
  • Strong understanding of security governance and compliance requirements for cloud and AI environments
  • Experience with service discovery technologies such as AWS Cloud Map and Consul
  • Demonstrated ability to mentor and develop engineers at multiple levels
  • Ability to communicate complex infrastructure and architecture concepts effectively to both technical and business stakeholders

Responsibilities

  • Architect cloud and AI platform strategies and clearly communicate proposed solutions to both technical and non-technical stakeholders
  • Provide technical leadership to engineering team members, manage work priorities, and establish technical direction for cloud infrastructure and AI platform practices
  • Collaborate with architecture leadership to define, design, and validate solution architectures for platform and AI application workloads running on AWS
  • Establish platform standards for AI workload deployment, including model versioning, inference-serving patterns, cost-per-inference governance, and AI API reliability
  • Administer and architect Amazon EKS environments at enterprise scale, including multi-cluster federation and AI workload scheduling policies
  • Design and implement enterprise observability strategies using Dynatrace, including service-performance SLOs, cost dashboards, AI workload monitoring, and incident response
  • Establish cloud-native architecture, Infrastructure as Code, and GitOps standards that support scalable and repeatable platform operations
  • Help ensure cloud and AI infrastructure meets enterprise requirements for security, reliability, performance, compliance, and cost efficiency
  • Mentor and develop engineers across multiple levels while promoting platform engineering and DevOps best practices

Benefits

  • Top performance is rewarded
  • Personal time is valued
  • Excellence is demanded at every level
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service