Senior DevOps Engineer

Computer Task Group, IncUNAVAILABLE, UNAVAILABLE
Remote

About The Position

We are seeking a Senior DevOps Engineer to join our team on a remote contract basis. This is a hands-on role for an engineer with strong experience in AWS, Kubernetes, Terraform, and CI/CD who can independently drive reliable infrastructure decisions in ambiguous environments. You will be instrumental in automating and operating cloud and Kubernetes platforms, improving deployment workflows, strengthening security and observability, and ensuring systems are scalable, reliable, and maintainable. This role also expects practical fluency with modern AI-assisted engineering workflows to accelerate analysis, troubleshooting, documentation, and delivery while maintaining high standards for correctness, security, and operational safety.

Requirements

  • 7+ years of professional experience in DevOps, platform engineering, SRE, or related roles, with a proven track record designing and operating scalable cloud infrastructure.
  • Strong expertise with AWS services, including EC2, S3, EKS, ECR, Route 53, IAM, and core networking and security concepts.
  • Advanced hands-on experience with Kubernetes, including cluster operations, upgrades, autoscaling, networking, ingress, RBAC, policy enforcement, and production workload management.
  • Strong experience with Terraform for infrastructure-as-code, environment management, and repeatable platform provisioning.
  • Experience with Kubernetes packaging and deployment tools such as Helm and Kustomize.
  • Experience with CI/CD systems such as GitLab CI/CD and familiarity with GitOps approaches using Argo CD, Flux, or similar tools.
  • Strong experience with observability and monitoring tools such as Prometheus, Grafana, Loki, or comparable technologies.
  • Experience evaluating or operating Kubernetes ecosystem tools such as ingress controllers, operators, Cluster Autoscaler, or Karpenter.
  • Strong experience with PostgreSQL/RDS, including performance tuning, reliability, and operational support.
  • Solid understanding of cloud and container security practices, including IAM least privilege, secrets management, image scanning, admission controls, and policy enforcement.
  • Strong scripting and automation skills in Python, Bash, or similar languages.
  • Practical fluency with AI-assisted engineering tools such as LLMs, code assistants, and workflow automation tools, along with strong judgment to evaluate and refine AI-generated outputs responsibly.
  • Excellent problem-solving, teamwork, and communication skills.
  • Ability to work independently in a remote environment and make sound technical decisions in ambiguous situations.

Nice To Haves

  • Certifications in AWS, Kubernetes, Terraform, or related platform technologies.
  • Experience with additional cloud platforms such as GCP or Azure.
  • Experience improving incident response, operational readiness, and platform standards for distributed engineering teams.
  • Experience designing internal platform capabilities that improve developer self-service and reduce operational toil.
  • Familiarity with compliance, governance, and risk management requirements in regulated or enterprise environments.

Responsibilities

  • Design, implement, and operate cloud infrastructure and automation using AWS and Terraform across services such as EC2, S3, EKS, ECR, Route 53, IAM, and core networking components.
  • Own EKS and Kubernetes platform operations, including cluster lifecycle management, upgrades, node group strategy, autoscaling, workload scheduling, capacity planning, performance tuning, and cost optimization.
  • Build and improve Kubernetes platform capabilities such as networking, ingress, DNS integration, service discovery, RBAC, policy enforcement, secrets management, and multi-environment consistency.
  • Develop and maintain deployment workflows using Helm, Kustomize, GitLab CI/CD, and GitOps approaches with tools such as Argo CD or Flux.
  • Deploy and evolve observability capabilities across infrastructure and Kubernetes environments, including metrics, logging, alerting, dashboards, and incident diagnostics.
  • Support and optimize platform and application workloads on Kubernetes, improving deployment patterns, scaling behavior, runtime efficiency, resilience, and day-to-day operational support.
  • Strengthen security across infrastructure, pipelines, and Kubernetes environments through IAM least-privilege access, secrets management, image and artifact scanning, admission controls, policy-as-code, workload isolation, and runtime hardening.
  • Support the reliability, backup integrity, availability, and operational performance of PostgreSQL/RDS environments in partnership with application teams.
  • Work closely with development and platform teams to improve deployment strategies, runtime reliability, developer experience, and operational standards for cloud-native systems.
  • Monitor system health, investigate incidents, troubleshoot infrastructure and application issues, and drive timely resolution through strong root cause analysis and preventative improvements.
  • Lead technical decision-making within the scope of the role by prioritizing work, evaluating tradeoffs, integrating stakeholder input, and driving issues through to completion with limited oversight.
  • Use modern AI tools to accelerate infrastructure design exploration, Terraform authoring, Kubernetes troubleshooting, CI/CD workflow development, observability analysis, and operational documentation while rigorously validating outputs for correctness, security, maintainability, and production readiness.
  • Document and continually refine DevOps methodologies, infrastructure standards, deployment workflows, operational procedures, and support runbooks.

Benefits

  • W2 contract assignment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service