DevOps & AI/ML Infrastructure Engineer

CreatorIQ•New York, NY
•$106,000 - $127,000•Hybrid

About The Position

The DevOps Engineer is responsible for supporting and improving cloud and ML/AI infrastructure, automating deployments, and maintaining CI/CD pipelines to ensure efficient, secure, and scalable development workflows. This role plays a crucial part in infrastructure automation, monitoring, and cloud security while collaborating with Software and ML Engineers, Product support, QA, and Security teams. As a key member of the DevOps team, the DevOps Engineer helps manage cloud environments, CI/CD pipelines, and Infrastructure as Code (IaC), ensuring high availability and compliance with security best practices.

Requirements

  • 3+ years of experience in DevOps, Cloud Engineering, Site Reliability Engineering (SRE), or a similar infrastructure-focused role.
  • 2+ years of hands-on experience with AWS services such as EC2, S3, RDS, Lambda, IAM, VPC, SQS, API Gateway, or similar services.
  • 2+ years of experience working with containerized environments and orchestration platforms such as Kubernetes and Amazon EKS.
  • Strong experience building and maintaining CI/CD pipelines using tools such as GitLab CI/CD or Jenkins.
  • Hands-on experience with Infrastructure as Code using Terraform, Terragrunt, CloudFormation, or similar technologies.
  • Strong Linux system administration and troubleshooting skills.
  • Solid understanding of networking fundamentals, including routing, load balancing, network security, and related concepts.
  • Scripting experience with Python, Bash, or similar languages to automate infrastructure and operational tasks.
  • Hands-on experience using AI tools to improve engineering workflows, automation, troubleshooting, or agentic use cases.
  • Experience supporting data, ML, or other compute-intensive production workloads.

Nice To Haves

  • Experience with Google Cloud would be valuable, particularly for candidates who have worked across multi-cloud environments.
  • Familiarity with Helm and service mesh technologies such as Istio, Linkerd, Traefik, or similar tools would be beneficial.
  • Experience with serverless and event-driven architectures using technologies such as AWS Lambda, API Gateway, and SQS is a plus.
  • Exposure to cloud and infrastructure security practices, including vulnerability management and tools such as Nessus, Prowler, Trivy, firewalls, or similar technologies, would be valuable.
  • Knowledge of security standards, compliance requirements, and cloud security best practices is beneficial.
  • Experience with observability, log analysis, and monitoring platforms such as Coralogix, Prometheus, Grafana, or similar solutions is a plus.
  • FinOps experience, including cloud cost monitoring, optimization, and accountability practices, would be valuable.
  • Experience with API gateways or API management platforms such as Kong, Apigee, or similar technologies is beneficial.
  • Experience with MLOps platforms and practices—particularly Databricks, model serving, ML pipelines, and model monitoring—would be an advantage.

Responsibilities

  • Support and maintain scalable, highly available, and secure cloud infrastructure in accordance with company policies and standards.
  • Provision and manage cloud resources using Infrastructure as Code (Terraform, Terragrunt, CloudFormation).
  • Implement cloud security best practices, including IAM/role-based access controls, encryption, vulnerability management, and secure infrastructure configurations.
  • Support containerized environments and orchestration platforms.
  • Apply DevSecOps principles across infrastructure and deployment workflows.
  • Participate in disaster recovery planning, testing, and recovery activities.
  • Maintain and optimize CI/CD pipelines using tools such as GitLab CI/CD and Jenkins, supporting application and ML model deployments.
  • Improve deployment reliability and support zero-downtime deployment strategies.
  • Automate configuration management, infrastructure provisioning, and routine operational processes.
  • Troubleshoot deployment and pipeline issues and implement improvements to prevent recurrence.
  • Develop scripts and automation to reduce manual work and improve engineering efficiency.
  • Help design, deploy, operate, and secure infrastructure supporting AI and agentic products, including MCP, agents, integrations, internal tooling, and customer-facing use cases.
  • Use AI-assisted engineering tools, coding copilots, and AI-driven troubleshooting to improve DevOps productivity and reduce repetitive operational work.
  • Evaluate and adopt practical AI-enabled workflows that improve infrastructure management, troubleshooting, and operational efficiency.
  • Operate and scale ML platform infrastructure, including Databricks interactive clusters, jobs compute, ML pipelines, and Model Serving endpoints.
  • Manage production model-serving infrastructure, including compute capacity, provisioned throughput, and autoscaling for high-throughput inference workloads.
  • Maintain infrastructure-level monitoring for model drift, data quality, inference performance, and serving health, while partnering with ML Engineering on model evaluation, quality thresholds, and model correctness.
  • Partner with ML Engineering to support reliable CI/CD and production deployment of ML models.
  • Maintain monitoring, logging, metrics, and alerting solutions using tools such as Prometheus, Grafana, Coralogix, and CloudWatch.
  • Support incident response and perform Root Cause Analysis (RCA) for infrastructure and deployment-related issues.
  • Improve system observability through effective log aggregation, metrics collection, monitoring, and alerting.
  • Partner with Software Engineers, ML Engineers, QA, and Software Engineers in Test to improve deployment workflows and integrate automated testing into CI/CD pipelines.
  • Collaborate with IT Security to maintain secure cloud operations and infrastructure policies.
  • Respond to engineering and Product Support requests in a timely manner and provide technical infrastructure support when needed.
  • Maintain accurate internal technical and operational documentation.
  • Collaborate effectively with international teams across multiple time zones.

Benefits

  • 15 days of vacation
  • floating and company holidays
  • wellness benefits
  • paid parental leave
  • Comprehensive medical, dental, vision, life, and disability insurance
  • additional wellness benefits
  • A 401(k) plan
  • Work from home stipend
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service