Senior DevOps Engineer

Maven AGIBoston, MA
$165,000 - $190,000

About The Position

Maven AGI is seeking a Senior DevOps Engineer to be responsible for the design, building, and operation of production systems across various environments including cloud providers, Kubernetes clusters, and on-premises setups. The role involves managing CI/CD pipelines to ensure the platform's reliable scaling for enterprise customers with specific deployment and security needs. This is a critical position that directly influences platform availability, developer efficiency, and customer trust.

Requirements

  • 3-7 years of professional DevOps/SRE/Infrastructure experience
  • Deep expertise with Kubernetes in production (AKS, EKS, or GKE)
  • Strong infrastructure-as-code skills (Pulumi, Terraform, or Bicep)
  • Experience operating CI/CD systems (GitHub Actions, ArgoCD, or Jenkins)
  • Proficiency in at least one scripting/programming language (Python, Go, TypeScript, or Bash)
  • Solid understanding of IaaS providers, networking, DNS, load balancing, and TLS
  • Experience with monitoring and observability stacks (Datadog, Prometheus, Grafana, or similar)
  • Experience with multi-cloud or hybrid (cloud + on-prem) deployments
  • Strong communication and cross-team collaboration skills
  • Organized, great attention to detail, comfortable operating in a ticketing environment
  • Thrives in fast-paced startup environments

Nice To Haves

  • Experience with GPU infrastructure and ML/LLM serving workloads (vLLM, TEI)
  • Familiarity with Temporal or other workflow orchestration systems
  • Security and compliance background (SOC 2, HIPAA, GDPR)
  • Cost optimization experience at scale

Responsibilities

  • Design, implement, and maintain both cloud and on-premise infrastructure (Azure, AWS, datacenter) using infrastructure-as-code (Pulumi, Bicep, Terraform)
  • Own Kubernetes cluster operations: deployments, scaling, monitoring, and incident response
  • Build and optimize CI/CD pipelines for a large-scale monorepo
  • Implement observability across services (metrics, logging, tracing, alerting)
  • Drive reliability practices: SLOs, capacity planning, disaster recovery, and runbook development
  • Operationalize and scale enterprise AI deployments on-premise, including GPU resource orchestration, model inference performance tuning, and high-concurrency platform management.
  • Collaborate with engineering teams to improve developer experience and deployment velocity
  • Manage secrets, access controls, and infrastructure security posture
  • Evaluate and adopt new tooling to reduce operational toil

Benefits

  • Competitive salary
  • Comprehensive benefits
  • Meaningful equity stakes
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service