Site Reliability Engineering (SRE) - Maryland, US

NTT DATA ServicesBaltimore, MD
$87,720 - $109,650Onsite

About The Position

We are seeking a Senior Site Reliability Engineer (SRE) to design, build, and support scalable, secure, and highly available cloud platforms. This role combines SRE, cloud engineering, platform engineering, and automation to improve reliability, performance, and developer productivity across enterprise applications. The position supports critical UPS.com services and participates in a 24x7 on-call rotation.

Requirements

  • 5+ years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or related roles.
  • 3+ years of hands-on Akamai experience, including CDN, WAF, and edge security services.
  • Strong experience with Kubernetes administration and containerized platforms.
  • Experience with GCP, Azure, Terraform, and Infrastructure as Code.
  • Experience building and supporting CI/CD and GitOps solutions.
  • Proficiency in Python, Bash, Shell scripting, or similar automation languages.
  • Experience with monitoring, observability, and incident response practices.
  • Strong troubleshooting, communication, and problem-solving skills.
  • Ability to lead initiatives and mentor engineers.
  • Must be able to support EST working hours and participate in a 24x7 on-call rotation.

Nice To Haves

  • GCP, GKE, OpenShift (OCP4)
  • Terraform, Google Config Connector
  • Azure Pipelines, Argo CD, Argo Workflows
  • Python, Shell Scripting, Ansible, AWX/Tower
  • Docker, Linux Administration
  • Istio, Kiali, Gatekeeper
  • Dynatrace, Grafana, Loki, Prometheus
  • HashiCorp Vault, SOPS, External Secrets Operator, Cert-Manager
  • Apache Kafka, Confluent Kafka, Strimzi
  • Apigee, WebMethods, Camunda
  • Google BigQuery, Firestore, Bigtable, AlloyDB, Pub/Sub
  • Apache Solr, Zookeeper
  • PagerDuty
  • Akamai Security Products and DigiCert Certificate Management

Responsibilities

  • Support and optimize mission-critical UPS.com applications and services.
  • Provide production support for Akamai CDN, WAF, bot management, and edge security solutions.
  • Manage certificate lifecycle processes, including DigiCert integrations, renewals, and automation.
  • Implement and maintain Akamai security policies, malware protection, and performance optimizations.
  • Design and operate highly available cloud infrastructure in GCP and Azure.
  • Drive SRE practices including SLOs, SLIs, incident management, root cause analysis, and reliability improvements.
  • Build and maintain Kubernetes platforms (GKE/OpenShift) and cloud-native services.
  • Develop Infrastructure as Code solutions using Terraform and Config Connector.
  • Implement CI/CD and GitOps workflows using Azure Pipelines and Argo CD.
  • Enhance observability through monitoring, logging, and alerting solutions (Dynatrace, Grafana, Prometheus).
  • Automate operational processes using Python, Shell scripting, Ansible, and AI-driven solutions.
  • Implement security best practices, including secrets management and policy enforcement.
  • Collaborate with cross-functional teams to resolve incidents and improve platform stability.

Benefits

  • medical, dental, and vision insurance with an employer contribution
  • flexible spending or health savings account
  • life and AD&D insurance
  • short and long term disability coverage
  • paid time off
  • employee assistance
  • participation in a 401k program with company match
  • additional voluntary or legally-required benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service