Senior Site Reliability Engineer, GovCloud

MedalliaMclean, VA
Hybrid

About The Position

Medallia is seeking a Senior Site Reliability Engineer to join its growing GovCloud team. This role will focus on operating and improving Medallia’s US public-sector cloud platform, built on AWS GovCloud and Kubernetes, to support federal agencies and other regulated customers in a highly available, secure, and compliant environment. The position is hybrid, based near Tysons, Virginia, requiring regular in-office collaboration with remote flexibility. The engineer will be hands-on, owning production systems, collaborating with engineering and security teams, and ensuring platform reliability during scaling. Medallia is the market leader in Experience Management, utilizing its SaaS platform, Medallia Experience Cloud, to manage experiences, insights, and actions across various sectors. The company's mission is to create a world where organizations are loved by their customers and employees, empowering exceptional people to create extraordinary experiences together.

Requirements

  • Must reside in the United States and be legally authorized to work in the US without sponsorship.
  • Bachelor’s degree or equivalent experience in Computer Science or a related field.
  • 5+ years of experience in Site Reliability Engineering, platform engineering, DevOps, or related production infrastructure roles.
  • Production experience with Kubernetes
  • Production experience with Core services (IAM, compute, object storage, encryption/key management) and cloud networking on AWS (strongly preferred), Google Cloud (GCP), Azure, or a similar public cloud platform.
  • Production experience with Terraform or comparable infrastructure-as-code tools
  • Production experience with Git and CI/CD pipelines
  • Production experience with Linux and foundational systems concepts (networking, DNS, TLS/certificates)
  • Production experience with PostgreSQL (or comparable relational databases) in production — replication, backups, and performance tuning
  • Proficiency in Python and/or Go experience to build automation scripts, operational tooling, and infrastructure services.
  • Experience troubleshooting production incidents, conducting root-cause analysis(RCA), and following change management processes.
  • Experience participating in a production on-call rotation.
  • Experience troubleshooting complex technical issues and writing clear documentation, runbooks and incident post-mortems.

Nice To Haves

  • Experience operating in FedRAMP, AWS GovCloud, or other regulated or compliance-heavy cloud environments.
  • Familiarity with security and compliance practices such as FIPS and vulnerability management.
  • Experience with observability and logging platforms in enterprise production environments.
  • Deep operational expertise with PostgreSQL (HA/replication, tuning, backup and recovery); familiarity with Redis and Kafka.
  • Experience supporting federal agencies or public-sector customers.
  • Experience with tools such as Jenkins, Argo CD, and GitHub Enterprise.
  • Strong collaboration skills and willingness to learn in a compliance-driven environment.

Responsibilities

  • Build, operate, and improve highly available, secure cloud infrastructure on AWS, including networking, access management, Kubernetes clusters, DNS, certificates, and shared platform services.
  • Design and operate AWS cloud networking end-to-end — VPC architecture, subnetting and routing, security groups/NACLs, VPC endpoints/PrivateLink, Transit Gateway, load balancing, and DNS — for secure, segmented, highly available connectivity.
  • Operate and tune production PostgreSQL — high availability and replication, backups and recovery, query and performance optimization, version upgrades, and capacity planning — as part of the platform's data tier.
  • Monitor production systems, respond to incidents, and drive fixes that improve reliability and reduce repeat issues.
  • Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and GitOps practices.
  • Improve observability across metrics, logs, and uptime monitoring; help tune alerts and operational runbooks.
  • Work with software engineering, security, and release management teams to deploy changes safely and resolve production issues.
  • Contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment.
  • Participate in an on-call rotation for production support.
  • Document systems and operational procedures clearly.
  • Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries.

Benefits

  • competitive health and wellness benefits, including but not limited to medical, dental, vision, 401(k), short-term and long-term disability, life and AD&D insurance, statutory leaves, paid parental leave, and paid holidays.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service