Site Reliability Engineer II, GovCloud

MedalliaMclean, VA
Hybrid

About The Position

Medallia is seeking a Site Reliability Engineer II to join its growing GovCloud team. This role will focus on operating and improving Medallia’s US public-sector cloud platform, built on AWS GovCloud and Kubernetes, to support federal agencies and other regulated customers in a highly available, secure, and compliant environment. The position is hybrid, based near Tysons, Virginia, offering a blend of in-office collaboration and remote flexibility. The engineer will be a hands-on contributor to production systems, working closely with engineering and security teams, with opportunities for growth in owning larger systems and ensuring platform reliability at scale.

Requirements

  • Must reside in the United States and be legally authorized to work in the US without sponsorship.
  • Bachelor’s degree or equivalent experience in Computer Science or a related field.
  • 2+ years of experience in Site Reliability Engineering, platform engineering, DevOps, or related production infrastructure roles (or equivalent hands-on production experience).
  • Production experience with: Core services (IAM, compute, object storage, encryption/key management) and cloud networking on AWS (strongly preferred), Google Cloud (CGP), Azure, or a similar public cloud platform
  • Production experience with Terraform or comparable infrastructure-as-code tools
  • Production experience with Git and CI/CD pipelines
  • Production experience with Linux and foundational systems concepts (networking, DNS, TLS/certificates)
  • Production experience with PostgreSQL or another relational database — basic operations (queries, backups)
  • Familiarity with Kubernetes concepts, container orchestration, and microservices management
  • Proficiency in Python and/or Go experience to build automation scripts, operational tooling, and infrastructure services.
  • Experience troubleshooting production incidents, conducting root-cause analysis(RCA), and following change management processes.
  • Experience participating in a production on-call rotation.
  • Experience troubleshooting complex technical issues and writing clear documentation, runbooks and incident post-mortems.

Nice To Haves

  • Experience operating in FedRAMP, AWS GovCloud, or other regulated or compliance-heavy cloud environments.
  • Familiarity with security and compliance practices such as FIPS and vulnerability management.
  • Experience with observability and logging platforms in enterprise production environments.
  • Deep operational expertise with PostgreSQL (HA/replication, tuning, backup and recovery); familiarity with Redis and Kafka.
  • Experience supporting federal agencies or public-sector customers.
  • Experience with tools such as Jenkins, Argo CD, and GitHub Enterprise.
  • Strong collaboration skills and willingness to learn in a compliance-driven environment.

Responsibilities

  • Build, operate, and improve highly available, secure cloud infrastructure on AWS, including networking, access management, Kubernetes clusters, DNS, certificates, and shared platform services.
  • Implement, operate and optimize AWS cloud networking, specifically managing VPCs, subnets and routing, security groups/NACLs, VPC endpoints/PrivateLink and load balancing.
  • Monitor, maintain and support production PostgreSQL — replication, backups and recovery, routine performance tuning, and upgrades — as part of the platform's data tier.
  • Monitor production systems, respond to incidents, and drive fixes that improve reliability and reduce repeat issues.
  • Develop and maintain Infrastructure-as-Code (primarily Terraform) and Kubernetes deployment workflows using Git, CI/CD, and GitOps practices.
  • Improve observability across metrics, logs, and uptime monitoring; help tune alerts and operational runbooks.
  • Partner with software engineering, security, and release management teams to deploy changes safely and resolve production issues.
  • Contribute to platform upgrades, security patching, and compliance-driven maintenance in a regulated cloud environment.
  • Participate in an on-call rotation for production support.
  • Use AI-assisted tooling responsibly, with attention to security, privacy, and customer data boundaries.
  • Learn the platform and grow your scope with mentorship from senior engineers.

Benefits

  • medical
  • dental
  • vision
  • 401(k)
  • short-term and long-term disability
  • life and AD&D insurance
  • statutory leaves
  • paid parental leave
  • paid holidays
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service