Senior Site Reliability Engineer - FedRAMP

F5•Reston, OR
•$161,900 - $242,900•Onsite

About The Position

We are seeking an experienced, security-focused Senior Site Reliability Engineer (Senior SRE) to drive the reliability, architectural design, and continuous compliance of our FedRAMP-authorized cloud platform. In this senior role, you will combine hands-on operational leadership with infrastructure architecture, technical governance, and audit readiness across AWS (Commercial and GovCloud) and Kubernetes environments. As a Senior SRE, you will design and operate highly resilient, multi-cluster Amazon EKS infrastructure, direct observability architectures using Prometheus and Grafana, maintain automated deployment pipelines using GitLab CI/CD, and lead technical response in a 24/7 rotational on-call schedule. You will also serve as a key technical liaison for Third-Party Assessment Organization (3PAO) FedRAMP audits, ensuring strict adherence to NIST SP 800-53 controls, architectural security patterns, and vulnerability management SLAs.

Requirements

  • This position will require you to be a US Citizen residing in the United States.
  • 5+ years of production experience operating high-availability systems in a Site Reliability Engineering, DevOps, or Systems Architecture role.
  • L3 Production Escalation: Proven track record of handling high-pressure L3 production escalations, driving live incident triage, and managing stakeholder communication.
  • Logging & Deep Troubleshooting: Advanced hands-on experience using Elasticsearch and Kibana to troubleshoot critical system events, write complex queries (KQL), parse logs, and construct operational dashboards.
  • FedRAMP & Audit Expertise: Direct experience implementing, maintaining, and defending systems during FedRAMP audits (Moderate or High), NIST SP 800-53, or SOC 2 Type II audits.
  • Architectural Design Skills: Strong capability in designing hybrid cloud network architectures, secure multi-tenant clusters, and formulating system security plans (SSPs).
  • Deep Linux Internals: Strong proficiency in Linux administration (RHEL, Rocky Linux, or hardened Ubuntu), networking (TCP/IP, iptables, DNS), and OS hardening.
  • Kubernetes & AWS EKS: Expert-level mastery of Kubernetes orchestration, pod autoscaling, ingress management, and cloud infrastructure on AWS/GovCloud.
  • Observability Expertise: Hands-on experience configuring and scaling Prometheus, writing PromQL queries, building Grafana dashboards, and managing Alertmanager rules connected to Slack.
  • IaC & GitOps: Proven track record implementing Terraform at scale and continuous deployment with ArgoCD and GitLab CI.
  • Software Development: Proficiency in writing production-grade automation scripts and tools in Go (Golang) and/or Python.
  • Edge & Network Engineering: Practical experience working with on-premises edge servers and configuring hardware/software load balancers (e.g., F5 BIG-IP, HAProxy, NGINX, AWS ALB/NLB) in secure zones.
  • Communication & Documentation: Clear, structured written communication skills with a proven habit of writing detailed runbooks and compliance-aligned documentation.

Responsibilities

  • Act as the final technical escalation tier (L3 support) for critical platform and production incidents, troubleshooting complex issues escalated by L1/L2 operations or customer support teams.
  • Lead high-priority incident response bridges, coordinating across development, security, and networking teams to drive rapid issue containment and resolution under strict service-level agreements (SLAs).
  • Troubleshoot deep, transient infrastructure errors (e.g., routing, network package drops, Kubernetes control-plane failures) that go beyond standard SOPs.
  • Establish clear, documented escalation pathways, translating complex L3 resolutions into actionable runbooks and empower L1/L2 teams.
  • Implement and enforce security controls aligned with the NIST SP 800-53 framework to achieve and maintain our FedRAMP Authorization to Operate (ATO).
  • Act as the technical lead for SRE during annual FedRAMP 3PAO audits, gathering evidence, demonstrating compliance, and proving operational control implementation.
  • Build and maintain secure, tamper-proof audit logging pipelines—forwarding application logs, system logs, API call records, and Kubernetes audit trails securely into Elasticsearch (and central SIEMs) with strict, compliant retention and index-lifecycle management policies.
  • Utilize Elasticsearch and Kibana as primary investigative tools to perform deep-dive troubleshooting of complex, distributed system anomalies and application errors across AWS and edge environments.
  • Build, customize, and curate high-signal Kibana dashboards, search queries (KQL/Lucene), and visualizations to provide real-time operational visibility and dramatically reduce Mean Time to Resolution (MTTR).
  • Troubleshoot log-ingestion pipelines (Vector, Fluentd) to resolve bottlenecks, parsing errors, or missing metadata in high-volume production environments.
  • Design and document highly resilient, secure-by-default architectural topologies for hybrid networks spanning AWS GovCloud and on-premises edge servers.
  • Architect and scale multi-tenant Kubernetes (EKS) clusters enforcing strict physical or logical boundaries, zero-trust network policies, and identity federation.
  • Configure and optimize high-availability L4/L7 load balancer networking, ingress controllers, and FIPS 140-3 validated traffic routing for secure, low-latency edge-to-cloud communication.
  • Manage physical bare-metal edge nodes, defining operating system hardening baselines (STIG compliance), secure boot processes, and automated physical host provisioning.
  • Author clean, modular, and secure Terraform code to provision cloud infrastructure, networking topologies, and security boundaries.
  • Build declarative deployment pipelines using GitLab CI/CD and ArgoCD to achieve true GitOps-driven delivery, ensuring all code modifications are traceable, signed, and fully audited.
  • Develop custom automation tools, controllers, and CLI utilities in Python or Go (Golang) to eliminate repetitive toil and automate continuous compliance reporting.
  • Design end-to-end monitoring and dashboarding systems utilizing Prometheus, Grafana, and Alertmanager.
  • Build actionable alerting pipelines integrated with Slack, ensuring notifications are high-signal and low-noise.
  • Facilitate blameless post-mortem reviews following high-severity incidents, utilizing evidence harvested from Kibana logs to document timelines, determine root causes, and programmatically prevent recurrences.

Benefits

  • incentive compensation
  • bonus
  • restricted stock units
  • benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service