Senior Monitoring and Observability Engineer

Peraton,
$104,000 - $166,000Remote

About The Position

Peraton is seeking a Monitoring and Observability Engineer with a strong observability background to build and operate the telemetry, monitoring, and alerting capability supporting an enterprise platform running in AWS and AWS GovCloud within a FedRAMP-authorized boundary. This is a hands-on engineering role, reporting to the Senior SRE / Observability Lead, focused on instrumentation, dashboarding, alert quality, and telemetry pipeline reliability across metrics, logs, and traces. This person will work closely with platform engineering and the broader SRE team to ensure every service is properly instrumented, every alert is actionable, and incident responders have the data they need to diagnose issues quickly — while keeping telemetry cost and cardinality under control.

Requirements

  • Must be a U.S. Citizen with the ability to obtain and maintain the required Public Trust level clearance.
  • Bachelor's Degree and 8 years of experience, a Master's Degree and 6 years of experience, or a High School diploma or equivalent and 12 years.
  • 4+ years hands-on experience with enterprise observability, monitoring, and alerting platforms.
  • Hands-on experience with Dynatrace, Datadog, and/or Splunk, including instrumentation, dashboarding, and alert design.
  • Working knowledge of metrics, logs, and traces, and experience with OpenTelemetry-based instrumentation.
  • Experience operating in AWS and/or AWS GovCloud; familiarity with containerized and cloud-native workloads.
  • Experience with Ansible / Ansible Tower / AAP for automated agent and configuration deployment.
  • Experience with GitLab CI/CD (and Jenkins) for building observability-as-code pipelines.
  • Scripting/automation proficiency in Python, Bash, or Go.

Nice To Haves

  • Observability platform certifications (Dynatrace, Datadog, or Splunk) or cloud architecture certifications.
  • Production experience with OpenTelemetry Collector deployment at scale.
  • Kubernetes and container observability, service-mesh telemetry, or eBPF-based collection experience.
  • Experience with AIOps or AI-assisted anomaly detection, event correlation, or incident summarization.
  • Experience in federal or regulated environments (FISMA, FedRAMP, NIST 800-53).

Responsibilities

  • Build and maintain telemetry pipelines end to end — collection, enrichment, routing, sampling, storage, and retention — across Dynatrace, Datadog, and Splunk.
  • Instrument applications and infrastructure for metrics, logs, and traces; support OpenTelemetry-based, vendor-neutral instrumentation standards.
  • Design observability for distributed systems: microservices, containers, and cloud-native workloads in AWS/GovCloud.
  • Build dashboards and golden-signal monitoring aligned to service criticality; enable correlation across metrics, events, logs, and traces so anomalies can be diagnosed without manual pivoting.
  • Monitor the health of the observability platforms themselves and remediate collector, agent, and pipeline issues.
  • Build and tune alert definitions with clear ownership, severity, and runbook links for every production alert; reduce false positives and alert flapping.
  • Support on-call rotations and participate in incident bridges, providing telemetry-driven triage and root-cause data.
  • Integrate alert routing, deduplication, and maintenance-window handling with ServiceNow and paging tools.
  • Contribute to post-incident reviews with detection-source analysis and MTTD/MTTA trend reporting.
  • Deploy and configure observability agents/collectors using Ansible, Ansible Tower, and AAP for consistent, automated rollout across environments.
  • Build observability-as-code pipelines in GitLab (GitLab CI/CD), including automated dashboard, alert, and collector-configuration deployment; support migration of remaining Jenkins jobs to GitLab.
  • Partner with platform engineering to instrument new Terraform-provisioned infrastructure as part of standard build patterns.
  • Advance configuration-as-code for pipelines, dashboards, and alert definitions so changes are versioned and peer-reviewed.
  • Support tagging conventions, dashboard/alert lifecycle management, and retention-tier governance.
  • Help manage telemetry cost, ingest volume, and cardinality growth across the observability platforms.
  • Support sensitive-data masking, audit, and continuous-monitoring requirements for observability data.
  • Maintain observability architecture documentation and runbooks; operate within SAFe using ServiceNow, Jira, and Confluence.

Benefits

  • Employees may be eligible for overtime, shift differential, and a discretionary bonus in addition to base pay.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service