Observability Platform Architect

Booz Allen HamiltonMcLean, VA
$86,800 - $198,000Remote

About The Position

The Opportunity: Are you looking for an opportunity to combine hands-on engineering with big picture thinking at the point where AI stops being a demonstration and starts running production operations? You understand that autonomous operations is not a tooling problem but a trust problem — what a system is permitted to do on its own, and how you prove that decision was sound. As an observability platform architect on our team, you'll set the technical direction for how the enterprise monitors its business services across multiple portfolios and separately accredited environments. Your customers will trust you not only to design the platform, but to build it alongside the team — owning instrumentation strategy and the entity model, automating an estate this size so it stays manageable by a team this size, and defining what agentic detection, triage, and remediation are permitted to do unattended. On our team, you'll broaden into Dynatrace, OpenTelemetry, and agentic operations in environments most engineers never get the opportunity to work in. This role sets the standards and then builds against them alongside the team, so you'll stay in the code. The ideal candidate comes from a Platform Engineering, Observability, Monitoring Infrastructure, or Site Reliability Engineering background with genuine ownership of an observability platform: its architecture, standards, and operation across an organization. Join us. The world can't wait.

Requirements

  • 8+ years of experience with engineering experience
  • 3+ years of experience maintaining responsibility for an enterprise observability or APM platform itself, including its architecture, deployment, upgrades, standards, and cost across an organization
  • Experience with an enterprise observability platform, such as Dynatrace, Datadog, New Relic, Grafana, Prometheus, or Elastic, including writing and optimizing complex queries against its data
  • Experience replacing recurring manual platform work with production automation, and shipping a system that makes decisions automatically, including agentic operations, ML in production, fraud or risk scoring, or security detection, including owning whether those decisions stayed correct over time
  • Experience with OpenTelemetry, including collectors, instrumentation, semantic conventions, and the practical limits of vendor-neutral instrumentation
  • Experience with infrastructure as code and CI/CD, such as Terraform, GitHub Actions, Azure DevOps, or GitLab, and deploying agents or collectors at scale across Windows, Linux, Kubernetes, cloud platforms, and network infrastructure
  • Ability to write, review, and ship code
  • Secret clearance
  • Bachelor's degree in CS, Information Systems, or Engineering

Nice To Haves

  • Experience with Dynatrace in depth, including Grail, DQL, Workflows, Site Reliability Guardian, Davis, and Dashboards, with configuration managed through the Terraform provider
  • Experience building automated remediation in production, including an automation that acted wrongly and what changed afterward
  • Experience with multiple observability platforms, such as Grafana, Prometheus, Datadog, or Elastic
  • Experience operating observability in DoD cloud environments such as Azure or AWS IL5/IL6, including working around feature gaps between commercial and authorized offerings
  • Experience operating one platform across more than one environment or security boundary, and managing the drift and duplication that creates
  • Knowledge of what observability platforms operating models costs to run
  • ITIL 4 Foundation certification

Responsibilities

  • Architect and build the observability platform across every accredited environment, and keep them from drifting into separately maintained estates.
  • Drive OpenTelemetry adoption hands-on, including collector configuration, instrumentation, and rollout, and set the standards, naming conventions, and onboarding patterns the team builds against.
  • Automate the platform itself, including service onboarding, agent deployment, tagging, dashboard and alert provisioning, and access management, delivered as code through CI/CD.
  • Design and deploy agentic detection, triage, and remediation, defining what an agent may do unattended, how we verify the service actually recovered, how it rolls back, and what evidence it leaves behind.
  • Own the platform's AI capability end to end, including what is enabled, what it is trusted to decide, how its accuracy holds up over time, and how incorrect actions get caught.
  • Build the entity model, dashboards, and SLOs that leadership depends on, and drive down cost per monitored service.

Benefits

  • health, life, disability, financial, and retirement benefits
  • paid leave
  • professional development
  • tuition assistance
  • work-life programs
  • dependent care
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service