Observability Engineer

Agile Business ConceptsFort Meade, MD
Onsite

About The Position

Observability Engineer Location: Fort Meade, MD (or as required by contract) About Agile Business Concepts Agile Business Concepts (ABC) is a leading provider of cybersecurity, information technology, engineering, and mission support services to Federal Government agencies. We deliver innovative, secure, and scalable technology solutions that strengthen our customers' operational readiness and mission success. Our team is committed to excellence, integrity, collaboration, and continuous innovation while supporting some of the nation's most critical national security initiatives. We are seeking an experienced Observability Engineer to join our growing team supporting a high-profile Federal customer. The successful candidate will play a key role in designing, implementing, and maintaining enterprise observability and monitoring capabilities that improve system performance, operational visibility, and mission resilience across complex production environments. Position Summary The Observability Engineer will be responsible for designing, implementing, and supporting enterprise monitoring and observability solutions for mission-critical systems. This role focuses on improving system visibility, reliability, and operational performance through enterprise logging, telemetry collection, dashboard development, alerting, and advanced analytics. The ideal candidate possesses experience supporting large-scale enterprise environments and collaborates effectively with DevOps, Site Reliability Engineering (SRE), Platform Engineering, Cybersecurity, and Infrastructure teams to enhance monitoring coverage, automate operational processes, and improve overall system health.

Requirements

  • Active Top Secret / Sensitive Compartmented Information (TS/SCI) security clearance.
  • Bachelor's degree in Computer Science, Information Technology, Cybersecurity, Systems Engineering, or a related technical discipline (or equivalent experience).
  • Minimum 8 years of experience supporting enterprise observability, monitoring, Site Reliability Engineering (SRE), Platform Engineering, Systems Engineering, Cybersecurity Engineering, or related technical environments.
  • Demonstrated experience implementing enterprise monitoring and observability platforms.
  • Experience with: Enterprise log management, Telemetry collection, Dashboard development, Alert configuration, Event correlation, Performance monitoring
  • Strong understanding of enterprise infrastructure monitoring, including: Compute, Storage, Networking, Operating Systems, Virtualization, Applications, Databases
  • Experience performing Root Cause Analysis (RCA) and supporting production incident response.
  • Experience working within Agile software development environments.
  • Ability to obtain SAFe Scrum Master Certification if not currently certified.
  • Excellent written, verbal, analytical, and problem-solving skills.

Nice To Haves

  • Experience with one or more of the following technologies: OpenObserve, Splunk Enterprise, Splunk Observability Cloud, Elastic / ELK Stack, Grafana, Prometheus, Datadog, Dynatrace, Fluentd, Fluent Bit, OpenTelemetry, Kibana, Logstash
  • Cloud monitoring (AWS, Azure, or Google Cloud)
  • Kubernetes and container observability
  • Infrastructure as Code (Terraform, Ansible)
  • CI/CD pipelines (Jenkins, GitLab, GitHub Actions)
  • API integrations and monitoring automation
  • Security Information and Event Management (SIEM)
  • Cybersecurity monitoring and security telemetry
  • Artificial Intelligence for IT Operations (AIOps)
  • Machine learning-assisted monitoring and anomaly detection
  • Python, PowerShell, or Bash scripting

Responsibilities

  • Design, implement, configure, and maintain enterprise observability and monitoring solutions.
  • Develop and optimize log collection, ingestion, normalization, correlation, and analysis across multiple enterprise data sources.
  • Build dashboards, visualizations, alerts, and performance reports to monitor system health, availability, and operational performance.
  • Analyze infrastructure, application, network, and cloud telemetry to proactively identify performance issues, anomalies, and operational trends.
  • Perform root cause analysis (RCA) for production incidents and recommend corrective actions to improve system reliability.
  • Collaborate with Engineering, DevOps, Site Reliability Engineering (SRE), Platform Engineering, Infrastructure, and Cybersecurity teams to improve enterprise monitoring and operational visibility.
  • Support Service Level Agreements (SLAs) and operational performance objectives through proactive monitoring, alerting, and automation.
  • Ensure monitoring data is properly collected, tagged, normalized, retained, and correlated across enterprise environments.
  • Develop monitoring standards, best practices, and operational documentation.
  • Automate monitoring workflows and operational processes through scripting and API integrations.
  • Support cloud-based monitoring initiatives across hybrid and multi-cloud environments.
  • Mentor junior engineers and provide technical leadership on enterprise observability initiatives.
  • Participate in Agile ceremonies including sprint planning, backlog grooming, daily stand-ups, sprint reviews, and retrospectives.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service