Observability Engineer

AFS•Exton, PA
•Onsite

About The Position

The position exists to close a known capability gap in log correlation, SLO and error-budget reporting, and early-warning detection. Success in this role directly protects the performance-service-level track record AFS markets to clients and reduces the time it takes to detect and resolve production issues across the client portfolio. The Observability Engineer owns the observability capability end to end across AFSVision HPC environments, including tool evaluation, log correlation architecture, and SLO/SLI definition. The role has authority over day-to-day monitoring and observability engineering decisions, and escalates architectural changes, tooling investments, and cross-team resourcing needs to the Managing Director, System Services Group. This position partners closely with Application Services, Database Administration, Network Engineering, and Information Security, and must design and operate all observability tooling within AFS’s data classification policy and applicable FFIEC, SOC 2, and NIST requirements, since the underlying telemetry can include client production data. As AFS’s Next Gen Datacenter project moves workloads toward public cloud infrastructure, this scope extends to CI/CD pipeline and Kubernetes observability alongside the existing on-premises HPC environment.

Requirements

  • Demonstrated experience standing up log correlation and SLO-based reliability practices matters more than years in a title.
  • Experience designing and operating log correlation platforms at production scale, ideally across Linux, middleware, and application-tier logs.
  • Background maintaining enterprise monitoring tools (SolarWinds, ManageEngine, Site24x7, or similar), including patching, compliance, and dashboard configuration.
  • Demonstrated ability to build SLO/SLI frameworks and error-budget reporting tied to contractual SLAs.
  • Experience producing client-facing performance and SLA reporting.
  • Strong Linux systems background, with experience monitoring middleware and application logs for correlation and root cause analysis.
  • Cloud infrastructure experience (Azure preferred) and strong log analysis and correlation skills.
  • Comfort operating in a regulated, audited environment; able to work within FFIEC, SOC 2, and NIST constraints.
  • Track record of reducing MTTD/MTTR through improved detection, alerting, and incident response practices.
  • Bachelor’s degree in Computer Science, Information Systems, or a related field, or equivalent hands-on experience.

Nice To Haves

  • Exposure to public cloud CI/CD pipelines and Kubernetes monitoring is a plus, as AFS’s Next Gen Datacenter project migrates workloads off-premises.

Responsibilities

  • Build log correlation capability – design and stand up correlated, searchable log correlation across Linux, middleware, and AFSVision application logs, replacing ad-hoc, siloed log search.
  • Deliver SLO/SLI reporting – convert raw monitoring data into SLO/SLI and error-budget reporting tied directly to client SLAs, giving Operations and clients defensible reliability metrics.
  • Publish client-facing SLA reports – produce accurate, recurring performance and SLA reports for clients, reflecting measured system metrics.
  • Implement early-warning indicators – for the conditions that most often drive incidents, including batch run-time drift, resource saturation, and connection-tier health across the Linux and middleware tiers supporting AFSVision.
  • Lead the observability roadmap – evaluate and deploy tracing and synthetic monitoring inside the HPC trust boundary under FFIEC, SOC 2, and NIST controls, and extend that roadmap to the public cloud CI/CD pipeline and Kubernetes environments introduced by AFS’s Next Gen Datacenter project.
  • Maintain the core monitoring stack – keep SolarWinds, ManageEngine, and Site24x7 patched, compliant, and configured with accurate per-environment monitors and SLA dashboards.
  • Reduce alert noise – measurably cut MTTD and MTTR, lowering after-hours escalations and incident cost.
  • Partner across teams – work with Application Services, Database Administration, Network Engineering, and Information Security to correlate signals across the full stack during incident investigation.
  • Report on reliability – build and maintain dashboards and reliability reporting reviewed on a regular cadence with Operations and System Services Group leadership.
  • Support incident response – participate in on-call rotation and lead root cause analysis for production incidents involving monitoring or observability gaps.
  • Document standards – maintain observability standards, runbooks, and tool configuration documentation for adoption across the System Services Group, including monitoring practices for the Next Gen Datacenter’s cloud and Kubernetes environments as they come online.
  • Evaluate new tooling – pilot new observability and monitoring tools and present cost/benefit recommendations to leadership.
  • Maintain compliance – ensure all observability practices and tooling comply with AFS data classification policy and FFIEC, SOC 2, and NIST requirements.

Benefits

  • Flexible schedule options are available in accordance with AFS company policy.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service