The position exists to close a known capability gap in log correlation, SLO and error-budget reporting, and early-warning detection. Success in this role directly protects the performance-service-level track record AFS markets to clients and reduces the time it takes to detect and resolve production issues across the client portfolio. The Observability Engineer owns the observability capability end to end across AFSVision HPC environments, including tool evaluation, log correlation architecture, and SLO/SLI definition. The role has authority over day-to-day monitoring and observability engineering decisions, and escalates architectural changes, tooling investments, and cross-team resourcing needs to the Managing Director, System Services Group. This position partners closely with Application Services, Database Administration, Network Engineering, and Information Security, and must design and operate all observability tooling within AFS’s data classification policy and applicable FFIEC, SOC 2, and NIST requirements, since the underlying telemetry can include client production data. As AFS’s Next Gen Datacenter project moves workloads toward public cloud infrastructure, this scope extends to CI/CD pipeline and Kubernetes observability alongside the existing on-premises HPC environment.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level