We are seeking a highly experienced and strategic Observability Lead-Sr. Infrastructure Engineer with a strong Forward Deployed Engineering (FDE), software development, event streaming, and technical leadership focus to define, implement, and scale enterprise observability capabilities. This role combines technical leadership in observability strategy with hands-on, customer-facing engineering, application development expertise, and Kafka/event-driven integration knowledge to deliver high-impact solutions across complex enterprise environments. As an Observability Lead, you will establish the vision and direction for metrics, logs, traces, telemetry pipelines, event streaming, and developer-enabled observability patterns using modern standards such as OpenTelemetry and a combination of open-source and commercial tooling. You will also operate as a forward deployed partner, working directly with engineering, SRE, platform teams, software development teams, and business stakeholders to solve real-world problems and deliver tailored observability implementations in production. In this role, you will drive a transition from reactive monitoring toward proactive, intelligence-driven observability. You will influence architectural decisions, embed observability into the software development lifecycle, guide code-level instrumentation and performance engineering practices, and ensure solutions are scalable and adaptable to diverse use cases, including event-driven and Kafka-based architectures. The role requires the ability to read, troubleshoot, and guide improvements to application code; design automation, APIs, integrations, collectors, dashboards, Kafka-aware telemetry flows, and reusable tooling; and help engineering teams adopt observability as part of day-to-day software delivery. Success in this position means improving system reliability, reducing mean-time-to-detect and resolve (MTTD/MTTR), enabling faster root-cause analysis, increasing engineering velocity, and creating a consistent, high-fidelity observability experience. You will also translate field learnings into reusable implementation patterns, code assets, automation, Kafka/event-streaming practices, developer standards, and enterprise platform capabilities that inform platform strategy and observability standards.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior