As part of the BQL (Bloomberg Query Language) Reliability Engineering team, you will build software and platform capabilities that improve the reliability, resilience, and transparency of BQL and the services it depends on. You’ll work on engineering problems at significant scale, using software and automation to make reliability a built-in property of the platform rather than a purely operational concern. You’ll be trusted to: Design, build, and maintain software and self-service platform capabilities that enable engineering teams to understand, operate, and improve the reliability of BQL at scale. Build tools and automated diagnostic capabilities that analyze telemetry and system behavior, helping engineers rapidly identify failures, regressions, and their root causes. Develop software that improves incident detection and diagnosis, reducing Mean Time to Detect (MTTD) and Mean Time to Resolve (MTTR) for high-severity incidents. Engineer observability capabilities that turn metrics, logs, traces, and other system signals into actionable insights across BQL’s distributed architecture. Partner with BQL engineering teams on system design, instrumentation, SLIs and SLOs, ensuring reliability is built into services throughout the Software Development Lifecycle. Improve platform resilience through engineering and experimentation, including load and stress testing, canary releases, controlled experiments, and failure testing. Identify recurring operational problems and eliminate them through software, automation, and improvements to platform architecture.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior