Reliability Engineer Responsible for the availability, performance, monitoring, and incident response, among other things, of the cloud platforms and services. Ensure that everything that goes to production complies with a set of general requirements like diagrams, dependencies of other services, monitoring and logging plans, backups and possible high availability setups. Manages uncaught exceptions, hardware degradation, networking problems, high usage of resources, or slow responses that could happen at any time. Uses metrics such as mean time to recover (MTTR) and mean time to failure (MTTF). Considered an emerging authority, who applies extensive technical expertise. Develops technical solutions to complex problems. Exercises considerable latitude in determining objectives and approaches to Assignment.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior