The Site Reliability Engineer II optimizes service performance, actively participates in reliability improvements, and conducts in-depth SLO and capacity analysis. This position exists to enhance system reliability and scalability while contributing to automation and self-service tool development. The role involves monitoring service performance, troubleshooting production issues, understanding system architecture, monitoring service reliability, participating in resolving basic issues, learning disaster recovery testing procedures, understanding SLO concepts, monitoring and analyzing SLO patterns, assisting in implementing SLO visualization and alerting, performing basic capacity analysis, identifying trends in system capacity, participating in capacity planning, deploying and maintaining existing automation tools, creating simple scripts, and troubleshooting automation scripts.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
Associate degree