Design, implement, and optimize components of highly distributed systems with a focus on scalability, reliability, availability, security, and operability. Build and test high-scale services, develop automation and Infrastructure as Code (IaC), and leverage distributed state and data plane technologies for large-scale data processing. Develop fault-tolerant systems using redundancy, replication, failover, retries, circuit breakers, and other resiliency patterns. Proactively monitor service health through testing, alarms, dashboards, and telemetry. Troubleshoot production issues, participate in incident response and root cause analysis, and maintain runbooks and operational procedures. Apply security controls and remediation practices while ensuring infrastructure, changes, and documentation meet applicable compliance and operational standards.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed