This role involves owning moderately complex components within Compute/Imaging platform services, leading team-level improvements to integration frameworks and developer tooling. The engineer will perform deep debugging across Operating System, Infrastructure Orchestration, and distributed services, driving fixes that protect downstream consumers and upgrade paths. Responsibilities include analyzing usage, performance, and error budgets for specific platform surfaces, and implementing targeted resilience and capacity optimizations. The role also requires authoring and curating team-scoped documentation, samples, and adoption guidance. The engineer will design, implement, and optimize components in distributed systems with an emphasis on scalability, resiliency, and operability. This includes delivering features and load/performance tests, leveraging data plane platforms and distributed state tools for high-volume retrieval, storage, and processing, and reviewing peers’ implementations for scalability compliance. Building fault-tolerant paths (redundancy, replication, automatic failover), applying recovery-oriented principles, and implementing retries, circuit breakers, and timeouts are key. Proactive issue detection and mitigation via tests, alarms, dashboards, and telemetry, along with authoring runbooks and participating in incident response and RCAs, are also expected. The role involves implementing standard replication and synchronization, developing automation/IaC for troubleshooting and maintenance, and applying advanced security controls (encryption, access, remediation) while ensuring change, compliance, and documentation standards are met.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed