The AMHS Site Reliability Engineer (SRE) is a critical member of the software operations organization, responsible for guaranteeing the uncompromising 24x7 reliability and scalability of the software controlling Samsung's automated material transport ecosystem. In this role, you will ensure the seamless performance of the Material Control System (MCS) and OHT Control System (OCS), tackling complex server-side bottlenecks and network communications (IPC, RPC) to prevent fab operational disruptions. Your primary objective is to transform reactive incident response into proactive reliability. Beyond immediate operational support, you will spearhead the modernization of AMHS infrastructure by leveraging Python, Java, and modern SRE practices. You will design scalable data pipelines, implement comprehensive observability frameworks (Prometheus, Grafana), and manage the orchestration of services via Docker and Kubernetes. Working at the intersection of software engineering and manufacturing execution systems (MES), you will build sophisticated automation tooling to reduce manual toil and optimize critical middleware (e.g., RabbitMQ, Redis). Ultimately, your data-driven optimizations will ensure the AMHS control software scales seamlessly, directly driving the throughput and productivity of our high-stakes semiconductor fabrication site.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level