About The Position

We exist to wow our customers. We know we’re doing the right thing when we hear our customers say, “How did I ever live without Coupang?” Born out of an obsession to make shopping, eating, and living easier than ever, we’re collectively disrupting the multi-billion-dollar e-commerce industry from the ground up. We are one of the fastest-growing e-commerce companies that established an unparalleled reputation for being a dominant and reliable force in South Korean commerce. We are proud to have the best of both worlds — a startup culture with the resources of a large global public company. This fuels us to continue our growth and launch new services at the speed we have been since our inception. We are all entrepreneurs surrounded by opportunities to drive new initiatives and innovations. At our core, we are bold and ambitious people that like to get our hands dirty and make a hands-on impact. At Coupang, you will see yourself, your colleagues, your team, and the company grow every day. Our mission to build the future of commerce is real. We push the boundaries of what’s possible to solve problems and break traditional tradeoffs. Join Coupang now to create an epic experience in this always-on, high-tech, and hyper-connected world. Role Overview: Implement SRE best practices to improve the reliability, scalability, and performance of datacenter services. Develop and maintain automation scripts for infrastructure provisioning, monitoring, and management. Conduct root cause analysis and post-mortem reviews to prevent recurrence of incidents.

Requirements

  • Bachelor's degree in Computer Science, Electrical Engineering, Math, or a closely related field
  • 8 years of experience in backend software development
  • Experience working in cloud environments, particularly AWS
  • Demonstrated experience in building and maintaining highly available, distributed systems

Nice To Haves

  • Proficiency in observability tools and technologies (e.g., Prometheus, Grafana, ELK Stack).
  • Experience with SRE practices and tools (e.g., Kubernetes, Docker, Terraform).
  • Strong programming and scripting skills (e.g., Python, Go, Bash).
  • Familiarity with cloud platforms (AWS, Azure, GCP) and their observability and reliability services.
  • Strong problem-solving skills and attention to detail.
  • Excellent communication and collaboration skills.
  • Ability to work in a fast-paced, dynamic environment.

Responsibilities

  • Design, implement, and maintain observability solutions for datacenter infrastructure.
  • Develop, deploy, and maintain the operational and reliability components of a large-scale Observability and Telemetry collection platform, emphasizing performance at scale, real-time monitoring, logging, and alerting.
  • Participate in and enhance the entire lifecycle of services, from inception and design to deployment, operation, and refinement.
  • Develop and optimize monitoring systems to ensure high availability and performance.
  • Create and manage dashboards, alerts, and reports to provide visibility into system health and performance.
  • Analyze and optimize the performance of datacenter systems and applications.
  • Implement best practices for resource utilization and efficiency.
  • Work closely with other engineering teams to understand and meet their observability and reliability requirements.
  • Collaborate with hardware and software vendors to evaluate and integrate new technologies.
  • Ensure that observability and reliability solutions comply with security policies and industry standards.
  • Implement and maintain security measures to protect data and infrastructure.
  • Provide support for observability and reliability-related issues, including debugging and resolving hardware and software problems.
  • Develop and maintain documentation for troubleshooting procedures and best practices.
  • Stay updated with the latest advancements in observability and SRE technologies and integrate them into the infrastructure.
  • Continuously improve the reliability, scalability, and performance of datacenter services.

Benefits

  • Medical/Dental/Vision/Life, AD&D insurance
  • Flexible Spending Accounts (FSA) & Health Savings Account (HSA)
  • Long-term/Short-term Disability
  • Employee Assistance Program (EAP) program
  • 401K Plan with Company Match
  • 18-21 days of the Paid Time Off (PTO) a year based on the tenure
  • 12 Public Holidays
  • Paid Parental leave
  • Pre-tax commuter benefits
  • MTV - [Free] Electric Car Charging Station
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service