Site Reliability Engineer

KlaviyoBoston, MA
Hybrid

About The Position

At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements. If you’re a close but not exact match with the description, we hope you’ll still consider applying. Design and develop systems and processes that enable highly available and scalable systems. Design, build, and deliver software to dramatically improve the availability, scalability, latency, and efficiency of services. Achieve breakthroughs in systems throughput by identifying and eliminating bottlenecks. Champion best practices by actively collaborating with other teams in a culture that values technical design review. Collaborate with other Engineers to build better software by focusing on performance, self-healing systems, configuration as code, defensive programming, and application security. Participate in periodic on-call duties with a focus on resolving issues quickly once discovered, preventing recurrences, and minimizing alert fatigue. Work closely with product-facing Engineers to ship impactful code. Perform quantitative analysis to understand and scale systems and manage the cross-functional efforts to resolve scalability issues. Produce and advocate for preventative, upstream solutions with internal stakeholders and external vendors and dependencies. Support informed, data-driven decision-making in a fast-paced environment with competing priorities. Promote Site Reliability best practices across the Engineering organization. Telecommuting permitted 2 days per week. Multiple positions. Full-time. EEO/fully supports affirmative action practices.

Requirements

  • Master’s degree in Computer Science, Computer Engineering, Software Engineering, Electrical Engineering, or a related field and 24 months of experience in an engineering occupation.
  • 24 months of experience in Designing and developing features for the Continuous Integration / Continuous Deployment (CI/CD) pipeline to enhance deployment strategies and developer productivity.
  • 24 months of experience in Python, Java, and Groovy.
  • 24 months of experience in Docker or Kubernetes.
  • 24 months of experience in EmberJS, New Relic, and SumoLogic.
  • 24 months of experience in AWS services including AWS Step Functions and AWS EC2.
  • 24 months of experience in Troubleshooting deployment failures and supporting systems during AWS outages.
  • 24 months of experience in Managing database migrations for multiple services and ensuring compliance with data security standards.
  • 24 months of experience in Maintaining CI/CD pipelines for microservices and optimizing deployment efficiency.
  • 24 months of experience in Networking concepts.

Responsibilities

  • Design and develop systems and processes that enable highly available and scalable systems.
  • Design, build, and deliver software to dramatically improve the availability, scalability, latency, and efficiency of services.
  • Achieve breakthroughs in systems throughput by identifying and eliminating bottlenecks.
  • Champion best practices by actively collaborating with other teams in a culture that values technical design review.
  • Collaborate with other Engineers to build better software by focusing on performance, self-healing systems, configuration as code, defensive programming, and application security.
  • Participate in periodic on-call duties with a focus on resolving issues quickly once discovered, preventing recurrences, and minimizing alert fatigue.
  • Work closely with product-facing Engineers to ship impactful code.
  • Perform quantitative analysis to understand and scale systems and manage the cross-functional efforts to resolve scalability issues.
  • Produce and advocate for preventative, upstream solutions with internal stakeholders and external vendors and dependencies.
  • Support informed, data-driven decision-making in a fast-paced environment with competing priorities.
  • Promote Site Reliability best practices across the Engineering organization.

Benefits

  • Annual cash bonus plan
  • Equity
  • Sign-on payments
  • Comprehensive range of health, welfare, and wellbeing benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service