About The Position

We are looking for a Platform SRE to join Laku6, part of Carousell Group. You will take real, end-to-end ownership of running our infrastructure. You keep our systems healthy and dependable, and you are the kind of engineer the team trusts to trace an issue and bring things back when they break. You have the judgment and drive to carry a problem through to done, with the whole team behind you. This is as much a platform role as a reliability one. You will not just keep our infrastructure running, you build it: the internal tooling our engineers depend on, and the roadmap that expands what we can do. In time you help shape where the platform goes and make the whole org faster. We hire problem solvers, not tool operators. We care that you reason from the problem to the right solution, not that you are loyal to a particular stack. The tools listed below are our current context, not a checklist and not a dogma. If a problem is better solved another way, you make that call and bring us with you.

Requirements

  • A problem solver first, tool-agnostic by default. You reason from the problem to the right solution, pick the right tool for the job, and let go of one when it stops earning its place. You are not attached to a technology for its own sake, and you can justify your choices from first principles.
  • A level head under pressure. When something breaks, you can trace it and drive it to resolution, and you would rather remove the cause than escalate the symptom.
  • High agency and the appetite to grow into owning outcomes at the company level, not just your tickets.
  • 2 to 4 years in SRE, DevOps, Systems Engineering, or Software Engineering with real systems-engineering exposure.
  • Comfort writing tools and automation in Bash, Python, or Go.
  • Ability to follow end to end development process
  • Hands-on experience with containerization and orchestration (Docker, Kubernetes).
  • Experience with a cloud environment (GCP or AWS; GCP is a plus).
  • Working experience with some of our stack: databases (Postgres, MySQL), caching and queues (Redis, Kafka), CI/CD, and IaC (Terraform/Opentofu, Helm, Ansible, ArgoCD).
  • Depth in all of it is not required; the willingness and ability to pick up what you are missing is.
  • User obsession and empathy.
  • Focus on impact and results: you work on the right things and get them done.
  • Drive and resourcefulness to persevere and overcome obstacles achieving challenging goals.
  • High integrity and the ability to positively collaborate with others.

Responsibilities

  • Own the day-to-day operation of our production infrastructure.
  • When there is a bug or something goes down, you can trace it, respond, and see the problem through to resolution.
  • Open capabilities we do not have yet.
  • Turn recurring toil into shared, self-service tooling that lifts every team, and build the foundations teams rely on.
  • Grow into setting technical direction: the standards, defaults, and platform choices for how we run infrastructure as we scale.
  • Keep production reliable and observable: run our Kubernetes architecture, maintain our monitoring (VictoriaMetrics, Jaeger, Grafana), and build alerts that catch problems before users do.
  • Keep our datastores healthy (PostgreSQL, MySQL, MongoDB, Elasticsearch, Kafka, Redis, Memcache): replication, failover, and performance.
  • Manage infrastructure as code (Opentofu, Ansible, ArgoCD, Helm) and keep delivery fast and safe through CI/CD (GitHub Actions, GitLab CI, CircleCI).
  • Keep production secure and cost-efficient: secrets, firewalls, access controls, and sensible cost guardrails.
  • Bring AI into how we work, in operations and in the platform alike, where it removes toil, cuts noise, and speeds up decisions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service