About The Position

Medallia is seeking a hands-on Senior Platform Software Engineer with experience operating cache, distributed query, and compute platforms at scale. As a key member of the Platform Services team, you will help ensure the availability, reliability, performance, and operational readiness of Redis, Kvrocks, Trino, and Spark. Medallia is the pioneer and market leader in Experience Management. Our award-winning SaaS platform, Medallia Experience Cloud, leads the market in the management of experiences, insights, and actions for candidates, customers, employees, patients, and residents alike. We believe that every experience is a memory that can last a lifetime. Experiences shape the way people feel about a company. And they greatly influence how likely people are to advocate, contribute, and stay. At Medallia, we are committed to creating a world where organizations are loved by their customers and their employees. We empower exceptional people to create extraordinary experiences together. Bring your whole self.

Requirements

  • 5 years of experience in software, systems, platform, DevOps, or reliability engineering.
  • 3 or more years of experience operating distributed platforms in production.
  • Hands-on operational experience managing and scaling at least two of the following platforms in production: Redis, Kvrocks, Trino, or Spark.
  • Experience leading production incident triage, root-cause analysis (RCA), and performance optimization for memory fragmentation, query execution bottlenecks, cluster failovers, and CPU/memory resource utilization.
  • Experience designing, testing and executing disaster recovery plans, automated node failovers, cluster re-sharding, and zero-downtime upgrades for high-throughput platforms
  • Demonstrated experience in Java, Go, Python, or a similar programming language.
  • Experience conducting formal architecture reviews, authoring technical design documents (e.g., RFCs/ADRs), and establishing operational runbooks across engineering teams.

Nice To Haves

  • Experience with additional platforms among Redis, Kvrocks, Trino, and Spark.
  • Experience operating distributed platforms on Kubernetes or cloud infrastructure.
  • Experience integrating Trino with object storage, Hive Metastore, or external data catalogs.
  • Experience validating topology, failover, recovery, and application-health behavior.
  • Experience with platform automation, observability, upgrades, capacity planning, and failure-mode testing.
  • Experience supporting large-scale, business-critical cache, query, or compute platforms.

Responsibilities

  • Operate and maintain Redis, Kvrocks, Trino, and Spark platforms.
  • Manage Redis Cluster and Sentinel deployments, replication, failover, persistence, upgrades, and performance.
  • Validate Kvrocks topology, replication, failover, recovery, and operational readiness.
  • Troubleshoot Trino coordinators, workers, catalogs, query execution, S3 access, and Hive Metastore integrations.
  • Operate Spark clusters and troubleshoot drivers, executors, scheduling, failover, resource usage, and application health.
  • Lead production incident triage, root-cause analysis, and performance tuning.
  • Plan and validate upgrades, capacity, failover scenarios, and recovery procedures.
  • Review application architectures and production-readiness requirements.
  • Build operational automation, monitoring, alerts, documentation, and troubleshooting playbooks.
  • Participate in a periodic on-call rotation to maintain 24/7 reliability and performance of production services.

Benefits

  • competitive health and wellness benefits, including but not limited to medical, dental, vision, 401(k), short-term and long-term disability, life and AD&D insurance, statutory leaves, paid parental leave, and paid holidays.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service