Big Data Support Engineer L1

Tookitaki Holding
Onsite

About The Position

We are seeking a Technical Support Agent who will play a critical role in maintaining high customer satisfaction by ensuring timely and effective resolution of client issues. This role is essential for Tookitaki's support services, catering to both on-premise and cloud-hosted (CaaS)clients. You will work closely with cross-functional teams to manage daily support issues, adhere to SLAs, and contribute to the continuous improvement of our customer support processes. The ideal candidate will have a strong understanding of Tookitaki's product and tech stack, the ability to triage issues effectively, excellent client management skills, and fluency in English. Knowledge of Cantonese is a plus.

Requirements

  • Apache Spark (Spark on Kubernetes preferred)
  • Kubernetes
  • Apache Kafka
  • Elasticsearch
  • AWS Services (EC2, EKS, S3, IAM, CloudWatch)
  • Docker
  • Airflow
  • Hive / Trino
  • Linux Administration
  • Prometheus
  • Grafana
  • Kibana
  • Elasticsearch Monitoring
  • CloudWatch
  • Kubernetes Logging & Monitoring
  • Log Analysis and Production Diagnostics
  • Strong SQL knowledge
  • Understanding of distributed systems architecture
  • Knowledge of networking fundamentals
  • Experience using REST APIs
  • Familiarity with Git and CI/CD concepts
  • Ability to analyze application logs and distributed system failures
  • Understanding of resource management, autoscaling, and Kubernetes scheduling
  • Excellent written and verbal English communication
  • Strong stakeholder management
  • Ability to communicate technical concepts to non-technical users
  • Experience managing customer escalations
  • Mandarin (spoken and written) is preferred
  • Cantonese is an added advantage
  • Strong analytical and troubleshooting skills
  • Ability to work under pressure during production incidents
  • Prioritize incidents based on business impact
  • Perform structured RCA and recommend preventive measures
  • Experience working with SLA-driven support environments
  • Familiarity with ITIL incident management processes
  • Strong documentation practices
  • Experience using Freshworks, Jira, ServiceNow, or similar platforms
  • 3–6 years of experience supporting large-scale production systems.
  • Experience supporting cloud-native applications running on Kubernetes.
  • Hands-on production support experience with Spark,Kubernatives, Kafka, Elasticsearch, and AWS.
  • Experience in Financial Services, FinTech, RegTech, SaaS, or Big Data platforms is highly desirable.

Nice To Haves

  • Knowledge of Cantonese is a plus.

Responsibilities

  • Handle and triage tickets related to incidents, service requests, and change requests via the Freshworks platform.
  • Provide technical support for Tookitaki's CaaS and on-premise clients post implementation, ensuring issues are resolved within SLA timelines.
  • Maintain ownership of client issues, ensuring resolutions align with SLAs and meet client expectations.
  • Collaborate with Tookitakiʼs Product Engineering and Infrastructure teams to escalate unresolved issues, secure workarounds, or deliver fixes for P1 to P4 tickets.
  • Act as a bridge between Services (onboarding team) and Support, ensuring a seamless transition when clients go live.
  • Triage technical issues effectively by diagnosing the problem, identifying the root cause, and determining the appropriate resolution path.
  • Develop a deep understanding of Tookitaki's product architecture and tech stack (AWS, Big Data technologies like Hive, ElasticSearch, Kubernetes, etc).
  • Build and maintain strong relationships with clients, demonstrating excellent communication skills and a customer-first approach.
  • Clearly explain technical resolutions to non-technical stakeholders, ensuring transparency and trust.
  • Maintain thorough documentation of all support tickets, including actions taken and lessons learned, in the Freshworks platform.
  • Proactively suggest process improvements to enhance support efficiency and client satisfaction.
  • Participate in rotational shifts and ensure availability during defined upgrade windows (e.g., second and fourth Saturdays) to support both infra-wide updates and tenant-specific changes.
  • Ensure 24/7 availability as part of the teamʼs support structure for critical escalations.
  • Own customer incidents, service requests, and change requests through Freshworks or equivalent ticketing platforms.
  • Troubleshoot production issues affecting Spark jobs, Kubernetes workloads, Kafka pipelines, Elasticsearch clusters, APIs, and cloud infrastructure.
  • Ensure all customer issues are resolved within SLA while maintaining high customer satisfaction.
  • Perform root cause analysis (RCA) and document preventive actions.
  • Monitor health and performance of Spark applications running on Kubernetes.
  • Investigate failed Spark jobs, executor failures, pod crashes, resource contention, and scheduling issues.
  • Monitor Kafka topics, brokers, consumer groups, lag, and message delivery.
  • Support Elasticsearch cluster health, indexing pipelines, shard allocation, and search performance.
  • Perform production validations after deployments and infrastructure upgrades.
  • Participate in planned maintenance activities, upgrades, and release support.
  • Diagnose issues across the complete cloud-native stack including Apache Spark on Kubernetes, Apache Kafka, Kubernetes (Pods, Deployments, Services, ConfigMaps, Secrets), Docker Containers, AWS Infrastructure, Elasticsearch, Airflow Workflow Orchestration, Hive / Trino / SQL-based Data Processing, REST APIs and Microservices, Linux-based Production Systems.
  • Work closely with Product Engineering to identify software defects and Infrastructure teams to resolve platform-related issues.
  • Coordinate with Product Engineering for bug fixes and product improvements.
  • Work closely with Infrastructure teams during production incidents.
  • Support onboarding teams during production go-live and customer transition.
  • Participate in Major Incident Management (P1/P2).
  • Provide timely updates to customers during incidents.
  • Communicate technical issues in a clear and business-friendly manner.
  • Maintain ownership until issue closure.
  • Prepare incident summaries, RCA documents, and customer communications.
  • Create and maintain SOPs, troubleshooting guides, and knowledge base articles.
  • Identify recurring issues and recommend automation opportunities.
  • Improve monitoring, alerting, and operational processes.
  • Contribute to platform reliability and operational excellence initiatives.
  • Participate in 24x7 production support rotation.
  • Provide support during scheduled maintenance windows, infrastructure upgrades, and customer go-lives.
  • Support weekend deployment activities when required.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service