Systems Resilience Engineer - Hybrid Cloud Storage

QumuloSeattle, WA
$140,000 - $210,000Hybrid

About The Position

Qumulo's cloud data platform manages exabytes of the world's most demanding data, unifying files, objects, and every workload across edge, core, and cloud. This role is for an engineer who thinks like a breaker, finding out how Qumulo breaks before customers do. The platform manages exabytes of data for over 1,100 customers across on-prem and every major cloud, handling mission-critical workloads where a missed edge case can lead to significant customer impact. You will put on the customer's hat, understand how a feature will be used in real-world scenarios, and design tests to push it beyond its limits across both hardware and cloud environments. Your responsibilities will include automating current manual testing processes, determining testing strategies (what, how often, why), and contributing to setting the quality bar for releases.

Requirements

  • 3+ years building and operating automated testing, validation, and/or certification for complex software systems.
  • Strong programming ability in C.
  • A 'breaker's instinct' – proactively seeking edge cases and testing potential failure scenarios.
  • Proven track record of building tests independently.
  • Hands-on experience with both on-premises infrastructure and cloud environments (AWS, GCP, or Azure).
  • Working fluency in Linux (Ubuntu) and Python.
  • A data-driven approach to test strategy.
  • Solid understanding of networks (routing, firewalls, security inspection devices, switch configuration) is a plus.

Nice To Haves

  • Experience with Qumulo’s distributed file system, or parallel file systems.
  • Experience with orchestration tools (Ansible, Terraform), containers, and Kubernetes.
  • Storage (IOPS, Latency, read/write patterns) or protocol experience (NFS, SMB, S3).

Responsibilities

  • Design and operationalize testing for new features, including understanding customer usage, scale-testing, and identifying failure points.
  • Automate manual, repetitive testing using Python and in-house frameworks on Jenkins and Argo.
  • Develop a data-driven plan for test execution frequency and rationale, including a framework for scheduling and rerunning tests.
  • Troubleshoot build and test failures across VM instances and Qumulo-qualified hardware, distinguishing between compile-time errors, integration failures, infrastructure issues, and actual bugs.
  • Implement monitoring and alerting using tools like OpenMetrics, Grafana, InfluxDB, and Prometheus.
  • Contribute to setting the quality bar for releases and have a say in what ships.
  • Participate in an on-call rotation for systems owned by the team.

Benefits

  • Pre-IPO stock options
  • Flexible time-off policy
  • HSA and PPO health insurance options
  • Dental and Vision insurance
  • 401(k) plan
  • Choice of an ORCA card or parking subsidy
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service