Principal Field Engineer (Hybrid)

DigiStor•Austin, TX
•Hybrid

About The Position

We are seeking an experienced enterprise SSD engineer to take technical ownership of complex drive and firmware issues. You will investigate failures, turn customer symptoms into repeatable tests and establish the evidence needed to verify a resolution. You will build a detailed understanding of DigiStor firmware behavior and work closely with controller and firmware partners to assess proposed changes and identify their impact. This is a senior individual contributor role with influence across engineering, product management and customer technical teams. It suits an engineer who enjoys working directly with logs, code and hardware, and can explain difficult findings clearly. This role will be remote, but you will need to spend at least 30% of the time onsite at client.

Requirements

  • Typically 10–15 years of engineering experience, with substantial focus on enterprise SSD technology, firmware diagnostics, drive validation or storage failure analysis.
  • Strong understanding of SSD architecture, including NAND, controllers, the flash translation layer, wear leveling, endurance, error correction and power loss protection.
  • Practical expertise in NVMe and PCIe, with the ability to investigate interactions between drives, server platforms, BIOS, drivers and operating systems.
  • Advanced log analysis and troubleshooting skills, with a record of converting intermittent failures into reproducible cases and defensible root cause conclusions.
  • Strong scripting skills in Python and Bash or equivalent languages. Experience building automated diagnostics, workload generators and reproduction tools.
  • Ability to interpret C or C++ firmware code, debug embedded behavior and assess how a change could affect reliability, performance or compatibility.
  • Sound technical judgment, clear written communication and the confidence to question a proposed fix while remaining constructive with customers and engineering partners.
  • A degree in electrical engineering, computer engineering, computer science or a related discipline, or equivalent practical experience.

Nice To Haves

  • Enterprise SSD OEM or controller supplier experience, including customer escalations, firmware qualification and collaboration with server manufacturers.
  • Practical use of Linux, nvme-cli, fio, firmware debug utilities, Git and PCIe protocol analyzers; Windows and PowerShell experience is also useful.
  • Familiarity with NVMe management interfaces, secure firmware updates, self-encrypting drives and TCG Opal, or storage certification requirements.

Responsibilities

  • Lead root cause investigations into enterprise SSD failures, timeouts, resets, data integrity concerns and performance regressions. Isolate the contribution of drive firmware, host software, PCIe connectivity and system configuration.
  • Analyze NVMe health and error logs, telemetry, firmware traces and crash dumps. Correlate events across the drive, operating system and server to establish the sequence leading to a failure.
  • Develop scripts to automate log collection, parsing, event correlation and diagnostic reporting, with clear version tracking and repeatable execution.
  • Create reliable reproduction scripts and targeted workloads. Document the hardware, firmware, configuration, command sequence and failure trigger so partners can reproduce the same behavior.
  • Build a detailed understanding of firmware architecture and operation, including command handling, the flash translation layer, garbage collection, power management and error recovery.
  • Review firmware changes, release notes and source code where available. Assess the proposed fix, design regression coverage and explain the risks before recommending adoption.
  • Work directly with controller and firmware suppliers to resolve defects. Provide complete evidence, challenge unsupported conclusions and track issues through to verified closure.
  • Evaluate performance, latency consistency, stability and data integrity under representative enterprise workloads. Use controlled comparisons to distinguish drive behavior from host or test limitations.
  • Maintain reusable diagnostic tools and technical records, mentor engineers and communicate investigation status and release recommendations to customers and internal stakeholders.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service