Principal Storage Architect

Graphcore•Milpitas, CA
•Hybrid

About The Position

As Principal Storage Architect, you will define the high-performance storage architecture for Graphcore's AI computing and data center infrastructures. You will lead the design of local and distributed storage tiers that provide the throughput, availability, and predictable tail latency required for large-scale training and inference. Working within Advanced Architecture, you will address storage requirements for saving model states, offloading key-value caches, large datasets, and high-speed data-loading pipelines. You will combine hands-on performance engineering with Principal-level technical leadership across hardware, software, networking, systems engineering, automation, supply chain, and external technology partners.

Requirements

  • A bachelor's or equivalent experience or master's degree in computer science, computer engineering, information technology, electrical engineering, or a related field, or equivalent practical experience.
  • Extensive hands-on experience in storage engineering and architecture for AI, high-performance computing, hyperscale, or large data center environments, including technical leadership at Principal, Staff, or Lead Architect scope.
  • Deep knowledge of NVMe, PCIe Gen5 or Gen6 architecture, NAND flash behavior, SSD form factors, endurance, firmware, and device lifecycle management.
  • Deep operational knowledge of Linux storage internals, including the block layer, I/O schedulers, direct I/O, kernel and driver behavior, and file systems such as XFS, ext4, or ZFS.
  • Hands-on experience with storage benchmarking and profiling tools such as fio, blktrace, iostat, and perf, including the ability to design representative workload models and isolate bottlenecks.
  • Programming and automation skills in Python and Bash, with experience using structured data formats and REST APIs to build monitoring, validation, or lifecycle workflows.
  • Demonstrated ability to diagnose complex hardware and software interactions, including kernel failures, PCIe errors, network issues, performance regressions, and SSD firmware defects.
  • Experience defining telemetry signals, alert thresholds, dashboards, and operational response criteria for storage health and performance at fleet scale.
  • Proven ability to lead cross-functional architecture decisions, influence senior engineers and leaders, and communicate technical risks, tradeoffs, and recommendations clearly.

Nice To Haves

  • Experience implementing, testing, or troubleshooting GPU Direct Storage or comparable direct accelerator-to-storage technologies.
  • Experience designing or operating NVMe over Fabrics deployments using RoCEv2 or TCP.
  • Experience with the Storage Performance Development Kit or other user-space storage frameworks.
  • Familiarity with high-performance parallel and distributed file systems such as Weka, VAST Data, Lustre, DAOS, DDN, or comparable platforms.
  • Knowledge of Compute Express Link and its implications for future memory and storage tiering in AI servers.

Responsibilities

  • Define the end-to-end architecture and technology roadmap for local, disaggregated, and distributed storage across Graphcore AI server and data center platforms.
  • Design storage topologies that optimize PCIe lane allocation, network-domain placement, and data paths among CPUs, AI accelerators, memory, and storage systems.
  • Lead the architecture of storage control-plane and data-plane solutions, evaluating commercial and open technologies against performance, resilience, manageability, and lifecycle requirements.
  • Profile and tune the Linux storage stack, block layer, I/O schedulers, direct I/O paths, file systems, and drivers to improve IOPS, bandwidth, and 99.99th-percentile latency.
  • Optimize storage for AI workloads, including model checkpointing, key-value cache offload, data ingestion, and direct data movement between NVMe storage and accelerator memory.
  • Set the technical direction for NVMe SSD lifecycle management, including qualification, provisioning, health monitoring, firmware rollout, failure handling, and warranty-return automation across E1.S, E3.S, U.2, and U.3 devices.
  • Define telemetry and alerting requirements for direct-attached and distributed storage, including endurance, wear, drive writes per day, thermals, capacity, performance, and latency anomalies.
  • Provide architecture requirements and technical guidance to automation teams building frameworks that characterize storage performance, reliability, and compatibility across AI platforms.
  • Lead root-cause analysis for complex storage failures and performance degradation across Linux kernels, drivers, PCIe, networks, SSD firmware, and third-party storage systems.
  • Partner with storage vendors, supply chain, networking, systems engineering, and product teams to select, integrate, and deploy storage components and platforms.
  • Evaluate emerging storage, interconnect, and memory-tiering technologies and translate relevant developments into platform requirements and future system designs.

Benefits

  • Medical, dental, and vision coverage, with options that may extend to eligible dependents.
  • Mental health, wellness, and employee assistance resources.
  • Retirement savings benefits and company contributions where applicable.
  • Paid vacation, sick time, company holidays, and parental or family leave in accordance with applicable plans and policies.
  • Life insurance and short-term or long-term disability coverage.
  • Flexible working hours and hybrid working arrangements where compatible with the role and team requirements.
  • Professional development resources, learning programs, office amenities, and team-led activities.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service