Neo Cloud - Principal AI Cloud Storage Engineer

Blaze Talent•Seattle, WA
•Remote

About The Position

Neo Cloud is seeking a Principal Software Engineer to define and build its AI cloud storage platform. This is a senior individual-contributor role for an experienced system architect who can operate at the intersection of customer needs, system architecture, and production operations. The role involves translating future AI/ML customer storage needs into a performant, resilient, and economical storage system at massive scale.

Requirements

  • 10+ years of professional software engineering experience building and operating cloud storage systems in production.
  • Direct experience with AI-focused storage platforms such as DDN, Weka, or VAST, including their architectural approaches to throughput, caching, and GPU-cluster integration.
  • Deep understanding of distributed systems fundamentals: consistency models, replication, consensus, partitioning, failure detection, and recovery.
  • Proven experience operating high-scale distributed systems in production, including on-call ownership, incident response, and driving systemic reliability improvements.
  • Strong systems programming skills (e.g., Go, C++, Rust, or Java) and comfort working across the stack from low-level I/O and networking to distributed control planes.
  • Excellent written and verbal communication skills.

Nice To Haves

  • Experience designing or tuning local node caching layers to accelerate AI training and inference data access.
  • Experience with high-performance networking.
  • Contributions to open-source storage projects, relevant patents, or published technical papers/talks.

Responsibilities

  • Engage directly with customers, solutions architects, and product teams to understand future storage requirements for AI/ML workloads.
  • Extrapolate from current usage patterns and industry trends to anticipate future requirements.
  • Partner with product management to prioritize platform investments based on near-term customer needs and longer-term strategic bets.
  • Design, implement, and operate AI cloud object and file storage systems, including the data path, metadata path, and control plane.
  • Take a system-level approach that accounts for the full characteristics of AI workloads, building end-to-end solutions — including local node caching strategies, the network hardware and protocols that move data between storage and compute (e.g., RDMA, high-throughput NICs, congestion control), and the underlying storage hardware and software (media, erasure coding, metadata services) — and understanding how decisions in one layer constrain or unlock the others.
  • Architect for the specific demands of AI workloads: very high aggregate throughput to keep GPU/accelerator clusters fed, support for massive numbers of small and large objects and files, efficient checkpointing at scale, and predictable tail latency under heavy concurrent load.
  • Drive core storage system design decisions, including durability and consistency models, erasure coding and replication strategies, metadata scalability, multi-tenancy and isolation, and S3-compatible and POSIX/file-protocol API design.
  • Take end-to-end ownership of services in production: build for observability and operability from day one, participate in on-call, lead incident response and root-cause analysis for critical issues, and drive long-term reliability and performance improvements.
  • Identify and eliminate performance bottlenecks and scalability limits before they become customer-facing problems; lead capacity planning for rapid growth.
  • Ensure deep understanding of data privacy and security and its implication on performance.
  • Partner with network engineers to deliver complete AI storage system.
  • Set technical direction and best practices for the storage organization; author and review design documents for significant architectural changes.
  • Provide deep technical mentorship to senior and staff engineers; raise the engineering bar across the team through code review, design review, and hands-on collaboration.
  • Influence technical strategy across adjacent teams.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service