Technical Product Manager, Infrastructure

Vast.aiLos Angeles, CA
Onsite

About The Position

Vast.ai is seeking a Technical Product Manager to drive the backend our GPU cloud marketplace runs on. This is the software behind every live GPU rental (over 700k transactions a month): the daemon on every host's machines, the orchestration that powers each instance, and the infrastructure and test systems that let us ship it all at high velocity. The Scale: With over 20k GPUs, our AI cloud platform powers thousands of bleeding-edge training runs and critical production workloads for 120k+ developers all over the planet. This is a product with real scale, real data, and real users from day one. The work you ship moves tangible revenue in weeks, not quarters. The Challenge: We don't own the GPUs. Thousands of independent hosts price and operate them, and no two machines are alike. Your job is to make that heterogeneous, decentralized supply behave like the top-tier cloud our customers expect: fast, secure, and reliable.

Requirements

  • 3+ years as a backend engineer AND 3+ years in product management
  • Hands-on experience with several of: security, scalability, reliability at scale, infrastructure tooling, distributed systems, observability, abuse/fraud prevention, compliance
  • Industry experience in one or more of the following areas: software infrastructure, developer tooling (APIs, SDKs, CLIs), AI/ML, cloud computing, GPUs, or two-sided marketplaces
  • Experience at a fast-paced startup or rapid-growth team
  • AI-native. You work with AI agents every day, and you want to build the compute layer the AI era runs on.
  • Deeply technical. You came up as an engineer, and you reason about scalability, security, and efficiency as product concerns, not afterthoughts.
  • Persuasive across the stack. You can make a technical tradeoff legible to c-suite and a thankless migration compelling to the team doing it.
  • Biased toward action. You'd rather ship the fix today than present the plan next week.

Responsibilities

  • Drive the infrastructure, security, and reliability roadmap: sequencing competing priorities, turning non-functional requirements into specs engineering can build, and seeing them through until they ship.
  • Own the metrics that prove it worked: platform uptime, fleet reliability, cost-to-serve.
  • Focus on Scalability & performance: Keep the core systems fast and ahead of demand as the marketplace grows, from database performance to end-to-end latency.
  • Focus on Security, trust & compliance: Harden the platform against attacks and abuse, and build the compliance roadmap enterprise customers need.
  • Focus on Observability & infra tooling: Give every engineer a clear view of platform health, and every host a clear view of their fleet. Own the metrics, tracing, and tooling behind it.

Benefits

  • Comprehensive health, dental, vision, and life insurance
  • 401(k) with company match
  • Meaningful early-stage equity
  • Onsite meals, snacks, and close collaboration with founders/tech leaders
  • Ambitious, fast-paced startup culture where initiative is rewarded
  • Ample AI agent budget
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service