About The Position

The Head of Infrastructure will be responsible for engineering the design of high-load distributed systems, spanning from bare-metal environments to cloud-based Kubernetes. This role involves owning the platform’s reliability, performance, resilience, and security. Key duties include troubleshooting bottlenecks and production incidents across various infrastructure components such as networking, Kubernetes, messaging systems, and data infrastructure. The position requires building and executing infrastructure roadmaps that encompass migrations, high availability, monitoring, security, and CI/CD. A significant aspect of the role is driving Kubernetes adoption across all environments, including production, and building and maintaining fully operational failover and disaster recovery capabilities. The Head of Infrastructure will also be responsible for writing and reviewing Infrastructure as Code. On the leadership front, this role owns the entire infrastructure area, including priorities, capacity planning, on-call processes, and collaboration with Engineering and Product/Project Management. The individual will plan and implement infrastructure changes in close collaboration with development teams, manage infrastructure projects end-to-end (breaking down initiatives, tracking progress, removing blockers), and own technical hiring processes (candidate screening, technical interviews, hire/no-hire recommendations). Additionally, the role involves developing technical leads, empowering them to make decisions independently, and managing infrastructure costs and capacity.

Requirements

  • Strong hands-on experience with high-load systems, including messaging queues, proxies, databases, networking, Kubernetes, CI/CD, and Infrastructure as Code.
  • Ability to independently troubleshoot and resolve complex production issues.
  • Strong experience with both bare-metal and cloud infrastructure, without being tied to a single technology stack.
  • Ability to design infrastructure for distributed systems and clearly explain and justify architectural decisions.
  • Hands-on experience leading and resolving production incidents.
  • Strong systems thinking: you understand the broader consequences of technical decisions and know how to simplify complex systems.
  • 5+ years of experience as a Head or Team Lead in a product company.
  • Proven ownership of people management, prioritization, hiring, and delivery.
  • Strong hiring skills: you can identify and attract strong engineers and confidently make hire/no-hire decisions.
  • Proven ability to develop technical leads while maintaining strong technical oversight.
  • Strong ownership mindset: you focus on solving problems rather than assigning blame.
  • Ability to define the vision and direction for your area and drive it forward without waiting for tasks or instructions from senior management.

Responsibilities

  • Engineer design of high-load distributed systems, from bare-metal environments to cloud-based Kubernetes.
  • Own the platform’s reliability, performance, resilience, and security.
  • Troubleshoot bottlenecks and production incidents across networking, Kubernetes, messaging systems, and data infrastructure.
  • Build and execute infrastructure roadmaps covering migrations, high availability, monitoring, security, and CI/CD.
  • Drive Kubernetes adoption across all environments, including production.
  • Build and maintain fully operational failover and disaster recovery capabilities.
  • Write and review Infrastructure as Code.
  • Own the entire infrastructure area, including priorities, capacity planning, on-call processes, and collaboration with Engineering and Product/Project Management.
  • Plan and implement infrastructure changes in close collaboration with development teams.
  • Manage infrastructure projects end-to-end: break down initiatives, track progress, and remove blockers.
  • Own technical hiring: candidate screening, technical interviews, and final hire/no-hire recommendations.
  • Develop technical leads and empower them to make decisions independently.
  • Manage infrastructure costs and capacity.

Benefits

  • 20 vacation days and 5 family days yearly
  • Flexible start to the workday
  • Support from a professional corporate coach and psychologist
  • Regular internal and external activities, workshops, trips, and corporate events
  • Access to our internal knowledge base, meetups, and team-building activities
  • Ongoing training in new technologies and continuous professional development support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service