Staff Infrastructure Engineer

Tabs•New York, NY
•$220,000 - $270,000•Onsite

About The Position

Tabs is seeking a Staff Infrastructure & Reliability Engineer to manage its AWS environment, software deployment processes, and monitoring systems. This role involves setting the company's infrastructure direction as a hands-on individual contributor, collaborating with platform and product engineering teams. The position is crucial for a growing company, focusing on building a robust, maintainable, and scalable infrastructure that meets compliance requirements for payments, billing, and revenue. The engineer will influence engineering and product decisions, working closely with supportive leadership and high-caliber engineers. The role requires owning outcomes and contributing to a culture of reliability and continuous improvement.

Requirements

  • Software engineer first, with infrastructure expertise built on that foundation.
  • Experience running production systems on AWS and leading platform-level change.
  • Systems thinking: understanding of risk, rollback strategy, blast radius, and feedback loops.
  • Treating CI/CD and environments as products that should be fast, reliable, and self-serve.
  • Ability to dig into logs and data independently when issues arise, especially under pressure.
  • Influencing through trust and clarity rather than control.
  • Balancing pragmatism with long-term system health.
  • Valuing learning from failure and improving processes over assigning blame.
  • Clear communication and ability to work well across teams.
  • 8+ years in software engineering, infrastructure, or SRE roles.
  • Experience in one or more modern languages such as TypeScript with a track record of writing production-quality scripts, tools, and services, and still hands-on today.
  • Deep hands-on experience running production workloads on AWS, including container platforms such as ECS/Fargate.
  • Expertise with infrastructure as code using Terraform, and ownership of Docker and Git workflows in production.
  • Solid working knowledge of Kubernetes and Helm.
  • Experience designing and running CI/CD systems such as GitHub Actions, including build parallelization and developer experience improvements.
  • Deep experience with observability tooling across metrics, logs, tracing, and alerting, including defining SLIs, SLOs, and error budgets.
  • Expertise operating distributed systems in production at scale, with an implementation-level understanding of messaging systems, partitioning, deploy strategies, and failure modes.
  • A track record of leading high-severity incidents, debugging live production issues from logs and data, and running blameless postmortems.
  • Experience proposing and evaluating multiple architectures, making trade-offs that fit the company's stage, and driving infrastructure decisions across teams.
  • Experience across more than one architecture or company environment, ideally including both larger companies and high-growth startups.
  • Comfortable navigating ambiguity and setting direction in a fast-moving environment.
  • Experience mentoring engineers and shaping infrastructure practices across an engineering org.

Nice To Haves

  • Experience owning broad infrastructure surface area at a high-growth startup, including as the primary infrastructure or SRE owner.
  • Experience operating Kafka or a similar distributed messaging system at scale.
  • Experience building developer tooling that engineers adopt and rely on.
  • Prisma expertise.

Responsibilities

  • Define and evolve reliability standards, SLIs, SLOs, and error budgets.
  • Improve observability, alerting, and incident processes across services.
  • Lead high-severity incidents hands-on and drive clear, actionable follow-ups.
  • Partner with engineering teams to design resilient, scalable systems.
  • Write production-quality code and automation to reduce toil and lower operational risk.
  • Mentor engineers and influence best practices across teams.

Benefits

  • Competitive compensation and equity
  • Unlimited PTO
  • Up to 100% employer covered monthly healthcare premium (medical, dental, vision)
  • Lunch provided via Sharebite, plus dinner for any later in office days.
  • Parental leave up to 12 weeks
  • Tax free commuter and parking benefits
  • Voluntary insurances (Life, Hospital, Critical Illness, Accident)
  • Employee Assistance Program (Rightway)
  • Free One Medical Membership
  • 401k
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service