About The Position

We're looking for a Senior Site Reliability Engineer, Platform Infrastructure to take hands-on technical ownership of the architecture, reliability, and scalability of our entire AWS infrastructure. Reporting to the Engineering Manager, Platform Infrastructure & SRE, you'll set technical direction, review designs, and raise the bar for reliability engineering across a growing and globally distributed engineering organization. This is a senior individual contributor role. You'll work side by side with our onsite SRE team, Software Engineers, and other Lead Engineers to support seamless 24/7 reliability. It's ideal for an AI-forward engineer with a strong software engineering background who uses AI-assisted development tools to move faster, has a passion for infrastructure-as-code, and a proven track record of mentoring engineers to build highly reliable, scalable, and performant systems.

Requirements

  • 4+ years of experience in a software engineering or site reliability engineering role.
  • An AI-forward mindset: fluency with AI-assisted development tools (e.g., Claude Code, GitHub Copilot, Cursor) to accelerate delivery, practical prompt engineering skills, and experience managing context (e.g., structuring prompts, memory, and retrieved data) to keep LLM-based workflows accurate and reliable, plus hands-on exposure to AI/ML infrastructure (e.g., model serving, vector databases, LLM operations).
  • A strong background in software engineering, ideally with experience in backend microservices (.NET is a strong plus).
  • Deep, hands-on expertise with AWS and its core services (e.g., EC2, S3, RDS, Kinesis, VPC, IAM).
  • Proven experience building and managing infrastructure with Infrastructure-as-Code (IaC) tools like Terraform or CloudFormation.
  • A solid understanding of SRE principles and a proven track record of improving site reliability.
  • Experience with modern observability, monitoring, and logging platforms, ideally Datadog and OpenTelemetry.
  • A Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent industry experience.

Nice To Haves

  • Communicates complex technical concepts with clarity—written and verbal—to diverse audiences across engineering, product, and leadership.
  • Mentors with intent, with a demonstrated ability to develop engineers and help them grow in their careers.
  • Influences without authority, building genuine alignment on reliability and infrastructure standards across teams that don't report to you.
  • Stays calm and decisive under production pressure, leading incident response and blameless postmortems that turn outages into durable fixes.
  • Stays genuinely curious, tracking advances in cloud infrastructure, reliability engineering, and AI tooling, and pulls the best of what's new into the team's everyday practice.

Responsibilities

  • Own the architecture, reliability, and scalability of critical AWS infrastructure, working hands-on across the full stack.
  • Partner with Software Engineers and other Lead Engineers to shape the roadmap and technical strategy for Cricut's platform infrastructure.
  • Take ownership of our AWS environment, driving best practices in security, cost management, and scalability.
  • Champion and expand our "infrastructure-as-code" philosophy across the organization.
  • Use AI-assisted development tools (e.g., Claude Code, GitHub Copilot) to accelerate delivery, applying prompt engineering and context management practices to get reliable results, and evaluate AI/ML infrastructure (e.g., model serving, vector databases, LLM tooling) as it becomes part of the platform.
  • Oversee production monitoring, incident response, and blameless post-mortem processes to continuously improve system reliability.
  • Develop and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical production systems.
  • Act as a key consultant for our feature-focused pillar and pod teams, ensuring they have the infrastructure resources and support required to deliver their projects successfully.
  • Mentor software engineers who have an affinity for infrastructure, helping them grow their skills in reliability engineering.
  • Collaborate closely with the onsite SRE team, sharing the on-call rotation to ensure seamless 24/7 reliability coverage.

Benefits

  • Competitive Medical, Dental, and Vision coverage
  • 401(k) match
  • Generous PTO
  • Tuition reimbursement
  • Yearly lifestyle stipend
  • Exclusive employee discounts
  • Relocation assistance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service