Senior Software Engineer (Cloud Infrastructure)

YipitData
$168,000 - $193,725Remote

About The Position

We are hiring a Senior Software Engineer to help modernize, scale, and operate the shared cloud infrastructure that powers YipitData’s data-intensive products, agentic customer experiences, and engineering organization. Reporting to the Head of Cloud Infrastructure, you will drive forward Site Reliability Engineering and DevOps initiatives on mission-critical systems that our customers rely on 24/7. You will work on a modern tech stack built with AWS Cloudformation, Databricks, Kubernetes, and Terraform. This is a hands-on engineering role with meaningful ownership. You will help build reliable, reusable infrastructure capabilities; improve how engineers develop, deploy, observe, and operate cloud technologies; and strengthen our approach to reliability, incident management, disaster recovery, and technology efficiency. You will collaborate closely with platform, application, and data engineers while contributing to a small, high-leverage infrastructure team. The infrastructure team owns shared platforms and capabilities, while adjacent Data Platform, Data Engineering, and application teams own their applications and data products. You will partner with those teams to provide smooth integrations, operational guidance, and dependable infrastructure that helps them move quickly and safely. Please note we expect the Senior Software Engineering to work East Coast hours.

Requirements

  • 4+ years of relevant software, infrastructure, SRE, platform engineering, or DevOps experience, with hands-on experience using AWS, Kubernetes, Datadog, Databricks, and Terraform in production environments.
  • A strong interest in building AI-native and data intensive applications and systems, working closely with Databricks, OpenAI, Anthropic, and open source AI tools to create reliable and efficient AI solutions.
  • Write reliable, maintainable software and automation and are comfortable working across infrastructure, systems, and application boundaries.
  • Operated business-critical production systems and understand observability, incident response, operational readiness, and root-cause analysis.
  • Experience improving CI/CD systems, infrastructure as code, deployment workflows, or internal developer platforms.
  • Understand reliability concepts such as SLIs, SLOs, error budgets, capacity planning, and disaster-recovery objectives. Familiarity with observability solutions like Datadog is a plus.
  • Can independently own projects, make sound technical tradeoffs, communicate clearly, and collaborate effectively with engineers across multiple teams.
  • Motivated by reducing engineering friction and creating reusable solutions rather than repeatedly solving the same problem manually.
  • Comfortable participating in an on-call or incident-response rotation and helping the team design a sustainable operational model.
  • Hold a Bachelor’s degree or have equivalent practical experience.

Nice To Haves

  • Familiarity with observability solutions like Datadog is a plus.
  • Experience supporting data-intensive SaaS, analytics, familiarity with big data analytics, or rapidly scaling technology environments is a plus.

Responsibilities

  • Design, build, and operate shared cloud infrastructure using AWS, Kubernetes, Terraform, Databricks, Cloudflare, and related cloud-native technologies.
  • Deliver SRE and DevOps initiatives that improve reliability, scalability, observability, deployment safety, and operational readiness.
  • Build reusable infrastructure modules, automation, and self-service workflows that reduce manual work and improve the developer experience.
  • Help define and implement service-level indicators, service-level objectives, monitoring, alerting, and error-budget practices for critical systems.
  • Participate in incident response and improve operational outcomes through clear runbooks, effective post-incident reviews, and durable corrective actions.
  • Strengthen disaster-recovery readiness through recovery planning, automation, testing, and remediation of identified gaps.
  • Improve CI/CD workflows and infrastructure delivery so engineering teams receive faster feedback and can deploy confidently.
  • Partner with application, Data Platform, and Data Engineering teams to understand infrastructure needs and help teams operate their workloads effectively.
  • Improve cloud efficiency through thoughtful architecture, capacity planning, Kubernetes resource optimization, cost visibility, and automation.
  • Contribute to technical standards, architecture decisions, documentation, and the evolution of the team’s sustainable 24/7 operating model.

Benefits

  • flexible work hours
  • flexible vacation
  • a generous 401K match
  • parental leave
  • team events
  • wellness budget
  • learning reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service