Sr. Software Engineer – Cloud Infrastructure and Devops

PayPalSan Jose, CA
$123,500 - $212,850Hybrid

About The Position

This job delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC). It involves advising management on project-level issues, guiding junior engineers, operating with little supervision, and applying knowledge of technical best practices. In this role, you will be a key contributor to Venmo’s chaos engineering and business continuity efforts, building and operating the systems that ensure our infrastructure can withstand and recover from any failure scenario. Our Business Continuity team is responsible for the resilience and recoverability of Venmo’s critical systems. We design, build, and operate the infrastructure, tooling, and processes that protect the company from disruption; spanning disaster recovery automation, chaos engineering, incident game days, and backup and restore systems. Our work ensures that Venmo can withstand and rapidly recover from any failure scenario, from infrastructure outages to region-wide incidents. We operate at the intersection of cloud engineering, SRE, and reliability, and our work is foundational to the availability of every product and service at Venmo. As a Senior Engineer on the Business Continuity team, you will deliver complete solutions spanning all phases of the SDLC. You are a hands-on contributor that teaches by example and mentors other engineers and product owners on how to improve their development, using our systems, by driving best practices. We have a lot of independence in our decisions, but we are kept accountable for the results. Our technical stack: AWS, EKS, Docker, GitHub Enterprise, Terraform, GitHub Actions, DataDog, Bash, Python, Go.

Requirements

  • 3+ years relevant experience and a Bachelor’s degree OR Any equivalent combination of education and experience.
  • Bachelor’s in computer science or related field of study
  • 5+ years’ experience in software development or a related field
  • 3+ years’ experience operating distributed applications 24x7x365, as part of a Cloud Engineering, DevOps, and/or SRE team
  • Extensive hands-on experience with designing, implementing, and supporting infrastructure (AWS experience preferred) to support global-scale services
  • Deep hands-on experience with IaaS and PaaS solutions from AWS (or similar cloud provider)
  • Hands-on programming and scripting (Python, Java, Bash, Go)
  • Hands-on experience with containers and container orchestration: Docker, Kubernetes
  • Strong communication skills with the ability to understand and explain technical issues to a non-technical audience
  • Experience with chaos engineering tools and frameworks (e.g., AWS Fault Injection Simulator, Gremlin, Chaos Monkey, Litmus)
  • Hands-on experience designing and implementing disaster recovery solutions for distributed systems
  • Experience developing and maintaining backup and restore tooling and strategies
  • Experience planning and executing game day exercises or incident simulations
  • Understanding of RTO/RPO requirements and how to design systems to meet them

Nice To Haves

  • GitHub Enterprise
  • Terraform
  • GitHub Actions
  • DataDog
  • Bash
  • Python
  • Go

Responsibilities

  • Delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC) (design, implementation, testing, delivery and operations), based on definitions from more senior roles.
  • Advises immediate management on project-level issues
  • Guides junior engineers
  • Operates with little day-to-day supervision, making technical decisions based on knowledge of internal conventions and industry best practices
  • Applies knowledge of technical best practices in making decisions
  • Act as a hands-on contributor, while leading by example
  • Mentor junior engineers
  • Enjoy a high degree of independence in decision-making while being accountable for results of yourself and your team
  • Be part of a team dedicated to driving the scalability and reliability of Venmo’s AWS cloud infrastructure
  • Contribute to initiatives: internal teams rarely have dedicated project managers. You will define the design, identify stakeholders, navigate risks and changes, coordinate colleagues’ work, implement solutions, and be accountable for timely and quality project delivery
  • Troubleshoot incidents, identify root causes, fix and document problems, and implement preventive measures
  • Lead by example, making meaningful contributions to the improvement of engineering teams’ production operations
  • Develop and improve tools and automation to manage infrastructure and application configuration as code
  • Enhance the quality, reliability, and stability of our infrastructure and operations
  • Design, implement, and operate chaos engineering experiments to proactively identify and remediate system weaknesses
  • Build and maintain disaster recovery automation, runbooks, and infrastructure across Venmo’s cloud environment
  • Plan and execute chaos and incident game day exercises to validate system resilience and team preparedness
  • Develop and maintain backup and restore tooling and processes for critical systems and data
  • Define and track resilience metrics, SLOs, and recovery time objectives for owned systems

Benefits

  • generous paid time off
  • healthcare coverage for you and your family
  • resources to create financial security
  • support your mental health
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service