About The Position

Every time someone taps, swipes, or clicks to pay- Visa infrastructure makes it happens in milliseconds, across 200+ countries. As a Software Development Engineer on the Product Reliability Engineering (PRE) team, you won’t just watch those systems run- you’ll be one of the engineers building, automating, and evolving them. PRE is not a traditional ops team. We are a software engineering organization that treats infrastructure as code, reliability as a product, and automation as a strategic advantage. You’ll write Python, build agentic AI tools, manage data platforms, and contribute to the distributed systems that process billions of real-time transactions. From day one, you are an engineer- and from day one, your work matters. If you are endlessly curious about how large-scale systems stay resilient, obsess over elegant automation, and want to launch your career at the intersection of AI, infrastructure, and global financial technology — this role was built for you.

Requirements

  • Bachelor’s degree with 2+ years of relevant professional experience, OR an advanced degree with at least 2 years of relevant experience, OR 5+ years of relevant work experience.
  • Hands-on software engineering or automation experience using Python, Java, Go, JavaScript/TypeScript, or a comparable language.
  • Experience supporting Linux-based systems and troubleshooting distributed applications or infrastructure.
  • Hands-on experience with one or more logging or search platforms such as Splunk, ClickHouse, OpenSearch, or Elasticsearch.
  • Hands-on experience with metrics and visualization technologies such as Prometheus, Thanos, Grafana, or Bosun.
  • Experience developing backend services, APIs, command-line tools, integrations, or operational automation.
  • Experience with cloud platforms, preferably AWS or GCP, and with cloud-native architecture and services.
  • Experience with containers and orchestration technologies such as Docker, Kubernetes, or equivalent enterprise container platforms.
  • Experience with infrastructure as code, configuration management, CI/CD pipelines, Git, automated testing, and deployment tooling.
  • Understanding of telemetry pipelines, log collection and parsing, metrics collection, alerting, dashboards, data retention, access controls, and platform integrations.
  • Understanding of distributed systems, scalability, high availability, disaster recovery, performance tuning, capacity planning, and reliability engineering.
  • Experience in production incident troubleshooting, root cause analysis, problem management, and implementing preventive remediation.
  • Experience with vulnerability remediation, secure configuration, certificate management, patching, upgrades, and software lifecycle management.
  • Ability to translate user and platform requirements into maintainable engineering solutions and clear technical documentation.
  • Strong problem-solving, communication, and collaboration skills, with the ability to work effectively across globally distributed teams.

Responsibilities

  • Design, develop, and deploy end-to-end automation for deployment pipelines, infrastructure provisioning, platform operations, and release orchestration across complex production environments.
  • Write clean, production-grade Python (and Go or Bash where it counts) to eliminate toil, reduce operational risk, and improve the reliability and scalability of critical engineering workflows.
  • Design and implement reusable frameworks for release scheduling, validation, rollback, reporting, and configuration management that support the software delivery lifecycle.
  • Drive automation initiatives that improve engineering efficiency, standardization, and operational excellence across teams.
  • Design, build, operate, and continuously improve relational database platforms supporting critical payment systems and high-volume transaction processing.
  • Contribute to architecture decisions, platform enhancements, and engineering solutions that improve scalability, resiliency, and performance.
  • Lead database health and lifecycle operations including upgrades, patching, backup and recovery strategies, and platform modernization efforts.
  • Analyze and optimize database performance through index tuning, execution plan analysis, replication monitoring, and capacity management.
  • Develop automation for database operations, configuration management, and schema deployments using tools such as Ansible, Liquibase, and CI/CD pipelines.
  • Build proactive monitoring, observability, and reporting solutions that identify reliability risks before they impact production services.
  • Design and build GenAI-powered engineering solutions that automate deployment orchestration, operational workflows, release governance, and platform management.
  • Integrate LLM-driven capabilities into observability, incident response, troubleshooting, and developer productivity workflows to improve operational effectiveness.
  • Evaluate and implement emerging AI, automation, and machine learning technologies that improve reliability, efficiency, and engineering velocity.
  • Contribute to agentic automation strategies that help evolve PRE into an increasingly intelligent and autonomous engineering organization.
  • Design and build dashboards, alerts, telemetry pipelines, and health indicators using tools such as Prometheus, Grafana, Splunk, or ELK to provide visibility across globally distributed systems.
  • Analyze platform performance, reliability, utilization, and availability data to identify trends and implement long-term improvements.
  • Lead troubleshooting efforts across infrastructure, applications, databases, and platform services, performing root cause analysis and driving durable corrective actions.
  • Design and implement self-healing, automated remediation, and auto-scaling capabilities that improve system resilience and reduce operational overhead.
  • Design and implement highly available, scalable infrastructure solutions that support business-critical payment systems operating at global scale.
  • Ensure platforms and services meet security, compliance, governance, and resiliency requirements across cloud-native and hybrid environments.
  • Drive vulnerability remediation, configuration hardening, patch management, and security automation efforts to improve platform security posture.
  • Partner with engineering teams to build reliability and security practices directly into the software development lifecycle.
  • Partner with software engineers, product managers, platform teams, and global PRE peers to design, deliver, and operate reliable engineering solutions.
  • Participate in architecture reviews, design discussions, code reviews, and technical planning activities, contributing engineering expertise and best practices.
  • Create and maintain technical documentation, runbooks, operational procedures, and engineering standards that improve team effectiveness and knowledge sharing.
  • Participate in on-call rotations and incident response activities, driving operational improvements and helping teams learn from production events.
  • Take ownership of assigned initiatives from design through implementation, deployment, and operational support while continuously seeking opportunities to improve systems and processes.

Benefits

  • Medical
  • Dental
  • Vision
  • 401(k)
  • FSA/HSA
  • Life Insurance
  • Paid Time Off
  • Wellness Program
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service