About The Position

Netspend Corporation is a global, vertically-integrated financial services and technology company dedicated to the delivery of innovative financial empowerment solutions to consumers worldwide. Netspend's financial products and services span prepaid, debit, cross-border payments, and loyalty solutions for consumers and enterprise partners. Netspend provides prepaid and debit account solutions that connect customers with secure, convenient access to global payment networks so they can manage their money and make everyday purchases. With a nationwide U.S. retail network, customers can purchase and reload Netspend products at 130,000 reload points and over 100,000 distributing locations. Since our founding in 1999 by industry pioneers, Netspend products have processed billions of dollars in transaction volume and served millions of customers worldwide. The company is headquartered in Austin, Texas with employees worldwide. We are seeking a visionary Director – Observability, Response & Reliability (ORR) with 15+ years of overall technical experience, including 5+ years in engineering leadership, to serve as our primary authority on system resilience, full-stack observability, and enterprise incident management across Netspend's global financial ecosystem. In this high-visibility role, you will lead our India-based ORR organization, driving the strategy that transforms how Netspend monitors, predicts, responds to, and resolves critical operational events. You will establish a world-class Observability and AIOps practice, institutionalize Site Reliability Engineering (SRE) principles across all product teams, and safeguard systems handling ACH processing, core payment rails, millions of daily card transactions, and bank-sensitive data.

Requirements

  • Bachelor’s or Master’s Degree in Computer Science, Software Engineering, Information Technology, or a related quantitative field.
  • 15+ years of total experience in SRE, Systems/Platform Engineering, Infrastructure Architecture, and Enterprise Observability.
  • Deep expertise architecting enterprise telemetry solutions using tools such as Splunk, Dynatrace, Datadog, Prometheus, Grafana, and OpenTelemetry.
  • Proven track record implementing AIOps tools, automated incident remediation, and AI/ML-based anomaly detection engines.
  • Expert knowledge of SRE best practices, error budget policy enforcement, SLO modeling, and incident response frameworks.
  • Strong hands-on architectural understanding of AWS services (EC2, ALB/NLB, Direct Connect, IAM, KMS, Transfer Family).
  • Practical experience executing "strangler fig" migration strategies off legacy tooling (Puppet, SVN, Xymon) to modern GitOps/IaC (GitLab CI/CD, Terraform, Ansible).
  • Experience monitoring and troubleshooting distributed middleware and database technologies (Apache Kafka, Cassandra) under heavy throughput.
  • Understanding of secure file transfer (SFTP, AS2), mTLS, cryptographic key management (HSMs, Virtucrypt), and high-availability payment processing environments.

Nice To Haves

  • Splunk Certified Architect, Dynatrace Master / Professional, or equivalent.
  • AWS Certified Solutions Architect – Professional
  • Certified Site Reliability Engineer (SRE), ITIL v4 (Incident/Problem Management focus).

Responsibilities

  • Design and execute the enterprise Observability roadmap, establishing full-stack, end-to-end visibility across microservices, legacy monoliths, cloud infrastructure, and data pipelines (Kafka, Cassandra).
  • Define unified logging, metrics, traces, and synthetics standards (OpenTelemetry, Prometheus, Splunk, Dynatrace) across all engineering groups to enable granular transaction-level tracing for payment flows.
  • Pioneer the adoption of AIOps, Machine Learning, and anomaly detection to shift the engineering culture from reactive alert firefighting to proactive noise reduction, predictive fault detection, and automated self-healing workflows.
  • Build custom monitoring, alert thresholds, and real-time dashboards for critical FinTech protocols, payment rails, ACH file watching, and bank connectivity channels (Axway, SFTP, AS2).
  • Institutionalize SRE paradigms (SLIs, SLOs, Error Budgets, Reliability Reviews) across all software product squads, embedding reliability directly into the software development lifecycle (SDLC).
  • Establish proactive resilience practices, including failure injection, Chaos Engineering, and regular multi-AZ/multi-region Disaster Recovery (DR) simulations for critical financial services.
  • Partner with cloud and finance leadership to balance extreme uptime demands with cost efficiency across Splunk, Dynatrace, AWS telemetry storage, and log ingestion limits.
  • Oversee the high-severity Incident Management framework, ensuring 24/7 incident readiness, rapid mean time to detect (MTTD), and swift mean time to resolve/recover (MTTR).
  • Champion a culture of psychological safety through rigorous, blameless post-mortems and Root Cause Analyses (RCAs) to drive systemic platform fixes and prevent recurring outages.
  • Establish transparent reporting frameworks to translate uptime metrics, availability SLAs, error budget consumption, and platform risks directly to C-suite leadership (CTO, CIO).
  • Provide reliability oversight for moving critical functions off legacy platforms (Xymon, Puppet, SVN) to cloud-native AWS architectures without risking live financial traffic or data loss.
  • Enforce strict Infrastructure-as-Code (Terraform/Ansible) and GitLab CI/CD pipeline reliability, integrating automated canary deployments, security scans, and instant rollback mechanisms.
  • Partner with security teams to ensure all observability logging, tracing, and automation frameworks comply strictly with PCI-DSS, SOC 2, mTLS, and banking audit standards.
  • Build, mentor, and scale a high-performing team of Site Reliability Engineers, Observability Specialists, and Incident Response Leads in India.
  • Define technical career paths, performance benchmarks, and continuous learning opportunities in observability and SRE for mid-level and senior engineers.
  • Manage strategic relationships, licensing, contractual negotiations, and technical roadmaps with key observability and cloud partners (Splunk, Dynatrace, AWS, GitLab).
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service