Site Reliability Engineer, Global Banking & Markets, Vice President

Goldman SachsNew York, NY
$150,000 - $300,000Onsite

About The Position

At Goldman Sachs, our Engineers don't just make things - we make things possible. Change the world by connecting people and capital with ideas. Solve the most challenging and pressing engineering problems for our clients. Join our engineering teams that build massively scalable software and systems, architect low latency infrastructure solutions, proactively guard against cyber threats, and leverage machine learning alongside financial engineering to continuously turn data into action. Create new businesses, transform finance, and explore a world of opportunity at the speed of markets. Within the firm's Global Banking & Markets business, the Site Reliability Engineering (SRE) team ensures the availability, resilience, and performance of core business services that underpin a global 24x7 trading operation. Working across Global Markets' front, middle, and back office functions, you will engineer reliability while balancing stringent non-functional demands for availability, latency, and resilience as well as complex, evolving business requirements. Want to push the limit of digital possibilities? Start here.

Requirements

  • 8+ years of professional software / reliability engineering experience, with strong command of at least one major language (Java 17+ preferred), including concurrency, collections, and modern language features.
  • Demonstrated risk acumen — the ability to identify, quantify, and mitigate operational and technical risk in a regulated financial services environment.
  • Excellent communication and stakeholder-coordination skills — proven ability to connect the right people quickly and drive resolution across geographically distributed, technical and non-technical audiences.
  • Proven experience running high-availability production environments: SLIs/SLOs, error budgets, on-call, incident command, and post-incident reviews.
  • Strong understanding of cloud infrastructure (GCP, AWS), container orchestration (Kubernetes, Docker), and infrastructure-as-code.
  • Working knowledge of AI models and AI-assisted engineering tools (e.g., Claude Code, GitHub Copilot Agent Mode, Devin, Gemini Code Assist), including the ability to govern AI agents, critically assess their output, and maintain quality over AI-generated work.
  • Experience building event-driven and distributed systems, including messaging platforms (e.g., Apache Kafka), delivery guarantees, and resilience strategies.
  • Strong SDLC and automation practices: version control, CI/CD pipelines, automated build/test/deploy workflows, and code quality tooling.
  • Solid observability discipline: application instrumentation, distributed tracing, structured logging, and metrics-driven operations.
  • Ability to rapidly navigate, understand, and debug large and unfamiliar codebases — with and without AI assistance.

Nice To Haves

  • Chaos engineering, capacity planning, load/performance testing, and production support in high-availability, latency-sensitive environments.
  • Spring Boot, gRPC / Protocol Buffers, integration/orchestration frameworks (e.g., Apache Camel, Spring Integration), and pipeline/adapter patterns (retry, dead-letter queues, error isolation).
  • Cloud platforms (GCP, AWS), Kubernetes/Docker, JVM tuning for containerized workloads, and infrastructure-as-code (Terraform, Helm).
  • Applying AI models to operational use cases — anomaly detection, log analysis, automated remediation, and agentic operations.
  • Prometheus, Grafana, OpenTelemetry, and SLO tooling.
  • Data modeling, SQL/NoSQL databases, caching strategies, and performance optimization in latency-sensitive systems.
  • Enterprise security patterns; authentication protocols, mutual TLS, secrets management, and certificate rotation.
  • Equities, post-trade, or financial services experience; trade lifecycle concepts, position management, reconciliation, and multi-system migration environments.
  • Asynchronous / non-blocking I/O frameworks (e.g., Vert.x, Netty), multi-region / BCP architectures, and open-source contribution experience.

Responsibilities

  • Define and defend Service Level Objectives (SLOs), error budgets, and reliability standards for critical trading services, with risk always front of mind.
  • Identify systemic risks before they materialize, automate away repetitive operational work, and strengthen the resilience posture of the platform.
  • Act as a trusted coordinator during incidents — rapidly mobilizing the right engineers, domain experts, and stakeholders across a globally distributed organization, and communicating clearly with both technical and non-technical audiences.
  • Orchestrate AI coding and operations agents to accelerate root-cause analysis, remediation, and automation while maintaining mastery, quality, and production fitness over all AI-generated work.
  • Design and operate high-availability, multi-region, event-driven services on a modern cloud-native platform, setting the reliability and architectural standard for years to come.
  • Design, build, and operate high-availability, multi-region, cloud-native services with security and comprehensive observability (metrics, distributed tracing, structured logging) built in at every layer.
  • Establish and manage SLIs, SLOs, and error budgets; drive blameless post-incident reviews and translate findings into durable engineering improvements.
  • Lead incident response for latency-sensitive, high-throughput trade lifecycle systems — quickly diagnosing issues, coordinating cross-functional responders, and communicating status to stakeholders.
  • Develop event-driven architectures, multi-stage processing pipelines, and optimized data paths for high-throughput trade lifecycle management.
  • Apply strong risk acumen to change management, capacity planning, and resilience testing (chaos engineering, failover, and BCP drills).
  • Partner with engineers, domain experts, and global stakeholders to understand production processes, challenge entrenched assumptions in a cloud-centric, AI-driven world, and drive modernization.
  • Multiply your impact with a modern, AI-centric toolchain, orchestrating AI agents across the SDLC and operations to rapidly comprehend large codebases, generate production-quality automation, and accelerate delivery.

Benefits

  • Discretionary bonus
  • Competitive benefits and wellness offerings
  • Training and development opportunities
  • Firmwide networks
  • Mindfulness programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service