Director, Performance Engineering and Test Automation

Blue Shield of CaliforniaOakland, CA
Hybrid

About The Position

The Stellarus Technology team builds and operates the platform, services, and products that support payer operations, member experiences, customer integrations, and internal product delivery. The Director, Performance Engineering and Test Automation will report to the VP, Engineering. In this role, you will create the automated performance and reliability validation systems that give engineering teams clear evidence before services reach production. You will own the strategy and implementation for performance testing, load testing, scalability validation, resilience testing, synthetic monitoring, test data automation, and release gates across the Stellarus platform. You will work directly with developers to define the standards, patterns, tooling, and expectations teams use when they design, build, test, and operate services, including where AI and AI-assisted engineering tools can improve quality, speed, and signal. This is a new role for a builder who can define the practice, implement the tooling, partner with engineering teams, and create a culture where performance and reliability are validated continuously. The role combines performance engineering, platform quality, automation, observability, and production readiness. Our leadership model is about developing great leaders at all levels and creating opportunities for our people to grow personally, professionally, and financially. We are looking for leaders energized by creative and critical thinking, building and sustaining high-performing teams, getting results the right way, and fostering continuous learning.

Requirements

  • Bachelor’s degree or High School Diploma/GED and 4 years of additional relevant experience in lieu of a degree.
  • 10 years of software engineering, quality engineering, performance engineering, platform engineering, SRE, or related technical experience.
  • 6 years leading teams or technical programs focused on performance, test automation, reliability, platform quality, or production readiness.
  • Strong hands-on background building automated test systems, load testing frameworks, performance pipelines, or reliability validation tooling.
  • Experience testing distributed systems, microservices, APIs, event-driven architectures, and cloud-native applications.
  • Practical understanding of latency, throughput, concurrency, queueing, backpressure, retries, caching, database performance, resource saturation, and failure modes in production systems.
  • Experience with tools such as k6, Gatling, JMeter, Locust, Playwright, Cypress, Jest, Grafana, Prometheus, OpenTelemetry, Datadog, New Relic, Azure Monitor, or similar platforms.
  • Experience integrating automated tests into CI/CD pipelines and making test results useful to engineering teams.
  • Strong understanding of Kubernetes, containers, cloud infrastructure, API gateways, message queues, relational databases, and service observability.
  • Ability to work with engineers at code, architecture, infrastructure, and operational levels.
  • Experience defining performance benchmarks, synthetic workloads, test data strategies, production-like environments, and release readiness criteria.
  • Strong communication skills. You can explain performance risk clearly and turn findings into practical engineering action.

Nice To Haves

  • Experience evaluating or applying AI-assisted engineering tools for test generation, analysis, automation, or developer productivity is preferred.
  • Healthcare, payer, regulated-industry, SOC 2, HIPAA, or security-sensitive platform experience is preferred.
  • Automated performance and load testing for APIs and distributed workflows
  • Synthetic monitoring and production-like traffic modeling
  • Observability-driven debugging and regression detection
  • AI-assisted test generation, failure analysis, and engineering workflow automation
  • CI/CD quality gates and release readiness automation
  • Database and message-queue performance analysis
  • Kubernetes workload tuning and capacity planning
  • End-to-end test automation for user-facing and service-level workflows

Responsibilities

  • Build the performance engineering and automated testing function for Stellarus' cloud-native platform.
  • Design and implement fully automated performance, load, stress, soak, regression, and scalability test systems for services, APIs, workflows, and products.
  • Establish automated quality gates in CI/CD so teams get clear feedback before code reaches production-like environments.
  • Define service-level performance expectations, test coverage standards, baseline thresholds, release criteria, and escalation paths for performance regressions.
  • Partner directly with developers to define practical standards for testing microservices, event-driven workflows, APIs, customer-facing applications, and internal platform services.
  • Use AI-assisted tools where they make sense to accelerate test generation, synthetic scenario creation, log analysis, regression detection, documentation, and developer feedback loops.
  • Create production-like test environments, test data strategies, traffic models, synthetic workloads, and repeatable test scenarios that reflect real healthcare payer workflows.
  • Use observability data to connect test results to system behavior, including latency, throughput, error rates, queue depth, resource saturation, database performance, cold starts, retries, and downstream failures.
  • Help teams identify bottlenecks in application code, APIs, database queries, infrastructure, Kubernetes configuration, service-to-service calls, and event processing.
  • Build dashboards and reporting that make performance trends visible and actionable for engineers, product leaders, and technology leadership.
  • Collaborate with developers to define standards for automated end-to-end testing, contract testing, integration testing, and performance regression testing across the Nx monorepo.
  • Work with platform engineering to integrate testing into GitHub Actions, AKS, Dapr, Azure Service Bus, APIM, Terraform-managed environments, and ArgoCD deployment workflows.
  • Lead incident and launch readiness reviews where performance, scalability, or reliability risk is material.
  • Mentor developers, architects, and managers on performance design, test automation, failure analysis, and production readiness.
  • Build a small high-leverage team over time, while remaining close enough to the tooling and technical details to set direction credibly.

Benefits

  • Opportunities for people to grow personally, professionally, and financially.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service