About The Position

For a single team delivering components of distributed systems. Translates goals into a 1–2 quarter execution plan, sets coding, testing, and scalability practices, and provides hands-on oversight of performance tuning and load/perf testing. Guides the team in building fault-tolerant, in-service-upgradable components (redundancy, replication, failover) and in applying resiliency patterns (retries, circuit breakers, timeouts). Ensures robust observability (tests, alarms, dashboards, telemetry) and operational readiness via reviewed runbooks and standard procedures. Manages delivery of scoped features and correctness testing (including fault-injection/brownouts), and directs implementation of data replication/synchronization to maintain integrity and availability. Leads team incident response and root-cause efforts, enforces no-customer-downtime practices, and drives use of automation/IaC for troubleshooting. Oversees team security implementation (encryption, access controls), tracks remediation plans, verifies compliance documentation, and coaches adherence to change-management plans for safe patching, updates, and rollbacks. Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing [email protected] [[email protected]] or by calling 1-888-404-2494 in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Requirements

  • Experience in managing a team delivering components of distributed systems
  • Ability to translate goals into a 1-2 quarter execution plan
  • Proficiency in setting coding, testing, and scalability practices
  • Hands-on experience with performance tuning and load/perf testing
  • Experience guiding teams in building fault-tolerant, in-service-upgradable components (redundancy, replication, failover)
  • Experience guiding teams in applying resiliency patterns (retries, circuit breakers, timeouts)
  • Experience ensuring robust observability (tests, alarms, dashboards, telemetry)
  • Experience ensuring operational readiness via reviewed runbooks and standard procedures
  • Experience managing delivery of scoped features and correctness testing (including fault-injection/brownouts)
  • Experience directing implementation of data replication/synchronization
  • Experience leading team incident response and root-cause efforts
  • Experience enforcing no-customer-downtime practices
  • Experience driving use of automation/IaC for troubleshooting
  • Experience overseeing team security implementation (encryption, access controls)
  • Experience tracking remediation plans
  • Experience verifying compliance documentation
  • Experience coaching adherence to change-management plans for safe patching, updates, and rollbacks

Responsibilities

  • Translates goals into a 1–2 quarter execution plan
  • Sets coding, testing, and scalability practices
  • Provides hands-on oversight of performance tuning and load/perf testing
  • Guides the team in building fault-tolerant, in-service-upgradable components (redundancy, replication, failover)
  • Guides the team in applying resiliency patterns (retries, circuit breakers, timeouts)
  • Ensures robust observability (tests, alarms, dashboards, telemetry)
  • Ensures operational readiness via reviewed runbooks and standard procedures
  • Manages delivery of scoped features and correctness testing (including fault-injection/brownouts)
  • Directs implementation of data replication/synchronization to maintain integrity and availability
  • Leads team incident response and root-cause efforts
  • Enforces no-customer-downtime practices
  • Drives use of automation/IaC for troubleshooting
  • Oversees team security implementation (encryption, access controls)
  • Tracks remediation plans
  • Verifies compliance documentation
  • Coaches adherence to change-management plans for safe patching, updates, and rollbacks

Benefits

  • Flexible medical
  • Life insurance
  • Retirement options
  • Volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service