Principal Software Engineer, Core Infrastructure

Oracle•Reston, VA
•$114,600 - $234,600

About The Position

As Oracle Cloud Infrastructure (OCI) continues its rapid expansion, we are seeking a skilled Software Engineer to join our newly established Cloud Performance Organization. This team plays a key role in addressing service inefficiencies, reducing cloud expenses, improving customer experience, and ensuring scalability. Your work will focus on optimizing the performance of OCI’s critical components, internal tools, and applications while fostering a culture of performance engineering. This is a greenfield opportunity to design and build new cloud services from the ground up. We are growing fast, still at an early stage, and working on ambitious new initiatives. You will be part of a team of smart, motivated, diverse people, and given the autonomy as well as support to do your best work. It is a dynamic and flexible workplace where you’ll belong and be encouraged. Leads development and architecture of scalable, elastic distributed systems for high-throughput, hyperscale workloads. Designs fault-tolerant, highly available systems with robust observability, testing, replication, and resilience mechanisms to meet SLOs. Drives operational readiness, production troubleshooting, and peer mentorship while implementing security, compliance, IaC, and automation for safe patching, upgrades, and rollbacks.

Requirements

  • Skilled Software Engineer to join our newly established Cloud Performance Organization.
  • Addressing service inefficiencies, reducing cloud expenses, improving customer experience, and ensuring scalability.
  • Optimizing the performance of OCI’s critical components, internal tools, and applications while fostering a culture of performance engineering.
  • Design and build new cloud services from the ground up.
  • Leads development and architecture of scalable, elastic distributed systems for high-throughput, hyperscale workloads.
  • Designs fault-tolerant, highly available systems with robust observability, testing, replication, and resilience mechanisms to meet SLOs.
  • Drives operational readiness, production troubleshooting, and peer mentorship while implementing security, compliance, IaC, and automation for safe patching, upgrades, and rollbacks.

Responsibilities

  • Lead development and architecture of scalable, elastic distributed systems for high-throughput, hyperscale workloads.
  • Define scalability requirements, optimize performance, and leverage distributed state management and data-plane platforms.
  • Design fault-tolerant, highly available systems using redundancy, replication, failover, load shedding, throttling, and rate limiting.
  • Establish SLOs, KPIs, telemetry, dashboards, and alerts to ensure reliability and performance.
  • Design performance, load, fault-injection, and brownout testing, and implement replication and synchronization for correctness and availability.
  • Proactively diagnose production issues, guide incident response and root cause analysis, and ensure operational readiness.
  • Enable in-service maintenance and upgrades with minimal customer impact.
  • Mentor engineers in troubleshooting and operational practices.
  • Implement encryption, access controls, and security remediation for multi-tenant environments.
  • Ensure compliance with applicable standards and maintain required documentation.
  • Develop and maintain IaC and automation for cloud infrastructure.
  • Enable safe and repeatable patching, updates, and rollbacks through effective change-management practices.
  • Manage moderately complex initiatives, prioritizing work, timelines, resources, and deliverables while providing technical oversight.
  • Collaborate across teams and stakeholders to align objectives and deliver solutions that meet business and customer needs.
  • Promote inclusive collaboration and diverse perspectives.
  • Analyze and resolve moderately complex issues, escalating critical concerns with clear assessments and recommended solutions.
  • Document and share effective problem-solving practices.
  • Stay current with industry trends and continuously develop technical skills.
  • Coach and mentor junior engineers and promote knowledge sharing.
  • Identify and implement improvements to processes, workflows, and team effectiveness.
  • Evaluate outcomes and incorporate stakeholder feedback.
  • Support talent development through candidate interviews, assessments, and hiring recommendations.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
  • Competitive benefits that support our people with flexible medical, life insurance, and retirement options.
  • Volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service