Principal Software Engineer, Core Infrastructure

OracleSanta Clara, CA
$114,600 - $234,600

About The Position

Leads development and begins architecting components of scalable, elastic distributed systems. Defines and enforces scalability requirements for owned components; optimizes code and data paths for high‑throughput, hyper‑scale workloads; and leverages data plane platforms for large‑scale retrieval, storage, and processing. Designs fault‑tolerant, in‑service‑upgradable systems using redundancy, replication, failover, and policies for partitions, applying load‑shedding, throttling, and rate‑limiting to handle network unreliability while meeting SLOs. Establishes KPIs and telemetry; builds proactive dashboards and alerts; and designs complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability. Proactively diagnoses and resolves production issues, mentors peers, and ensures operational readiness. Implements robust security controls, executes remediation, maintains compliance documentation, and develops IaC and automation that enable safe patching, updates, and rollbacks within change‑management plans. Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives. True innovation starts when everyone is empowered to contribute. That’s why we’re committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs. We’re committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing [email protected] [[email protected]] or by calling 1-888-404-2494 in the United States. Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans’ status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Requirements

  • Scalable, elastic distributed systems architecture
  • Scalability requirements definition and enforcement
  • Code and data path optimization for high-throughput, hyper-scale workloads
  • Leveraging data plane platforms for large-scale retrieval, storage, and processing
  • Design of fault-tolerant, in-service-upgradable systems
  • Use of redundancy, replication, failover, and policies for partitions
  • Application of load-shedding, throttling, and rate-limiting
  • Meeting SLOs under network unreliability
  • Establishment of KPIs and telemetry
  • Building proactive dashboards and alerts
  • Design of complex validation (fault injection, brownouts), replication, and synchronization
  • Proactive diagnosis and resolution of production issues
  • Mentoring peers
  • Ensuring operational readiness
  • Implementation of robust security controls
  • Remediation execution
  • Maintenance of compliance documentation
  • Development of IaC and automation for patching, updates, and rollbacks
  • Adherence to change-management plans

Responsibilities

  • Leads development and begins architecting components of scalable, elastic distributed systems.
  • Defines and enforces scalability requirements for owned components.
  • Optimizes code and data paths for high‑throughput, hyper‑scale workloads.
  • Leverages data plane platforms for large‑scale retrieval, storage, and processing.
  • Designs fault‑tolerant, in‑service‑upgradable systems using redundancy, replication, failover, and policies for partitions.
  • Applies load‑shedding, throttling, and rate‑limiting to handle network unreliability while meeting SLOs.
  • Establishes KPIs and telemetry.
  • Builds proactive dashboards and alerts.
  • Designs complex validation (fault injection, brownouts), replication, and synchronization for correctness and durability.
  • Proactively diagnoses and resolves production issues.
  • Mentors peers.
  • Ensures operational readiness.
  • Implements robust security controls.
  • Executes remediation.
  • Maintains compliance documentation.
  • Develops IaC and automation that enable safe patching, updates, and rollbacks within change‑management plans.

Benefits

  • Flexible medical
  • Life insurance
  • Retirement options
  • Volunteer programs
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service