Director, Core Infrastructure Engineering

Oracle•Seattle, WA
•$169,800 - $355,400

About The Position

Leads a multiple teams to implement strategies for the architecture and delivery of interdependent, scalable distributed systems that meet organizational and customer demands. Orchestrates cross-group optimization for high‑throughput, large‑scale data processing; aligns stakeholders on scalability requirements; and oversees elastic designs and effective use of data plane platforms. Provides strategic oversight for fault‑tolerant, in‑service‑upgradable architectures, sets direction for partition‑aware design choices, and leads initiatives to harden networks via load‑shedding, throttling, and rate‑limiting. Establishes expectations for formal verification and peer reviews, and sets SLO‑aligned durability and availability standards across the department. Drives KPI and telemetry strategies; directs creation of complex dashboards and alerting for proactive health assurance; and ensures functional/correctness validation, data replication, and synchronization meet organizational needs. Guides organization‑wide incident management and operational readiness, eliminating customer maintenance windows and ensuring consistent SOPs. Provides strategic security guidance (encryption, access controls), oversees remediation and compliance documentation, and sponsors automation (IaC) and change‑management alignment so systems can be safely patched, updated, and rolled back at scale.

Requirements

  • Experience leading multiple teams to implement strategies for the architecture and delivery of interdependent, scalable distributed systems.
  • Experience orchestrating cross-group optimization for high-throughput, large-scale data processing.
  • Experience aligning stakeholders on scalability requirements.
  • Experience overseeing elastic designs and effective use of data plane platforms.
  • Experience providing strategic oversight for fault-tolerant, in-service-upgradable architectures.
  • Experience setting direction for partition-aware design choices.
  • Experience leading initiatives to harden networks via load-shedding, throttling, and rate-limiting.
  • Experience establishing expectations for formal verification and peer reviews.
  • Experience setting SLO-aligned durability and availability standards.
  • Experience driving KPI and telemetry strategies.
  • Experience directing creation of complex dashboards and alerting for proactive health assurance.
  • Experience ensuring functional/correctness validation, data replication, and synchronization meet organizational needs.
  • Experience guiding organization-wide incident management and operational readiness.
  • Experience eliminating customer maintenance windows.
  • Experience ensuring consistent SOPs.
  • Experience providing strategic security guidance (encryption, access controls).
  • Experience overseeing remediation and compliance documentation.
  • Experience sponsoring automation (IaC) and change-management alignment.
  • Experience managing complex projects or initiatives, monitoring timelines, deliverables, and budgets.
  • Experience delegating work, setting priorities, and ensuring alignment with business needs.
  • Experience adjusting resources or project timelines in anticipation of business changes.
  • Experience leading cross-functional collaborative efforts.
  • Experience building and maintaining partnerships with business leaders, stakeholders, and/or customers.
  • Experience analyzing highly complex data and/or information to identify solutions to ambiguous issues.
  • Experience identifying root causes to prevent recurrence of issues.
  • Experience pursuing strategic learning opportunities to maintain expertise and apply best practices.
  • Experience creating opportunities for team members and leaders to build expertise.
  • Experience identifying skill gap trends across the organization.
  • Experience evaluating the efficiency of learning strategies and recommending adjustments.
  • Experience empowering teams to own the development and implementation of ideas that increase efficiency and effectiveness of processes, protocols, and workflows.
  • Experience coaching teams to gain buy-in for ideas and seek feedback.
  • Experience prioritizing and reviewing the roadmap of improvement initiatives.
  • Experience driving performance across teams through tailored feedback and coaching.
  • Experience ensuring consistency in the application of talent development procedures.
  • Experience socializing performance expectations across the organization.
  • Experience aligning individual development goals with organizational strategic initiatives.
  • Experience collaborating with HR to implement talent strategy through hiring and promotion processes.

Responsibilities

  • Implements strategies across multiple teams or groups for the architecture and design of interdependent scalable distributed systems, including the use of distributed state management tools, ensuring organizational and system demands are met.
  • Spearheads code and/or system optimization initiatives for large-scale data processing and high-throughput requirements across multiple areas, driving improvements that support hyper-scale systems.
  • Facilitates collaborations to define system scalability requirements, ensuring the defined requirements meet customer expectations.
  • Oversees the design of interdependent systems to scale with elasticity (e.g., effectively scaling both up and down).
  • Drives the effective use and implementation of data plane platforms for large-scale data operations.
  • Provides strategic oversight for the architecture of fault-tolerant interdependent systems capable of withstanding in-service updates by overseeing implementation across teams of redundancy, replication, and automatic failover mechanisms.
  • Influences and sets direction for designing systems to effectively handle service disruptions (e.g., network partitions) by prioritizing consistency, availability, or partition tolerance.
  • Leads strategic optimization initiatives for handling network unreliability, including directing the design of load-shedding, throttling, and rate-limiting techniques.
  • Holds teams accountable for leveraging formal verification techniques to verify system designs and conduct peer reviews across teams.
  • Drives the design of systems that are durable and adhere to service level objectives (SLOs), developing standards for availability and durability of other computing services across the department.
  • Drives strategies for defining key performance indicators (KPIs) and telemetry to identify risks, gaps, or cyclical dependencies in running systems, ensuring alignment with organizational goals.
  • Directs the creation and customization of complex dashboards, telemetry systems, and alerting mechanisms that proactively monitor and ensure optimal system health across teams.
  • Implements strategies to effectively determine if systems are meeting functional and correctness requirements, and encourages teams to identify improvement opportunities.
  • Provides thought leadership on processes for formally verifying complex features to ensure system design correctness.
  • Oversees the implementation of data replication and synchronization techniques, ensuring data integrity and availability across the organization.
  • Provides strategic oversight for diagnosing, debugging, and resolving issues in active systems to support ongoing operation.
  • Directs strategies within teams to prevent interruptions, ensuring no maintenance windows are required for customers and users when resolving issues.
  • Drives alignment across teams for operational readiness protocol and standard operating procedures.
  • Provides expert guidance for complex incident response and root cause investigations.
  • Provides strategic guidance in architecting robust security measures to protect data and applications in multi-tenant environments, ensuring encryption techniques and access controls are implemented.
  • Oversees execution of remediation plans to address identified security gaps, promoting significant improvements and continuous advancement of security measures.
  • Drives documentation efforts and ensures cloud infrastructure compliance with industry standards and regulations.
  • Provides strategic guidance across teams on developing and maintaining automation scripts and tools (e.g., Infrastructure as Code (IaC)) to manage cloud infrastructure.
  • Drives strategic alignment of change management plans for patching, updating, and rolling back applications, and oversees that system designs allow for automation of these processes.
  • Oversees and guides multiple teams on managing complex projects or initiatives, monitoring timelines, deliverables, and budgets (when applicable) to ensure strategic objectives are met.
  • Serves as a role model for appropriately delegating work, setting priorities, and ensuring alignment with business needs.
  • Coaches others on adjusting resources or project timelines in anticipation of business changes.
  • Role models leading cross-functional collaborative efforts to ensure alignment of expectations and strategic objectives.
  • Empowers teams to build and maintain partnerships with business leaders, stakeholders, and/or customers to address barriers and contribute to organizational success.
  • Drives transparency and inclusivity by modeling actively seeking, listening to, and leveraging diverse perspectives.
  • Shares problem-solving strategies across teams, providing oversight on complex operational and/or technical issues, as needed.
  • Coaches teams on analyzing highly complex data and/or information to identify solutions to ambiguous issues.
  • Provides direction on identifying root causes to prevent recurrence of issues.
  • Pursues strategic learning opportunities to maintain expertise and apply best practices at the organizational level.
  • Creates opportunities for team members and leaders to build their expertise in new areas, coaching them to build innovative skills.
  • Identifies skill gap trends across the organization and upholds a culture that places significant emphasis on sharing knowledge and pursuing learning opportunities that advance the organization.
  • Evaluates the efficiency of learning strategies and recommends adjustments as needed.
  • Empowers teams to own the development and implementation of ideas that increase the efficiency and effectiveness of processes, protocols, and workflows across the department.
  • Coaches teams to gain buy-in for ideas and to seek feedback on approaches and methods for continued improvement.
  • Prioritizes and reviews the roadmap of improvement initiatives to ensure alignment with strategic direction and maximize return on investments.
  • Serves as a role model for driving performance across teams through tailored feedback and coaching in alignment with performance management processes, guidelines, and expectations.
  • Drives consistency in the application of talent development procedures and socializes performance expectations across the organization.
  • Ensures that individual development goals are aligned with organizational strategic initiatives.
  • Collaborates with HR to implement talent strategy through hiring and promotion processes.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service