About The Position

This position will be full-time on-site at Oracle's offices located in Nashville, TN. Relocation assistance may be available in accordance with Oracle’s relocation policies. Candidates should expect a minimum of 25% travel, with additional travel as business needs require. As Director of Building Automation, you will provide strategic and organizational leadership for the reliability engineering function supporting Oracle Cloud Infrastructure’s mission-critical data center portfolio. You will own the vision, operating model, engineering standards, and portfolio programs that improve infrastructure availability, maintainability, resilience, and lifecycle performance at scale. This role requires significant hands-on and leadership experience within mission-critical environments. Direct data center experience is preferred. Candidates must have demonstrated experience supporting the electrical, mechanical, controls, and operational systems required to maintain continuous operations in mission-critical environments. You will lead managers, engineers, analysts, and technical programs responsible for reliability engineering, asset performance, predictive maintenance, failure analysis, defect elimination, and lifecycle risk management across OCI's data center infrastructure. You will partner with senior leaders across Data Center Operations, Engineering, Design, Construction, Commissioning, Automation, Procurement, and other infrastructure organizations to translate operational experience and engineering data into long-term reliability strategy. You will ensure lessons learned from incidents, equipment performance, maintenance activities, and portfolio trends result in durable improvements to standards, designs, operating practices, and investment priorities. Success in this role requires the ability to operate at both strategic and technical levels—setting multi-year direction for the reliability organization while maintaining sufficient engineering depth and data center operational knowledge to challenge assumptions, assess complex infrastructure risks, and drive disciplined decision-making across a rapidly scaling global portfolio.

Requirements

  • 10+ years of progressive engineering, reliability, maintenance, critical facilities, or infrastructure experience, including significant experience directly supporting mission-critical environments.
  • Demonstrated experience working within mission-critical operations or engineering environments where infrastructure availability, redundancy, maintenance execution, and operational risk directly affect service continuity.
  • 5+ years of progressive leadership experience, including responsibility for engineering managers, senior technical professionals, or large multi-site technical organizations and programs.
  • Demonstrated technical knowledge of mission-critical infrastructure, including experience with electrical distribution, UPS systems, generators, mechanical cooling systems, controls/automation, and integrated facility operations.
  • Demonstrated experience developing and implementing reliability, maintenance, asset-management, or operational excellence strategies across mission-critical infrastructure.
  • Strong working knowledge of reliability engineering methodologies, including structured root cause analysis, FMEA/FMECA, RCM, criticality analysis, defect elimination, reliability metrics, and lifecycle risk management.
  • Demonstrated experience evaluating infrastructure failures, operational events, equipment performance, maintenance effectiveness, and systemic reliability risks within mission-critical environments.
  • Demonstrated ability to use operational and engineering data to identify systemic risks, establish priorities, and influence significant technical or business decisions.
  • Experience leading cross-functional initiatives involving Data Center Operations, Engineering, Design, Construction, Commissioning, Procurement, OEMs, vendors, and other technical stakeholders.
  • Experience establishing engineering governance, standards, KPIs, and management mechanisms across multiple sites, regions, or infrastructure programs in mission-critical environments.
  • Demonstrated ability to communicate complex technical risks, tradeoffs, and investment recommendations to senior and executive leadership.
  • Bachelor’s degree in Electrical Engineering, Mechanical Engineering, Industrial Engineering, Systems Engineering, Reliability Engineering, or a related technical discipline; or equivalent relevant industry experience.

Nice To Haves

  • Direct data center experience supporting mission-critical infrastructure and operations is preferred.
  • Experience leading reliability engineering, critical facilities engineering, or asset-management organizations within hyperscale, colocation, or large-scale data centers.
  • Experience supporting geographically distributed or global data center portfolios.
  • Deep technical expertise in one or more critical infrastructure domains, with broad working knowledge across electrical distribution, UPS, standby generation, mechanical cooling, controls/automation, and integrated facility operations.
  • Experience developing and scaling predictive maintenance, condition-based monitoring, failure trend analysis, asset health modeling, and equipment risk-ranking programs within data center environments.
  • Advanced knowledge of reliability, availability, and maintainability analysis; maintenance strategy optimization; spare parts planning; lifecycle modeling; and total cost of ownership.
  • Experience governing commissioning, operational acceptance, maintenance program design, and readiness of new or modified mission-critical data center infrastructure.
  • Experience with CMMS/EAM, DCIM, EPMS, BMS, monitoring, telemetry, analytics, and automation platforms used to manage critical data center infrastructure.
  • Experience developing portfolio-level KPI frameworks, executive dashboards, reliability reviews, and risk-governance mechanisms.
  • Experience partnering with OEMs and strategic suppliers to address systemic equipment issues, improve product reliability, and influence equipment roadmaps or specifications.
  • Experience incorporating operational lessons learned into engineering standards, design requirements, equipment specifications, and capital investment decisions.
  • Demonstrated experience managing organizational growth, workforce planning, resource prioritization, and technical capability development across geographically distributed teams.
  • Certified Maintenance & Reliability Professional (CMRP) preferred.
  • Certified Reliability Engineer (CRE) preferred.
  • ASQ, SMRP, or equivalent reliability, maintenance, engineering, or quality certifications are a plus.
  • Data center or critical-environment credentials, including relevant Uptime Institute training or certifications, are a plus.
  • OEM, controls, analytics, asset-management, or condition-monitoring training applicable to critical data center infrastructure is a plus.
  • Advanced training or certification in FMEA/FMECA, RCA, RCM, Lean, Six Sigma, or structured problem-solving methodologies is a plus.
  • Advanced technical or business degree is a plus.

Responsibilities

  • Lead the Reliability Engineering organization supporting multiple regions, sites, and infrastructure programs across OCI's mission-critical data center portfolio.
  • Define the multi-year reliability engineering strategy, organizational roadmap, operating model, and investment priorities required to improve data center infrastructure resilience and support OCI's continued growth.
  • Establish and govern portfolio-wide reliability engineering standards and methodologies, including FMEA/FMECA, RCA, Reliability-Centered Maintenance (RCM), criticality assessment, defect elimination, reliability growth, and continuous improvement practices.
  • Build and lead a high-performing organization of managers, engineers, analysts, and technical specialists, establishing clear accountability, technical expectations, career development, and succession plans.
  • Own portfolio-level programs that improve reliability across critical data center infrastructure, including electrical distribution, UPS systems, generators, mechanical cooling systems, controls, automation, and supporting facility systems.
  • Establish a comprehensive reliability measurement framework that provides leadership with visibility into asset health, failure trends, systemic risks, repeat events, corrective actions, maintenance effectiveness, and lifecycle exposure.
  • Define reliability KPIs, targets, governance mechanisms, and executive reporting that enable data-driven prioritization of operational and engineering investments.
  • Establish governance for corrective and preventive actions resulting from incidents, root cause analyses, audits, equipment failures, and reliability trend reviews, ensuring actions are completed, verified for effectiveness, and sustained.
  • Drive systematic identification and elimination of recurring and systemic failure modes across the data center portfolio rather than relying solely on site-specific remediation.
  • Sponsor the development and adoption of predictive and condition-based maintenance capabilities, including monitoring, analytics, automation, asset health modeling, and emerging technologies that improve early detection of equipment degradation and failure risk.
  • Partner with Data Center Operations leadership to continuously improve maintenance strategy, operational readiness, troubleshooting practices, procedures, failure response, and infrastructure risk management.
  • Provide reliability governance and technical leadership for commissioning, acceptance testing, operational handover, major maintenance, retrofits, capacity expansion, and infrastructure lifecycle decisions.
  • Establish portfolio approaches to asset lifecycle management, including equipment health, utilization, failure history, remaining useful life, obsolescence, spare parts strategy, replacement planning, and end-of-life risk.
  • Partner with Design, Construction, Engineering, and Procurement leadership to ensure lessons from operating facilities influence equipment specifications, design standards, redundancy strategies, maintainability requirements, vendor selection, and total cost of ownership.
  • Develop mechanisms to convert site-level events and engineering findings into portfolio-wide standards, design changes, maintenance improvements, and risk-reduction programs.
  • Lead technical and business reviews of significant reliability risks and provide clear recommendations regarding mitigation strategies, priorities, investment requirements, and residual operational risk.
  • Develop strong partnerships with equipment manufacturers, service providers, and technology partners to improve equipment performance, failure intelligence, serviceability, and long-term reliability.
  • Establish effective operating rhythms for the organization, including portfolio reviews, technical reviews, risk escalation, program governance, resource prioritization, and executive communications.
  • Represent Reliability Engineering in senior leadership discussions involving infrastructure risk, operational performance, capacity growth, capital planning, and long-term data center strategy.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service