Site Reliability Engineer - Grid Orchestration Software

GE Vernova•Bellevue, WA
•$93,840 - $140,760•Onsite

About The Position

The GridOS Site Reliability Engineer is responsible for supporting the delivery, implementation, and reliability of GridOS software solutions across utility and energy distribution environments. This role combines software engineering and operations expertise to enable efficient deployment in customer environments while maintaining quality and delivering a high standard of customer satisfaction. The ideal candidate operates with a high degree of autonomy, applying sound technical judgment within established policies and procedures. This individual will play a key role in software implementation, troubleshooting, customization, and integration within customer environments, while effectively balancing project scope, timelines, and resource commitments.

Requirements

  • Bachelor's degree from an accredited university or college (or a high school diploma / GED with at least 6 years of experience in Job Family Group(s)/Function(s)).
  • 3+ years of experience in site reliability engineering, DevOps, systems engineering, or software engineering
  • 3+ Proficient with containerization and orchestration tools such as Docker and Kubernetes

Nice To Haves

  • Strong problem-solving skills with the ability to perform effectively under pressure
  • Excellent cross-functional communication and collaboration skills
  • Deep familiarity with Linux/Unix system administration and networking fundamentals (TCP/IP, DNS, HTTP).
  • Experience with monitoring and observability tools such as Prometheus, Grafana, or ELK
  • Strong oral and written communication skills.
  • Experience operating and supporting large-scale distributed systems
  • Familiarity with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps
  • Familiarity with database reliability, backup, and recovery processes

Responsibilities

  • Integrate Grid Orchestration Software (GridOS) solutions and perform comprehensive system testing.
  • Lead cross-functional incident response to resolve production issues and minimize customer impact.
  • Conduct reviews to drive corrective actions that decrease outage frequency and severity.
  • Design and maintain monitoring, logging, and alerting platforms (e.g., Prometheus, Grafana).
  • Collaborate with software teams to improve system resilience, release quality, and operational readiness.
  • Drive continuous improvement in disaster recovery, capacity planning, and system architecture design.
  • Develop scripts and workflows to streamline operations and reduce manual effort.
  • Mentor team members and serve as a technical resource for the organization.
  • Resolve customer technical issues, provide updates, and manage on-call responsibilities as required.

Benefits

  • medical, dental, vision, and prescription drug coverage
  • access to Health Coach from GE Vernova, a 24/7 nurse-based resource
  • access to the Employee Assistance Program, providing 24/7 confidential assessment, counseling and referral services
  • GE Vernova Retirement Savings Plan, a tax-advantaged 401(k) savings opportunity with company matching contributions and company retirement contributions, as well as access to Fidelity resources and financial planning consultants
  • tuition assistance
  • adoption assistance
  • paid parental leave
  • disability benefits
  • life insurance
  • 12 paid holidays
  • permissive time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service