Senior Engineer, Reliability

LPL FinancialFort Mill, SC
$101,558 - $169,229

About The Position

The Senior Engineer, Observability and Platform Stability is responsible for ensuring the reliability, availability, and performance of enterprise observability platforms and supporting applications. This role drives operational excellence through proactive monitoring, incident response, platform maintenance, automation, and continuous improvement initiatives. The position partners closely with Site Reliability Engineering (SRE), Platform Engineering, DevOps, Operations, and Incident Management teams to enhance platform stability, mature CI/CD practices, and support strategic technology initiatives. The ideal candidate brings strong cloud technology experience and a proven ability to support production environments while delivering actionable insights through observability capabilities.

Requirements

  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or a related field.
  • 6+ years of experience in SRE, platform engineering, production support, enterprise application operations, or related technology environments.
  • 5+ years of experience supporting enterprise platforms, including AWS, Dynatrace, ELK, ServiceNow, and SolarWinds.
  • Experience troubleshooting complex production incidents within enterprise-scale technology environments.
  • Experience executing technology changes through formal change management and release management processes.

Nice To Haves

  • Experience within financial services or another regulated industry.
  • Experience implementing or supporting CI/CD pipelines, release automation, and deployment processes across development, testing, and production environments.
  • Experience with observability platforms, monitoring tools, performance dashboards, or application monitoring solutions.
  • Experience collaborating with offshore, nearshore, or global delivery teams.

Responsibilities

  • Partner with Product Owners and engineering teams to provide operational and release support for technology initiatives.
  • Ensure platform changes meet operational readiness requirements, including rollback procedures, runbook documentation, integration standards, and support handoffs.
  • Maintain production stability throughout platform upgrades, enhancements, and enterprise initiatives.
  • Support platform ownership transitions and operational readiness activities across global delivery teams.
  • Troubleshoot application and platform issues to restore services and minimize business impact.
  • Serve as an escalation point for complex production incidents and operational challenges.
  • Participate in incident triage, root cause analysis, corrective action planning, and resolution activities.
  • Collaborate with engineering teams to implement long-term solutions that reduce recurring incidents.
  • Support mission-critical environments through timely incident response and service restoration.
  • Conduct health checks, configuration reviews, and performance assessments to identify operational risks.
  • Support reliability initiatives focused on improving recovery times, reducing incident recurrence, and minimizing change-related defects.
  • Validate vendor releases, hotfixes, and configuration changes prior to production deployment.
  • Partner with observability and analytics teams to enhance monitoring, alerting, and issue detection capabilities.
  • Execute platform changes through established SDLC, change management, and release management processes.
  • Collaborate with Platform Engineering, DevOps, Quality Engineering, and Scrum teams to improve release and deployment practices.
  • Support the adoption of source control, environment separation, release automation, and CI/CD capabilities.
  • Ensure solutions are testable, deployable, and operationally supported before and after production implementation.
  • Maintain runbooks, support documentation, configuration records, incident playbooks, and release procedures.
  • Document incident findings, lessons learned, and process improvement opportunities.
  • Contribute to the development of standardized, repeatable, and scalable operational practices.

Benefits

  • 401K matching
  • health benefits
  • employee stock options
  • paid time off
  • volunteer time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service