This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack. This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership. You will join a team responsible for ensuring the stability, reliability, and operational excellence of a critical UI service running in production. The role involves supporting the deployment, operation, and reliability of production services running on Kubernetes. You will monitor service health and investigate production incidents across distributed applications. Participation in on-call support, incident response, root cause analysis, postmortems, and reliability improvements is expected. You will also troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams, and support CI/CD, GitOps-based deployments, observability, and production monitoring. You will work within a client-directed backlog and established priorities.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed