The Splunk Agent Resilience team is defining the future of AI resilience. Together our team provides scalable, cost-effective evaluation and guardrails that ensure AI agents behave as intended, improving reliability and reducing risks. This unified approach empowers our customers to expertly deploy and manage AI-powered applications with enhanced observability and control. As a Staff Site Reliability Engineer (SRE), you will provide technical leadership for the reliability, scalability, and operational architecture of Splunk Agent Observability's platform. You will define the long-term reliability strategy, lead major infrastructure initiatives, and drive engineering excellence across deployment automation, production operations, and platform resiliency. In addition to owning complex production systems, you will influence engineering direction across teams and establish best practices for operating large-scale cloud and on-prem deployments.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior