Sr Engineer, SRE TechOps CICD (Remote)

CrowdStrikeUSA CA Remote, CA
$140,000 - $215,000Remote

About The Position

CrowdStrike's internal SRE team owns the automation, reliability, and observability of the internal developer platform that thousands of CrowdStrike engineers rely on to build and deploy software rapidly, efficiently, and at scale. We provide the resilient infrastructure and operational rigor that let product, platform, and application teams ship with confidence, without needing to think about the systems underneath them. We're looking for a Senior Site Reliability Engineer on the Technical Operations team to bring deep, hands-on expertise across load balancers, relational and non-relational databases, message queues (Kafka, Pulsar, RabbitMQ, RedPanda), and caching layers (Redis/Valkey, Varnish), paired with strong SLI/SLO instincts. This is a technical-anchor role: you're the person the team routes hard problems to, and you set the bar for operational excellence through the quality of your own work and the judgment you bring to design and incident reviews. You'll build relationships with technical leaders across the organization, contribute to architectural direction for the services you own, and help position those services for the company's next level of scale.

Requirements

  • Must be eligible for CJIS clearance (requires U.S. citizenship or Green Card/permanent resident status).
  • 10+ years of experience working in large-scale production SRE or infrastructure environments.
  • 3+ years of experience leveraging and integrating AI-assisted workflows to increase engineering efficiency.
  • Bachelor's degree in computer science or another highly technical, scientific discipline, or equivalent work experience.
  • On-premise and cloud expertise deploying and operating CI/CD tools (Bazel, Jenkins, GitLab CI, GitHub Actions), IaC provisioning (Ansible, Chef, Puppet, Salt, Terraform), source code management (Bitbucket, GitHub, GitLab), and monitoring/observability platforms (Datadog, Grafana, Humio/LogScale, Honeycomb, New Relic, Prometheus, Splunk).
  • Experience creating, deploying, operating, and scaling applications on Kubernetes.
  • Extensive experience deploying and managing data infrastructure at scale (Cassandra, Postgres, MySQL, MongoDB, OpenSearch, Kafka, Redis/Valkey).
  • Proficiency in common scripting languages (Python, Go, Bash, PowerShell).
  • Experience with storage technologies (SAN, NAS, NFS, Object Storage).
  • Experience architecting and deploying big data systems.
  • Security-first mindset with a working understanding of cybersecurity principles.
  • Proven ability to make well-informed, timely decisions under ambiguity.
  • Ability to balance short-term operational needs against long-term strategic goals.
  • Self-directed learner who takes initiative in fast-moving environments.
  • Must be able to work with a distributed team across multiple time zones.
  • Meticulous attention to detail.
  • Proven experience utilizing AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency and drive business outcomes.

Nice To Haves

  • Knowledge of networking patterns and general network troubleshooting (Load balancers, DNS, VIPS, Routing, Firewall rules)
  • Knowledge and proven operation ability across multiple cloud hyperscalers such as AWS, Azure, GCP, Oracle
  • Experience building self-service / provisioning-automation platforms that reduce operational toil.
  • Experience with Active Directory / Windows Server and hybrid on-prem + cloud environments.
  • Experience with data science, machine learning, and ETL, principles and tooling such as Apache Airflow, Apache Spark, ect

Responsibilities

  • Own the availability and health of key services within the CICD environment, maintaining a holistic view of system health across the platform.
  • Build software and systems to manage platform infrastructure and applications, and drive automation for service deployment and operational workflows.
  • Carry on-call responsibility for owned services; drive incident response and blameless postmortems to root cause.
  • Gather and analyze metrics from operating systems and applications to support performance tuning and root cause analysis.
  • Lead system design discussions, production readiness reviews, and capacity planning exercises.
  • Evaluate and integrate agentic and AI-assisted workflows into existing team processes, and help teammates adopt them.
  • Mentor mid-level and junior engineers through code review, design pairing, and incident retrospectives.
  • Investigate and evaluate emerging technologies, and provide recommendations that support future roadmap goals.
  • Build and maintain automated reporting on service health and compliance.
  • Provide technical feedback and guidance on projects outside your core area of ownership, helping raise the bar across the broader engineering organization.
  • Partner with peer senior engineers and engineering leaders to drive cross-team reliability improvements.
  • Contribute to the Embedded SRE model, helping strengthen partnerships between SRE and the services teams.

Benefits

  • Market leader in compensation and equity awards
  • Comprehensive physical and mental wellness programs
  • Competitive vacation and holidays for recharge
  • Paid parental and adoption leaves
  • Professional development opportunities for all employees regardless of level or role
  • Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
  • Vibrant office culture with world class amenities
  • Great Place to Work Certified™ across the globe
  • health insurance
  • 401k
  • paid time off
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service