Caddi-posted 21 days ago
Full-time • Mid Level
Chicago, IL
Fabricated Metal Product Manufacturing

As a Site Reliability Engineer at CADDi, you will build and secure infrastructure supporting our AI platform with special attention to safeguarding US customer data and supporting the Aerospace and Defense Industrial Base. You'll have strong ownership of US operations while collaborating with a global team of 150+ engineers in a fast-paced, high-growth environment.

  • Infrastructure & Architecture: Design, implement, and operate highly available, scalable, and fault-tolerant infrastructure primarily on GCP, but to include multi-cloud deployments. Optimize system performance, manage disaster recovery, and ensure cost-effectiveness.
  • Infrastructure as Code: Lead Terraform-based infrastructure development with security best practices, encrypted state management, and governance tools.
  • CI/CD & DevSecOps: Build robust pipelines supporting hundreds of developers and AI engineers. Integrate automated security testing, vulnerability scanning, and compliance checks throughout the development lifecycle.
  • Monitoring & Incident Response: Implement comprehensive observability strategies using Prometheus, Grafana, and ELK. Define SLOs/SLIs, manage error budgets, and lead incident response with blameless post-mortems.
  • Compliance & Security: Navigate complex regulatory requirements for U.S. Aerospace and Defense Industrial Base. Collaborate with security and legal teams on expanding compliance standards.
  • Automation & Collaboration: Reduce operational toil through Python, Go, or Bash automation. Work in a follow-the-sun model with global teams while taking primary responsibility for US platform partition incidents and operations.
  • For security reasons, the candidate must be a US Citizen, or a Permanent Resident (Green Card)
© 2024 Teal Labs, Inc
Privacy PolicyTerms of Service