Site Reliability Engineer II

Kalmbach Feeds Inc•Upper Sandusky, OH
•Hybrid

About The Position

Join a small SRE team at the start of an exciting build: we are shaping the next generation of our hybrid platform across company data centers and Microsoft Azure—not simply maintaining someone else’s established setup. You’ll have direct ownership and a real voice in the architecture, standards, and technologies we put into production. The platform will include production RKE2 clusters managed with Rancher, Azure Kubernetes Service (AKS), hybrid workloads, and disaster recovery, while storage, networking, and data-platform designs are still open to influence. This hands-on role spans racks to cloud, and what you help design will become what the company runs. You need not know every tool on day one, but should learn quickly and bring a practical, curious approach.

Requirements

  • 3+ years in SRE, platform, systems engineering, or DevOps, with hands-on experience operating and troubleshooting Kubernetes in production.
  • Hands-on data-center, colocation, or equivalent experience with servers, virtualization, storage, and networking.
  • Strong Linux fundamentals and working knowledge of TCP/IP, DNS, and TLS.
  • Experience with monitoring, logging, alerting, and incident response; clear communication and a calm, methodical approach during outages.
  • Scripting experience in Bash, Python, or Go; Git experience; and familiarity with CI/CD, GitOps, or IaC practices using any toolset.

Nice To Haves

  • RKE2/Rancher
  • AKS in a hybrid environment
  • Kyverno/OPA
  • 25/100GbE or Cisco Nexus
  • SAN/NVMe/object storage
  • stateful data platforms on Kubernetes
  • GPU/AI workloads
  • GitOps pipelines
  • Terraform or Ansible

Responsibilities

  • Operate production and non-production RKE2/Rancher Kubernetes clusters, including upgrades, node lifecycle, networking, ingress, DNS, certificates, and capacity; troubleshoot control-plane, scheduling, CoreDNS, and memory (OOM) issues, and improve observability, alerting, and runbooks.
  • Run Kubernetes on physical HPE servers and virtual machines in company data centers, including hardware, firmware, RAID, and out-of-band management; partner with infrastructure teams on Cisco networking, SAN/NVMe/object storage, failure domains, and capacity planning.
  • Support AKS, Azure Container Registry (ACR), and connectivity between data centers and Azure. Implement, test, and document disaster-recovery plans, and verify that backups can be restored.
  • Join the on-call rotation; investigate and respond to incidents, escalating to the Staff SRE when appropriate. Reduce recurring toil through automation, infrastructure as code (IaC), safer CI/CD, SLIs/SLOs, and attention to single points of failure.

Benefits

  • Take meaningful ownership of a platform being built for its next chapter.
  • You’ll help shape it from the ground up, influence foundational decisions, and see your work become the systems the company relies on—from physical servers to cloud.
  • If you want to build, improve, and own real infrastructure rather than inherit a ticket queue, this is your opportunity.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service