CloudOps Engineer - Software

Gallagher GroupKansas City, KS
Onsite

About The Position

The CloudOps Engineer - Software role sits within the eShepherd R&D team, reporting to the Regional Lead for North America, with day-to-day collaboration with the Software Team Lead for Cloud and Apps and engineers in Melbourne and Hamilton. eShepherd is a virtual fencing solution for cattle, utilizing solar-powered neckbands, a phone app, and a platform that enables farmers to manage livestock remotely. The company has grown from a startup to a scale-up within Gallagher, a long-standing farm fencing company. This role is crucial for evolving the platform's architecture to handle significant growth in device count, regional expansion, and new product development. The position involves investigating performance under load, identifying potential breaking points, and recommending solutions to the R&D team. Key areas of focus include query performance, data modeling at scale, ingest paths, caching, connection handling, storage strategy, regional topology, and failover. The role also encompasses improving monitoring, incident response, and capacity planning to ensure smooth growth. As part of the DevOps function, the engineer will extend it into the North American region, managing pipelines, infrastructure as code, environments, automation, and deployment tooling in collaboration with the global team. This includes setting up regional environments, deploying releases, and ensuring automation supports multi-region operations. The role contributes to the platform's next major extension, involving design work for new products and scaling. The CloudOps Engineer will serve as the R&D presence in North America, acting as the on-the-ground engineer during Melbourne's off-hours for incident response and providing timely answers to North American customer success and operations teams. The job emphasizes a cycle of design, test, deploy, and iterate, prioritizing measurement and weekly improvements over lengthy development cycles. The ideal candidate is results-oriented, data-driven, comfortable working independently across time zones, and possesses strong written communication skills. An understanding of the physical constraints of devices on farms (e.g., battery life, cellular coverage) is important for designing robust systems. The role requires commercial experience running production cloud infrastructure on AWS at scale, deep database competence with MySQL, a track record of improving system performance and reliability, experience with monitoring and incident response, and proficiency in infrastructure as code and CI/CD practices. Familiarity with IoT architectures and protocols is also beneficial.

Requirements

  • Commercial experience running production cloud infrastructure on AWS at meaningful scale, including compute, managed databases, and container orchestration (RDS, ECS, EKS).
  • Deep database competence with MySQL, covering schema design, query performance, indexing strategy, and tuning.
  • A track record of making systems faster and more reliable, with the ability to describe measurements, changes, and outcomes.
  • Proficiency in profiling, load testing, capacity planning, and cost per transaction analysis.
  • Real experience with monitoring, alerting, and incident response, including writing postmortems and implementing preventative changes.
  • Infrastructure as code as a default way of working.
  • Solid CI/CD and automation practice, covering build, test, and deployment workflows.
  • Sufficient development capability in Python or Node.js to write necessary tooling and services.
  • Strong communication skills to connect teams across different time zones.
  • Employment authorization in the U.S.

Nice To Haves

  • Familiarity with IoT architectures and protocols (e.g., MQTT, LTE-M, NB-IoT, LoRa).
  • Understanding of the impact of large fleets of constrained devices on cellular networks on a backend.
  • Experience with multi-region deployments.
  • Knowledge of data residency requirements.
  • Experience with high-volume time series or telemetry workloads.
  • Familiarity with DevSecOps practices.
  • Experience with Terraform.
  • Experience with Grafana or Prometheus.
  • Experience with embedded systems.

Responsibilities

  • Investigate platform performance under load, identify bottlenecks, and recommend solutions.
  • Analyze query performance, data modeling, ingest paths, caching, connection handling, storage strategy, regional topology, and failover.
  • Improve monitoring to catch problems before they become support tickets.
  • Enhance incident response to produce fixes rather than restarts.
  • Develop capacity plans to manage growth effectively.
  • Extend the DevOps function into the North American region.
  • Build and maintain CI/CD pipelines, infrastructure as code, environments, automation, and deployment tooling.
  • Manage regional environment setup and release deployments.
  • Contribute to the design of the next major platform extension and new product development.
  • Act as the R&D presence in North America, providing on-the-ground support and acting as the primary contact during off-hours for the Melbourne team.
  • Communicate findings and system behavior from the field back to the R&D team.
  • Implement and maintain systems using infrastructure as code principles.
  • Develop and maintain CI/CD and automation workflows.
  • Write tooling and services using Python or Node.js.
  • Collaborate with global engineering teams across significant time differences.

Benefits

  • Competitive salary and performance-based incentives
  • Health, dental and vision insurance
  • Life, accident, disability and critical illness coverage
  • 401(k) with employer contributions
  • Employee Assistance and wellness programs
  • Modern tools and technology to support your work
  • Opportunities for professional development and career growth
  • A collaborative engineering culture with global reach and startup energy
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service