Senior Devops Engineer

Tomorrow.ioWashington, DC
$160,000 - $180,000Hybrid

About The Position

Tomorrow.io's engineering department is focused on building life changing software and products at scale, from infrastructure that handles massive amounts of data to outstanding customer-centric user experiences in B2B, B2C and B2D products that change billions of lives worldwide. We're looking for a Senior DevOps Engineer to power the reliability, security, and efficiency of the world's most impactful weather platform. Your work will be guided by four core pillars- Security, Cost, SLOs, and Developer Experience- with AI acting as a force multiplier across each. You'll build self-service platforms that give developers and weather scientists true independence, weave AI into how we operate, and work side-by-side with R&D to push performance and scale further. Our environment spans two worlds: cloud-native product services on Kubernetes, and scientific computing on HPC clusters- and you'll help both thrive. You'll evolve our cloud infrastructure to match the pace of the business, hold the line on cost, and stay close to production through on-call. The people who thrive here bring a product mindset, take ownership without waiting to be asked, and leave the people and systems around them better than they found them.

Requirements

  • At least 6 years of experience as a Platform/DevOps/SRE Engineer in a containerized cloud environment experienced with AWS, GCP, or Azure and IaC, such as Terraform or Crossplane
  • Experience in fast-growing, cloud-native startup or scale-up environments
  • Strong sense of ownership and accountability for service reliability
  • Daily, hands-on use of AI coding agents; experience building agentic workflows is a plus
  • 10X mindset - always looking for the fastest, smartest path to a high-quality result
  • Daily use of AI coding agents (Claude Code, Copilot, etc.)- must; building agentic DevOps workflows- a plus
  • Experience with CI/CD tools and deployment methodologies in Kubernetes
  • Experience implementing and customizing monitoring systems (Datadog, Prometheus, Grafana, ELK Stack)
  • Experience working in an agile environment with high-velocity teams
  • Proficiency with scripting languages like Python, Node.js, and Go
  • Adaptable problem-solving mindset - thriving in changing environments and requirements
  • Excellent written and verbal communication skills, with the ability to collaborate effectively across distributed teams, time zones, and multiple R&D stakeholders

Nice To Haves

  • Experience with HPC / scientific computing- e.g., Slurm, AWS ParallelCluster, Azure CycleCloud
  • Familiarity with parallel filesystems- e.g., Lustre, NFS

Responsibilities

  • Develop and adopt AI-powered tools to make Development and Operations processes more efficient
  • Collaborate with weather scientists, engineers, and Spacecraft Mission Operations Engineers to optimize service performance, reliability, scale, security, and cost
  • Evolve and maintain adaptive cloud infrastructure to support our business strategy and enable smooth growth at scale
  • Build self-service platforms for scientists and developers to work independently
  • Support scientific computing workloads on HPC clusters (SLURM) alongside our cloud-native Kubernetes platforms
  • Introduce and integrate MLOps practices for GPU-based model deployment on Kubernetes
  • Maintain Production availability by participating in DevOps on-call shifts

Benefits

  • Comprehensive health benefits
  • Unlimited paid time off
  • Other benefits included
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service