Platform Engineer

Perry WeatherDallas, TX
Onsite

About The Position

At Perry Weather, we build weather safety technology that organizations depend on when conditions change. Thousands of users trust our fast-growing platform across school districts, cities, universities, golf courses, professional sports, construction, and manufacturing. Our software and connected hardware let organizations make rapid decisions when weather disrupts operations, protecting people and keeping activities running. We're looking for a Platform Engineer to help build and run the infrastructure the rest of engineering depends on. You'll partner closely with our Lead Platform Engineer, designing and delivering the work together, and own real surface area from the start: infrastructure as code, our Kubernetes environments, delivery pipelines, observability, and the guardrails that keep all of it safe to change. Engineering is your customer. Success in this role looks like fast service setup, uneventful deploys, and engineers who can diagnose a failing workload without asking for help. Engineering here also works with AI coding agents daily, which adds to the platform's job: checks that catch generated infrastructure errors before they apply, scoped credentials and audit trails for automated actors, and cost reporting that accounts for agent usage. You'll use agents in your own work and build these controls for everyone else.

Requirements

  • 4+ years in platform, infrastructure, DevOps, or backend engineering, with real ownership of production infrastructure
  • Terraform you've written and maintained. You've reviewed a plan and caught something you didn't want applied, and you know how state drift happens
  • Kubernetes in production. You can take a failing workload, trace it through configuration, networking, and resource limits, and explain what went wrong to the engineer who owns the service
  • CI/CD you've built and debugged, ideally GitHub Actions. You've fixed a pipeline engineers had stopped trusting, and you can separate a failing test from failing infrastructure
  • Working depth in a major cloud, including managed databases, networking, identity, and secrets. We're primarily on Azure, and AWS or GCP experience transfers
  • Scripting in Python, Go, or Bash good enough to automate a manual process end to end and leave it maintainable for someone else
  • Observability work you've done yourself: instrumentation you added, a dashboard other people used, and an alert you tuned because it fired too often
  • On-call experience for production systems, including at least one incident you drove to resolution and wrote up afterward
  • Product instincts about internal tooling. You've built something for other engineers, watched them use it, and changed it based on what you saw
  • Fluency with AI-native engineering tools and agentic workflows (e.g., terminal-native coding agents, LLM-assisted code refactoring and generation) to multiply technical output and speed up development cycles
  • Least-privilege credential design for automated systems: CI service accounts, scoped tokens, short-lived credentials, audit trails. Agents are the newest consumers of that work, and the principles carry over
  • A view on how automated changes should be reviewed. More code and configuration now arrive generated, which puts weight on the checks that run before a merge or an apply, and you should have opinions about what those checks need to cover

Nice To Haves

  • Experience with policy-as-code
  • Cloud cost management or FinOps practices
  • API gateway operations
  • DNS and CDN configuration
  • Helm chart authoring
  • Managed time-series or high-volume databases
  • Load and performance testing
  • Running AI or ML workloads on shared infrastructure
  • Exposure to fleets of connected hardware

Responsibilities

  • Write and maintain the Terraform that defines our cloud footprint, review changes for blast radius, and keep our modules usable by engineers outside the platform team.
  • Run our Kubernetes clusters and the Helm-deployed services on them, including resource tuning, scaling behavior, and the failure modes that only show up under load.
  • Build and maintain CI/CD in GitHub Actions so that builds, tests, and deploys are fast and dependable enough that engineers rely on them.
  • Improve the instrumentation, dashboards, and alerting that tell us how the platform is behaving, and make alerts specific enough that people act on them.
  • Track where our cloud spend goes, find the waste, and add controls that catch expensive changes before they ship.
  • Implement secrets management, least-privilege access, dependency and image scanning, and policy checks that catch mistakes during review.
  • Give AI coding agents and automation scoped identities and audit trails, add checks that catch generated infrastructure mistakes before they apply, and track their usage costs.
  • Share the on-call rotation, respond when the platform degrades, and make the change that prevents a repeat.

Benefits

  • Competitive health insurance plans
  • 401(k) with employer matching
  • Suite of voluntary benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service