Senior Site Reliability and Infrastructure Engineer

Treeswift IncNew York, NY
$160,000 - $220,000Hybrid

About The Position

Treeswift empowers energy companies to modernize their field work to meet the growth and challenges ahead by building physical AI for field workers. Their technology, described as an 'ironman suit' for engineers, linemen, and vegetation crews, aims to increase productivity tenfold. The platform integrates cutting-edge hardware, sensors (LiDAR, camera), AI, and software to revolutionize work in challenging environments. Since its first pilot in June 2024, Treeswift has rapidly expanded, now collaborating with three of the five largest utilities in the US. Their technology has proven effective in reducing wildfire risk, regulatory and outage risks from vegetation, preventing delays and cost overruns in new construction, and accelerating storm recovery. The company is assembling a team of mission-driven experts with deep industry experience in robotics and enterprise software development, backed by leading investors. Headquartered in midtown Manhattan with offices in San Francisco and Philadelphia, Treeswift is experiencing accelerating growth and seeks ambitious individuals to contribute to the future of field work.

Requirements

  • Experienced software engineer with the last 7-10 years requiring significant time on observability, systems/infrastructure engineering, SRE, or DevOps (ideally in a cloud environment).
  • Ability to reason about architecture end-to-end and articulate thoughts with product impact in mind (data movement, execution, failure handling, and operational visibility).
  • Hands-on experience with infrastructure-as-code (Terraform and similar) and using it to deliver reliable environments.
  • Experience with container orchestration and debugging in practice (Kubernetes and/or ECS/container-based deployments).
  • Strong Linux debugging skills and demonstrated ability to investigate production issues with logs/metrics and clear hypotheses.
  • Empathy and communication: ability to collaborate effectively with engineers across teams (especially the data platform team) and explain tradeoffs clearly.

Nice To Haves

  • Experience working in early-stage or fast-moving environments where ownership and processes evolve quickly.
  • Experience with Apache Airflow and/or Astronomer.
  • Experience with AWS, although other cloud providers are fine. (DuploCloud experience is also helpful.)
  • Experience with geospatial/imagery/lidar/point-cloud style domains.
  • ML Ops skills (model deployment/inference reliability, packaging, CI/CD for model artifacts, and operational observability for inference pipelines).

Responsibilities

  • Partner with the data platform and engineering teams to understand how changes propagate across pipeline execution (Astronomer-hosted Airflow DAGs), containerized workers (Kubernetes), and AWS services (S3, SQS, Lambda, Step Functions, ECS).
  • Design and implement reliability and observability for high-volume pipeline operations, including actionable monitoring/alerting for DAG/task failures and reruns, visibility into operational workflows like flight orchestration (including DLQ/failed-message alerting and notification pathways), and dashboards and SLO/SLI definitions focused on correctness, throughput, and pipeline health.
  • Own CI/CD guardrails for production changes: build/deploy validation and safe rollout mechanics for Astronomer deployments (image builds pushed to ECR, and Airflow configuration updates via Astronomer CLI variable updates).
  • Make machine learning inference operations more reliable and observable: instrument inference runs executed inside pipeline runners (model checkpoint resolution, S3 sync behavior, thresholds and fallback behavior, and output correctness), and add operational visibility for inference outcomes (e.g., unknown classification rates, fallback usage, and failure modes).
  • Create operational tooling and continuously improve systems (‘leave it better than you found it’), including runbooks, incident learnings, and engineering standards for debugging at scale, and automate away toil in deployment and operations workflows as we learn what hurts most.
  • On-call / incident response: help lead reliability improvements and operational readiness so the team has faster diagnosis, better alerts, and safer releases when issues do occur.

Benefits

  • Total compensation for this position is determined by skills, qualifications, relevant work experience, location, and other factors.
  • This salary estimate excludes the value of any potential bonuses; the value of any benefits offered; and the potential future value of any long-term incentives.
  • Treeswift is proud to be an equal opportunity employer. We provide employment opportunities without regard to age, race, color, ancestry, national origin, religion, disability, sex, gender identity or expression, sexual orientation, veteran status, or any other protected status in accordance with applicable law.
  • If you require any accommodations during the recruitment process, whether it be alternate forms of material, accessible meeting rooms, etc., please let us know and we will work with you to meet your needs.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service