Senior DevOps / Infrastructure Engineer

Category LabsNew York, NY

About The Position

Category Labs (formerly known as Monad Labs) is seeking a Senior DevOps / Infrastructure Engineer to operate the infrastructure behind Monad, a high-performance, EVM-compatible Layer 1 blockchain. This role involves managing a globally-distributed fleet of nodes, owning infrastructure-as-code and observability, and building AI-driven tooling for operations. The engineer will play a key role in designing workflows and guardrails for autonomous AI agents operating infrastructure, and will also manage the infrastructure for the company's model workloads. The company has recently raised $225M in series A funding and its public mainnet is now live.

Requirements

  • 5+ years in DevOps, SRE, or Infrastructure Engineering, operating production systems at scale.
  • Strong Linux, systemd, networking, and shell fundamentals, and comfortable debugging live systems over SSH.
  • Deep, hands-on infrastructure-as-code experience with Ansible and Terraform.
  • Experience with observability stacks (Prometheus, Grafana, Loki, or equivalents).
  • Hands-on fluency with AI-assisted engineering: use coding agents and LLM tooling in daily workflow and have judgment on where it helps and where it's risky.
  • Experience designing automation with safe guardrails.
  • Calm, methodical incident response.
  • Programming and scripting experience (e.g., Python, bash).

Nice To Haves

  • Experience with Kubernetes and GitOps (Flux or Argo).
  • Experience building AI agent tooling, MCP servers, or agent orchestration frameworks.
  • Experience serving inference, either locally or as a service.
  • Previous experience with blockchain clients or node operations.
  • A Bachelor of Science in Computer Science, Engineering, or a related field.

Responsibilities

  • Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response.
  • Own infrastructure-as-code: Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (GitOps) for platform services.
  • Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives.
  • Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes).
  • Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight.
  • Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse.
  • Harden nodes and services, manage secrets, and continuously drive down manual toil.

Benefits

  • Competitive salary and equity package.
  • Private health insurance options.
  • Flexible paid time off.
  • Monthly wellness reimbursement.
  • Paid parental leave.
  • World-class benefits package with 100% paid medical, dental, and vision insurance including 75% coverage for dependents and HSA + FSA options (US employees).
  • 401(k) with company match (US employees).
  • Lunch and dinner stipend (in-office NYC).
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service