Staff/Senior Distributed Systems Engineer

Helios Intelligence PlatformsNew York City, NY
Onsite

About The Position

Helios is building a new kind of company to solve America’s hardest problems, starting with the government interaction layer. Government shapes every consequential market, but the infrastructure connecting public institutions and private organizations remains fragmented, manual, and difficult to navigate. Helios is rebuilding that layer. Our core platform, Proxi, gives organizations the intelligence they need to understand what government is doing, why it matters, and what to do next. From that foundation, we design and deploy secure, mission-specific systems for government agencies, enterprises, and institutions operating in complex and highly regulated environments. We bring together frontier AI, deep public-sector expertise, and forward-deployed execution. Our team includes leaders and builders from the White House, U.S. Department of State, Datadog, and Microsoft. We are backed by leading institutional investors and trusted by organizations working on high-stakes problems across government and industry.

Requirements

  • Familiarity with network crawling, compute grids, memory-heavy document processing, OCR, inference, and latency-sensitive jobs.
  • Experience with queue-, log-, workflow-, and actor-based distributed execution systems, including Kafka, Redpanda, Pulsar, SQS, Pub/Sub, RabbitMQ, Temporal, and equivalent technologies.
  • Experience establishing at-least-once delivery with effectively once-only business outcomes through transactional outbox and inbox patterns, sagas, reconciliation, dead-letter handling, replay and historical backfill procedures.
  • Experience with CPU-, memory-, disk-, network-, and GPU-aware scheduling, including workload classification, priority allocation, starvation prevention, tenant and dependency concurrency limits, placement constraints, resource quotas, and noisy-neighbor isolation.
  • Experience with PostgreSQL, object storage, Redis or equivalent caches, search indices (Elasticsearch or Typesense), vector stores, change-data-capture, and graph-storage systems.
  • Experience with application of distributed coordination and concurrency controls, including leader election, locks, leases, fencing tokens, optimistic concurrency, conflict resolution, and partition recovery.
  • Experience with provisioning and operation of AWS, GCP, or Azure infrastructure through terraform or similar IaC, vulnerability scanning, CI/CD and deployment methods.
  • Experience with definition and administration of SLO/Is, error budgets, release controls, metrics, OpenTelemetry, Datadog, fault injection and load/failure testing.
  • Experience with evaluation of system performance and cost via e2e latency decomposition, resource and dependency profiling.
  • Experience with enforcement of platform security and tenant isolation policies.
  • Ability to support the real-time access and availability of our entire data corpus as well as supporting the continued construction of the Helios Rapid Ontology System (H.R.O.S.), our long horizon memory data plane.

Nice To Haves

  • Experience with frontier AI, deep public-sector expertise, and forward-deployed execution.
  • Experience working with leaders and builders from the White House, U.S. Department of State, Datadog, and Microsoft.
  • Experience working in a fast-moving startup with ambitious goals.
  • Flexibility during critical deployment periods, customer incidents, product launches, and other company-critical work.

Responsibilities

  • Expand and operate the distributed execution, storage, scheduling, and reliability primitives that support data ingestion, document processing, search indexing, model inference, agents, and continuously running research workflows.
  • Orchestrate critical infrastructure to meet mission-critical performance and reliability guarantees.
  • Manage the deployment of our GovCloud and Air-Gapped resources for sensitive environments.
  • Manage the orchestration of long-running agent research tasks including scheduling, lease management and retention.
  • Queue optimization and cross-cloud information pipeline scalability.
  • Build out dedicated resource-aware autoscaling architecture for fast search and document processing resources.
  • Manage CI/CD and compliance operations across the entire Helios platform.
  • Support global forward embedded customer infrastructure efforts.

Benefits

  • Unusual ownership
  • Direct access to consequential institutions
  • Opportunity to build systems that affect how major decisions are made
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service