Founding Platform Engineer, Agent Runtime

ZingageNew York, NY
Onsite

About The Position

Home care is a $130B industry that runs on the telephone, and we are the AI that answers it. Zingage's agents coordinate scheduling, on-call, and intake for 120+ home care agencies - 300K+ patient visits a month ride on our decisions, across every major EMR in the industry. We recently closed our Series A, backed by Bessemer, Bertelsmann Investments, and Yosemite (Reed Jobs), and the team is deliberately small and dense - operators and engineers from Ramp, Uber, Tennr, Datadog, and Verkada. Headquartered in New York. Our agents act in the physical world: they fill shifts, move visits, and write to the system of record that governs care for frail patients. An agent is only as good as the runtime beneath it - and the runtime for production agents in a regulated industry does not exist yet. Building it means solving four problems your best colleagues would agree are open: Distributed systems where the component is the source of nondeterminism. Forty years of systems design assumes deterministic components in an unreliable world. An agent runtime inverts the assumption: the actor itself is stochastic. What does idempotency mean for an agent action? What is a transaction when one participant is an LLM and the other is a twenty-year-old EMR with no locking semantics? Freshness as a per-decision SLO. "How stale is too stale to act on?" has a different answer for a 2AM call-off than for next week's schedule change. We need staleness budgets tied to real-world risk, enforced against four EMRs with no webhooks, brutal rate limits, undocumented semantics, and silent failure modes - CAP theorem where the partition is permanent and political. These upstreams were never designed to be built on. That is not the obstacle; it is why this layer, once built, is the moat. Replay of a stochastic actor. You cannot replay the model, so you must replay the world: capture every agent trace - inputs, tool calls, upstream state - completely enough that a new model can be interviewed against two years of production reality overnight. Time-travel debugging for agents. The trace corpus this produces is the most valuable asset the company will ever own, because it is how we adopt every new model first while competitors spend a quarter stabilizing. Safety as unrepresentability. A production outage taught us the principle the hard way: no prompt can make the model emit a field the schema forbids. Your job is to generalize it - a capability and governance system where unsafe actions against a patient's record are not discouraged or detected but unrepresentable at the boundary between a stochastic planner and a real-world effector. Zero unauthorized writes, audit-verified, forever, in a HIPAA-regulated industry where the audit trail must satisfy payers and regulators, not just engineers.

Requirements

  • Technical judgment
  • Ownership mindset
  • Role clarity
  • Experience with distributed systems.
  • Understanding of nondeterminism in system components.
  • Experience with LLMs and their integration into systems.
  • Knowledge of idempotency and transactional concepts in complex systems.
  • Experience with EMR systems or similar complex, legacy systems.
  • Understanding of CAP theorem and its implications.
  • Experience with data freshness SLOs and real-world risk assessment.
  • Ability to design and reason through system trade-offs.
  • Experience with HIPAA-regulated industries is a plus.
  • Experience with audit trails and regulatory compliance is a plus.

Nice To Haves

  • Experience with systems where components are sources of nondeterminism.
  • Experience with systems requiring per-decision freshness SLOs.
  • Experience with replaying stochastic actors.
  • Experience with safety mechanisms in AI systems.
  • Experience with EMRs that have no webhooks, rate limits, undocumented semantics, and silent failure modes.
  • Experience with building systems that must satisfy payers and regulators.

Responsibilities

  • Build the first version of each plane (Data, Action, Governance, Actuation).
  • Hire the team that will own these planes.
  • Generalize safety as unrepresentability for the agent runtime.
  • Ensure zero unauthorized writes, audit-verified, forever.
  • Develop a system for replaying stochastic actors by capturing agent traces.
  • Implement per-decision freshness SLOs for data.
  • Build a data plane that polls and streams EMR data across four vendors.
  • Create an action plane where every agent action is traced, auditable, and replayable.
  • Develop a governance plane for protected-field enforcement and capability-scoped writes.
  • Build an actuation plane that acts as a universal adapter for EMRs without APIs.

Benefits

  • Competitive base salary
  • Meaningful equity
  • Equipment stipend
  • Luxury gym membership in NYC
  • Daily lunch
  • Dinner when work runs past dark
  • Time off as needed
  • Happy hours
  • Poker nights
  • Builder events
  • Snacks in office
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service