Member of Technical Staff, Forward Deployed

Chakra Labs•New York City, NY
•Hybrid

About The Position

Chakra Labs' mission is to encode human taste into intelligence. We build high-fidelity environments, evals, and datasets for frontier AI research, working with several of the top labs. Our work sits at the frontier of post-training, agent environments, data quality, and research infrastructure. We care about building systems that make models better in ways that are measurable, useful, and hard to fake. The hardest problems at the frontier. A new environment modality, an eval targeting a failure mode nobody's measured, a dataset that doesn't exist yet. You take problems like these from a researcher's hunch to a shipped deliverable, working at the edge of what agents can currently do. Environments, evals, and datasets. One project is a high-fidelity environment, the next is a task distribution with grading logic, the next is a dataset built to a demanding spec. The bar is frontier-lab quality and the pace is relentless - you're writing whatever the deliverable needs: environment code, task specs, scoring harnesses. Pulling the frontier into the platform. The best one-offs don't stay one-offs. You'd recognize when a custom build proves out a capability worth generalizing, and help fold it into the core product - so it compounds instead of sitting on a shelf.

Requirements

  • Strong generalist engineer comfortable across backend services, data pipelines, and frontend to ship a usable interface.
  • Proficiency in TypeScript/Python or similar.
  • Ability to own a whole deliverable rather than just a layer of one.
  • Ability to turn a loosely-defined research question into a concrete environment, task set, eval, or dataset.
  • Ability to validate that a deliverable measures what was meant, not just what was easy to build.
  • Clear communication with researchers and technical customers.
  • Customer instincts to push back when they're asking for the wrong thing.
  • Genuine interest in how agents fail and how to measure it.
  • Ideally at least 3 years shipping production software, but less works if the above sounds like you.

Nice To Haves

  • Experience with ML research.
  • Experience with task design.
  • Experience with LLM-judged scoring.
  • Experience with reward hacking detection.

Responsibilities

  • Turn loosely-defined research questions into concrete environments, task sets, evals, or datasets.
  • Validate that deliverables measure what was intended, not just what was easy to build.
  • Communicate clearly with researchers and technical customers.
  • Push back when customers are asking for the wrong thing.
  • Differentiate between customer requests and customer needs.
  • Write environment code, task specs, and scoring harnesses.
  • Generalize custom builds into the core product.

Benefits

  • Work with the latest and greatest technologies across the data, AI, and infrastructure stack.
  • Work with a small team of experienced professionals from top tech companies.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service