About The Position

Own, operate, and improve a major vertical of CaseGuild’s distributed infrastructure. CaseGuild runs large-scale services that execute millions of queries, process hundreds of millions of tokens every minute, and ingest, transform, store, and retrieve substantial volumes of structured and unstructured data. The difficult part is not simply adding capacity. Workloads vary enormously: one customer’s matter may be 1,000 times larger or more demanding than another’s while both run on shared infrastructure. You will design systems that remain fair, isolated, observable, and predictable under contention. This is an end-to-end ownership role. There is no separate platform, SRE, infrastructure, or database team responsible for finishing the work. When you design a system, you will also own its infrastructure definitions, deployment configuration, production promotion, observability, operational behavior, and incident response. This is not primarily an architecture or advisory position. You will write production code, investigate performance and reliability problems, operate what you build, and establish technical patterns that other engineers can use. Success means One customer’s workload cannot degrade another’s. Large jobs are isolated, admission is fair, and tail latency remains predictable under contention. Critical services have clear ownership, strong observability, understood failure modes, and reliable recovery paths. The system handles extreme variance in matter size, query patterns, ingestion volume, and processing demand without requiring manual intervention. Bottlenecks across ingestion, storage, retrieval, orchestration, and AI-processing pipelines are identified and removed. Infrastructure, application code, deployment configuration, and production operation are treated as one engineering responsibility rather than separate functions. The engineering team makes better architectural decisions because you contribute both technical leadership and working implementations.

Requirements

  • Built and operated high-volume, low-latency services on shared infrastructure.
  • Experience with distributed systems.
  • Experience with workload isolation and multi-tenancy.
  • Experience with admission control, queuing, scheduling, and backpressure.
  • Experience with tail-latency management.
  • Experience with reliability and failure recovery.
  • Experience with large relational and NoSQL data stores.
  • Experience with production infrastructure and deployment automation.
  • Experience rebuilding a system that most people are afraid to touch.
  • AI-native enough that if you can dream it, you can drive AI to build it.

Responsibilities

  • Own, operate, and improve a major vertical of CaseGuild’s distributed infrastructure.
  • Design systems that remain fair, isolated, observable, and predictable under contention.
  • Own infrastructure definitions, deployment configuration, production promotion, observability, operational behavior, and incident response for designed systems.
  • Write production code.
  • Investigate performance and reliability problems.
  • Operate what you build.
  • Establish technical patterns that other engineers can use.
  • Ensure one customer’s workload cannot degrade another’s.
  • Ensure large jobs are isolated, admission is fair, and tail latency remains predictable under contention.
  • Ensure critical services have clear ownership, strong observability, understood failure modes, and reliable recovery paths.
  • Handle extreme variance in matter size, query patterns, ingestion volume, and processing demand without requiring manual intervention.
  • Identify and remove bottlenecks across ingestion, storage, retrieval, orchestration, and AI-processing pipelines.
  • Treat infrastructure, application code, deployment configuration, and production operation as one engineering responsibility.
  • Contribute technical leadership and working implementations to enable better architectural decisions.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service