Research Infrastructure - Member of Technical Staff

SimileSan Francisco, CA
$200,000 - $400,000Remote

About The Position

Simile is seeking a Member of Technical Staff for their Research Infrastructure team. This role involves building and owning the platform that researchers use for training, evaluating, and deploying models. The platform supports the full model lifecycle, from data exploration to production services serving millions of agent calls. The ideal candidate will be energized by optimizing performance, reducing costs, and ensuring the reliability of these critical systems. This role requires a blend of deep systems and ML proficiency, experience with production ML platforms, and the ability to debug and optimize distributed systems. The position is for someone who can own problems end-to-end, including the deployment phase, and who is driven by the leverage their work provides across the entire organization.

Requirements

  • High proficiency in Python and hands-on experience with modern ML frameworks (e.g., PyTorch, JAX).
  • Ability to refactor complex codebases for performance and architectural integrity.
  • Understanding of modern ML architectures to optimize them, with an intuition for time and memory usage.
  • Comfort with NVIDIA GPUs and the surrounding stack (NCCL, InfiniBand and NVLink topology, CUDA, Triton) or ability to learn quickly.
  • Experience building production ML platforms or MLOps systems (training orchestration, experiment tooling, model serving, LLM application platforms).
  • Experience architecting, observing, and debugging production distributed systems, especially performance-critical ones.
  • Understanding of the training and fine-tuning lifecycle and ability to architect continuous ingestion, monitoring, and high-availability serving.
  • Research- and data-literate: ability to navigate the ML research frontier, reproduce complex papers, tackle messy data, and write with rigor.
  • Self-directed and pragmatic: ability to identify important problems, acquire necessary knowledge, and balance ideal solutions with practical adjustments.
  • Humble attitude, eagerness to help colleagues, and desire for team success.
  • Strong quantitative foundation, typically a degree in Computer Science, Mathematics, Statistics, or a related field, or demonstrated equivalent ability.

Nice To Haves

  • Experience optimizing inference for multi-agent or agentic environments where requests are interdependent.
  • Experience with distributed training at scale.
  • Experience building and shipping production AI agents.
  • Familiarity with LLM serving frameworks (vLLM, SGLang, TensorRT-LLM) and inference-time optimization.
  • Interdisciplinary background in social science modeling or behavioral economics.

Responsibilities

  • Build the ML platform our researchers live in, designing and operating services, libraries, and tooling for the full lifecycle: data exploration, feature generation, experiment tracking, training orchestration, evaluation, and deployment.
  • Increase experiment velocity and streamline the researcher’s path from ideation to a fully validated, production-ready model.
  • Own throughput end to end for training and data pipelines, focusing on model FLOPs utilization, tokenization cost, and ingestion paths.
  • Profile and fix bottlenecks in time and GPU memory usage, including implementing observability for future issues.
  • Own the inference path for simulations, focusing on batching, scheduling, KV cache reuse, quantization, and unique request patterns for population-scale runs.
  • Ensure cost per simulation and latency per agent are treated as research constraints.
  • Lead the redesign of the data architecture to handle complexity and volume of simulation models, defining schemas and ingestion logic.
  • Scale training pipelines to meet the demands of society-scale modeling.
  • Own the GPU cluster, ensuring a multi-node fleet is healthy and saturated, managing node bring-up, NCCL and RDMA configuration, scheduling, storage lifecycle, checkpoint capacity, autoscaling, and alerting.
  • Keep the entire stack portable across multiple compute providers.
  • Build evaluation tooling with rigorous statistical frameworks to prove simulation fidelity and ensure fast execution.
  • Reproduce, critique, and improve upon academic work in simulation, training, and inference optimization.
  • Translate theoretical breakthroughs into production improvements with academic-level rigor.

Benefits

  • Competitive compensation packages
  • Base salary
  • Equity
  • Comprehensive benefits
  • Comprehensive medical, dental, and vision coverage
  • Flexible time off policies
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service