AI Platform Lead

Docusign•San Francisco, CA
•Hybrid

About The Position

We're building the foundation that lets every engineering team at Docusign build, ship, and operate AI-native and agentic systems safely and at scale. This isn't a traditional infrastructure or DevOps role: it's an architecture and orchestration role for someone who has already designed and run production LLM and agentic systems at scale. As the AI Platform Lead, you own the architecture of our AI platform — the infrastructure, model gateway and routing layer, the agent orchestration frameworks, and the inference and GPU serving infrastructure that other teams build on top of. You set the technical direction, write the reference implementations and standards, and partner directly with product and platform engineering teams to get them adopted. You are deeply hands-on: you prototype, you write production code, and you use AI coding assistants and agentic tooling as a core part of how you work — but your primary output is the architecture and platform other engineers build against, not a single service. The successful candidate has real, hands-on experience building or operating LLM-powered and agentic services in production — not just experimenting with them. You're self-directed, comfortable owning ambiguous technical direction, and able to explain complex AI-infrastructure tradeoffs to both engineers and non-technical stakeholders. This position is an individual contributor role reporting to the Sr. Director of Cloud Services.

Requirements

  • Basic Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent experience
  • 12+ years of experience in Cloud, Platform, or Infrastructure Engineering with a Bachelor's degree or 8+ years of experience with a Master's degree
  • Experience architecting and operating LLM-powered or agentic systems in production (not just prototypes or POCs)
  • Experience replacing manual, ticket-driven operations with automated or self-service systems — building the systems that act on what they observe, not just dashboards and monitors
  • Experience designing and operating an LLM gateway / model routing layer across multiple providers (e.g., OpenAI, Anthropic, Bedrock, Azure OpenAI)
  • Experience building and operating agent orchestration frameworks and agentic workflows (e.g., LangGraph, CrewAI, AutoGen, or a custom orchestration runtime) in production
  • Experience with GPU/inference-serving infrastructure and Kubernetes at scale (e.g., vLLM, KServe, Triton, model autoscaling, capacity planning)
  • Experience writing, reviewing, and testing production-quality code (Python and/or Go) — this is a build role, not a slide-deck role
  • Experience with cloud platforms (AWS, Azure, or GCP) and Infrastructure-as-Code (Terraform or Ansible)
  • Experience setting technical direction across teams and driving adoption of a shared platform, with strong written and verbal communication

Nice To Haves

  • Experience with RAG and vector database infrastructure (embedding pipelines, retrieval services) in production
  • Experience running LLMOps practices — model evaluation, prompt/version management, guardrails, and agent monitoring
  • Experience with containerization and microservices architecture (Docker, Kubernetes) beyond AI workloads
  • Understanding of network architecture and security best practices (VPNs, firewalls, load balancing) in cloud environments
  • Experience presenting architecture and roadmap to senior technical leadership and driving org-wide adoption

Responsibilities

  • Own the end-to-end design, build, and integration of the AI/agentic platform that compliments human-operated execution with systems that route, remediate, and provision themselves — model gateway, agent orchestration, inference serving — and the standards other teams build against.
  • Architect, build, and operate a multi-provider LLM gateway and routing layer (e.g., LiteLLM, Portkey, Bedrock, or equivalent) so requests reach the right model automatically — secure, cost-aware, and reliable — without a person routing them by hand
  • Design and build the orchestration layer and frameworks (e.g., LangGraph, CrewAI, or a custom runtime) that let agents carry out work autonomously, and define the rules — capability schemas, guardrails, policy enforcement — those agents operate under
  • Architect and operate scalable model-serving and GPU infrastructure (e.g., vLLM, KServe/Triton, Kubernetes-based inference serving, autoscaling, capacity planning) that balances latency, throughput, and cost
  • Replace manual provisioning and ticket-driven execution with automated, self-service systems — teams get what they need and common failures resolve themselves, without a person in the loop
  • Build golden paths for fine-tuning, model/agent registries, evaluation, and safe deployment — including CI/CD for models, prompts, and agents with canary, A/B, and shadow rollouts
  • Define and evangelize architecture patterns, security guardrails, and reusable accelerators/reference architectures that reduce time-to-value for teams building on the platform
  • Partner closely with product and platform engineering teams to drive adoption of the AI platform, unblock their AI/agentic use cases, and translate platform capabilities into their roadmaps
  • Implement security controls and Zero Trust practices — including SBOMs, image signing, and policy-as-code — partnering with InfoSec to maintain a strong security posture across model access, data handling, and agent permissions
  • Establish observability for AI infrastructure — token spend, model latency, GPU utilization, agent behavior — and continuously verify that what you've automated performs as intended

Benefits

  • Bonus: Sales personnel are eligible for variable incentive pay dependent on their achievement of pre-established sales goals. Non-Sales roles are eligible for a company bonus plan, which is calculated as a percentage of eligible wages and dependent on company performance.
  • Stock: This role is eligible to receive Restricted Stock Units (RSUs).
  • Paid Time Off: earned time off, as well as paid company holidays based on region
  • Paid Parental Leave: take up to six months off with your child after birth, adoption or foster care placement
  • Full Health Benefits Plans: options for 100% employer paid and minimum employee contribution health plans from day one of employment
  • Retirement Plans: select retirement and pension programs with potential for employer contributions
  • Learning and Development: options for coaching, online courses and education reimbursements
  • Compassionate Care Leave: paid time off following the loss of a loved one and other life-changing events
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service