Engineering Manager, Serving/API

Positron Corporation
•$225,000 - $350,000•Remote

About The Position

Positron is seeking an Engineering Manager to lead Serving/API. This team owns the software between a customer's API request and the tokens Positron accelerators generate: what the model sees, and how its output becomes a correct, well-formed response. You will lead a strong team and grow it as Positron adds models, modalities, and serving capabilities. The team owns: The OpenAI-compatible serving layer. HTTP endpoints, request validation, SSE streaming, usage accounting and API-spec fidelity. Tokenization and chat templates. HuggingFace tokenizers, chat-template rendering, and model-specific conversation formats such as OpenAI Harmony for GPT-OSS. Tool calling and structured output. Tool-schema handling, tool-call parsing and serialization, and grammar-constrained decoding (llguidance) for JSON and function calls. Reasoning and budgets. Reasoning-channel parsing, reasoning-effort controls, token limits and per-request budgets. Speculative decoding. The request-side plumbing for draft models and draft trees, acceptance metrics, and the policies that decide when speculation pays off. New modalities and frameworks. Vision-language model (VLM) input handling and integration paths with SGLang-style serving. This is a technical leadership role with real product accountability. You will set direction, build the team, create clear ownership, and improve the systems behind API fidelity, tool calling, reasoning, speculative decoding, and new-model readiness. You will work closely with Production Platform and Orchestration, Compiler/Executor, hardware, and customer-facing partners.

Requirements

  • Demonstrated success managing and growing engineering teams responsible for LLM inference, model serving, ML systems, or a closely related domain.
  • Systems depth in C++ and working proficiency in Python; comfort reviewing performance- and correctness-critical code.
  • A record of turning ambiguous product demands into a coherent roadmap, explicit ownership, and measurable engineering outcomes.
  • Ability to recruit, coach, and retain engineers across experience levels while maintaining a high technical bar.
  • Clear written and verbal communication, sound prioritization, and the ability to make tradeoffs visible to technical and executive stakeholders.
  • A hands-on leadership style: close enough to architecture and code to ask the right questions without becoming the team's bottleneck.

Nice To Haves

  • Contributing to vLLM, SGLang, llguidance, XGrammar, Hugging Face tokenizers, or similar projects.
  • Strong technical judgment across LLM serving: tokenization, chat templates, sampling, streaming APIs, and the OpenAI Chat Completions and Responses APIs.
  • Hands-on understanding of tool calling, function-calling formats, and structured output, and of how they fail in practice.
  • Familiarity with open-source serving stacks such as vLLM, SGLang, or TensorRT-LLM, and the judgment to know when to adopt rather than build.
  • Serving vision-language models: image preprocessing, vision encoders, and multimodal token handling.
  • Supporting reasoning models and their output formats, such as OpenAI Harmony.
  • Serving on GPU, FPGA, ASIC, or other accelerators, especially memory-bandwidth-bound inference.
  • Customer-facing API products where engineering teams own compatibility, escalations, and release readiness.

Responsibilities

  • Lead, coach, and grow a team of engineers spanning API serving, tokenization, structured generation, and model integration.
  • Establish a clear technical and organizational roadmap for the serving layer: API surface, chat templates, tool calling, reasoning, budgets, and new-model support.
  • Ensure API fidelity: OpenAI-compatible behavior, streaming, usage accounting, and error handling that customers and their tooling can rely on.
  • Lead reasoning-model support and request budgets, including reasoning formats, reasoning-effort controls, token limits, and per-request accounting.
  • Partner with Compiler/Executor on speculative decoding: draft-model integration, request-side plumbing, acceptance metrics, and policies for when speculation pays off.
  • Lead the plan for vision-language model (VLM) support and SGLang interoperability, deciding what to adopt, what to build, and how it fits Positron's engine.
  • Keep serving-layer overhead off the critical path for time-to-first-token and streaming latency.
  • Keep the API layer independent of hardware topology as Positron adds platforms that run inference across multiple hosts.
  • Work with Production Platform and Orchestration to define clear ownership of the shared host-level load-balancing layer, including where request policies and protections live.
  • Build release gates (API conformance tests, tool-call and reasoning evals, regression suites) so new models and features ship with confidence.
  • Collaborate with Production Platform and Orchestration to turn new serving capabilities into supportable production endpoints.
  • Translate customer and business priorities into sequenced engineering work while protecting the team from reactive, unstructured requests.
  • Hire thoughtfully, develop emerging leaders, and create ownership boundaries that remain effective as the organization scales.

Benefits

  • Fully company-paid medical, dental, and vision insurance for you and your dependents
  • Company-paid life and disability coverage, with voluntary options to add more
  • Supplemental hospital, critical illness, and accident coverage available
  • Unlimited paid time off, we encourage everyone to truly unplug and recharge
  • 13 paid company holidays
  • Remote-first culture with a company-provided computer and home office setup
  • Competitive salary and equity
  • 401(k) with company matching, eligible from day one
  • Visa Support
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service