Positron is seeking an Engineering Manager to lead Serving/API. This team owns the software between a customer's API request and the tokens Positron accelerators generate: what the model sees, and how its output becomes a correct, well-formed response. You will lead a strong team and grow it as Positron adds models, modalities, and serving capabilities. The team owns: The OpenAI-compatible serving layer. HTTP endpoints, request validation, SSE streaming, usage accounting and API-spec fidelity. Tokenization and chat templates. HuggingFace tokenizers, chat-template rendering, and model-specific conversation formats such as OpenAI Harmony for GPT-OSS. Tool calling and structured output. Tool-schema handling, tool-call parsing and serialization, and grammar-constrained decoding (llguidance) for JSON and function calls. Reasoning and budgets. Reasoning-channel parsing, reasoning-effort controls, token limits and per-request budgets. Speculative decoding. The request-side plumbing for draft models and draft trees, acceptance metrics, and the policies that decide when speculation pays off. New modalities and frameworks. Vision-language model (VLM) input handling and integration paths with SGLang-style serving. This is a technical leadership role with real product accountability. You will set direction, build the team, create clear ownership, and improve the systems behind API fidelity, tool calling, reasoning, speculative decoding, and new-model readiness. You will work closely with Production Platform and Orchestration, Compiler/Executor, hardware, and customer-facing partners.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Manager
Education Level
No Education Listed