Applied AI Engineer

Velixo
Remote

About The Position

Velixo is seeking an Applied AI Engineer to own the performance and cost of Velixo Intelligence, our MCP-based gateway that enables finance teams to query their ERP in natural language and perform governed writebacks using AI models like Claude, ChatGPT, and Copilot. This is a measurement-first engineering role focused on replacing anecdotal quality assessments with evidence-based evaluations and driving down cost per action without degrading quality. The role involves building evaluation infrastructure, instrumenting actions for cost tracking, reducing costs through various optimization techniques, and collaborating with the finance team to provide unit cost data for pricing and margin analysis. The successful candidate will have substantial experience building on LLM APIs in production, experience with evaluation harnesses and tooling, strong analytical skills, and proficiency in Python or TypeScript. The role offers the opportunity to tackle genuinely hard evaluation problems in a high-stakes environment and gain significant visibility into product and pricing decisions.

Requirements

  • Substantial experience building on LLM APIs in production.
  • Experience building an evaluation harness and understanding its limitations.
  • Hands-on with current tooling landscape: tracing and observability platforms (e.g., Langfuse), eval frameworks (e.g., promptfoo).
  • Comfortable with token accounting, context window management, and caching strategies.
  • Strong analytical instincts and fluency with data.
  • Ability to write clearly about tradeoffs for a non-technical audience.
  • Python or TypeScript proficiency.

Nice To Haves

  • Experience with MCP or other tool-calling frameworks.
  • Background in ERP, accounting, or financial systems.
  • Prior exposure to usage-based or credit-based pricing models.
  • Experience at a small company where you set your own priorities.

Responsibilities

  • Build and maintain regression suites for GIQL query generation, writeback correctness, and tool selection.
  • Define and measure what constitutes "good" performance for each action type.
  • Run structured evaluations before model upgrades or prompt changes, providing clear ship/no-ship recommendations.
  • Track quality over time to detect degradation before customers do.
  • Instrument every Gateway action for tokens, cache behavior, model, latency, and tenant.
  • Establish and maintain the cost-per-action baseline.
  • Reduce cost through prompt compression, prompt caching, context pruning, and routing work to smaller models.
  • Identify the expensive tail of usage (customers, query shapes, failure loops).
  • Track new model releases and evaluate them against internal workloads.
  • Maintain a current view of the price and performance frontier for AI models.
  • Recommend when to migrate models, wait, or run models in parallel.
  • Supply unit cost data to the finance team for credit pricing and plan allowances.
  • Model the margin impact of proposed pricing changes.
  • Provide the numbers that pricing and product decisions depend on.

Benefits

  • Equal opportunity employer.
  • Celebrates diversity.
  • Committed to creating an inclusive environment.
  • Fully remote position.
  • Priority to candidates located in Quebec.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service