Senior AI Platform & Agentic Infrastructure Engineer

OKXSan Jose, CA
$178,000 - $321,000Remote

About The Position

OKX’s Internal Audit function has an early but working AI-native capability: a multi-agent platform (Hive Mind), agentic workflows, data pipelines, and AI-enabled tools that let a very small team punch far above its weight. Your job is not to rebuild it at today’s maturity. Your job is to take it to a regulator-grade production standard and well beyond, and to drive AI Enablement across the department. You own the foundation, the agent runtime, and the harness: a resilient cloud platform, the agentic runtime and evaluation harnesses that make agents trustworthy in a regulated setting, and the governed data and AI infrastructure everything else depends on. We hire on demonstrated building, not on claims or credentials. We expect you to be more capable than the hiring manager in your domain: you will co-own and challenge the tooling strategy, not just implement it. This is a two-person engineering team: you deploy, debug, and hotfix your counterpart’s stack when needed.

Requirements

  • 7+ years building and operating resilient backend or platform systems in production, including on-call ownership over time.
  • Proven brownfield migrations: you have taken a founder-built or prototype system to production grade while it stayed in daily use.
  • Strong engineering fundamentals: data structures and algorithms, fluent Python, strong SQL and data modeling, plus one additional systems language (TypeScript/Node or Go), with the full-stack range to connect the infrastructure yourself.
  • Agentic runtime and harness engineering, proven by building: the field is too young to demand years of it, so we weigh real systems shipped over tenure. You have genuinely built model-agnostic agent orchestration, model routing and evaluation across providers, agent SDKs (Claude Agent SDK or equivalent), MCP servers, skill- and hook-based agent tooling, and the evaluation and red-team harnesses that grade agent behavior, with responsible-AI controls (hallucination, bias, drift).
  • Data engineering with provenance and lineage, and security and data-protection engineering by default (encryption, identity and access management, secrets, retention). You build systems that could withstand external audit, by design.
  • Cloud (AWS or GCP) plus resilience engineering: infrastructure-as-code (Terraform), CI/CD, HA, DR, SLOs, and observability, backed by automated testing and documentation.
  • You ship inside locked-down enterprise environments (TLS-intercepting proxies, endpoint detection and response tooling, restricted installs, security guardrails, OAuth admin consent) without treating security as someone else’s problem.
  • Partnership: you translate audit needs into systems, explain technical risk to non-engineers, and drive adoption.

Nice To Haves

  • Multi-agent orchestration frameworks, MCP servers, and tooling and plugin development across the Anthropic, OpenAI, and Google model ecosystems.
  • Model-risk or AI-governance program experience; frameworks such as the NIST AI Risk Management Framework.
  • Continuous-auditing or continuous-controls-monitoring platforms; streaming and real-time data at scale; statistical anomaly detection and applied machine learning beyond LLMs; vector stores and retrieval; LLM cost engineering.
  • Experience passing external audit, SOC 2, or SOX (helpful, not required).
  • Crypto and blockchain literacy; regulated financial-services, fintech, or crypto experience.

Responsibilities

  • Inherit, operate, and progressively migrate the working prototype estate (Google Workspace–native automation across Apps Script, Drive, and the Docs/Sheets/Slides APIs; locally scheduled jobs; and Claude Code agent tooling) to the target platform without interrupting daily and board-cycle workflows.
  • Re-architect Hive Mind into a resilient AWS or GCP platform with high availability (HA), disaster recovery (DR), defined service-level objectives (SLOs), and full observability, and own it in production.
  • Build the agentic runtime and harness: orchestration, multi-model routing that sends each task to the right model, tool, MCP, and plugin integration across model providers, and the evaluation, red-team, and regression harnesses that grade agents before and after production, with evaluation gates, judge calibration, and cost controls that hold for each model, including prompt-injection and data-exfiltration threat modeling for agents that read untrusted content.
  • Build the Responsible-AI and model-governance layer: hallucination, bias, and drift controls, output validation, guardrails, and complete logging. Extend the existing governance design (human-calibrated judge gates, golden-set regression, provenance registry) rather than replacing it.
  • Build data infrastructure with provenance and lineage: immutable audit trails, versioned evidence, reproducible pipelines, and traceability from source to report.
  • Engineer data protection: encryption, key and secrets management, least-privilege access, sensitive-data handling, residency, and defensible retention.
  • Stand up the computer-assisted audit technique (CAAT) and continuous-monitoring data foundation: analytics over full populations, with exceptions streamed in real time.
  • Own the cloud foundation: infrastructure-as-code, CI/CD, identity, networking, observability, and cost controls, and make the platform examinable.

Benefits

  • Competitive total compensation package
  • L&D programs and Education subsidy for employees' growth and development
  • Various team building programs and company events
  • Wellness and meal allowances
  • Comprehensive healthcare schemes for employees and dependants
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service