Director of AIML Engineering

ELLKAY, LLC• US,
•$200,000 - $230,000•Hybrid

About The Position

Ellkay is seeking a highly technical Director of AI/ML Engineering to lead, scale, and mentor a top-tier team of Staff AI Systems Integration Engineers and MLOps/GenAIOps Engineers. In this role, you will own the technical architecture, execution, and operational performance of enterprise-grade, autonomous Agentic AI systems, LLM workflows, and core machine learning platforms. As a technical leader with high-fidelity thinking, you will bridge the gap between complex engineering execution and strategic vision. You will partner closely with Product Management to define product scope, translate business goals into scalable technical strategies, and build robust, end-to-end solutions that operate seamlessly within production healthcare ecosystems.

Requirements

  • 12+ years of total software engineering and machine learning experience, with at least 4+ years in engineering leadership (Director, Senior Engineering Manager, or Principal Lead) managing high-performing AI/ML engineering teams.
  • Deep Agentic AI & LLM Systems Expertise: Hands-on experience designing and operating production-grade agentic workflows, dynamic tool/function-calling, multi-agent frameworks (LangGraph, AutoGen, CrewAI, Bedrock Agents), and state persistence engines.
  • Strong MLOps / GenAIOps Foundation: Direct experience overseeing model/agent lifecycle automation, including AWS SageMaker (Pipelines, Model Registry, Model Monitor), prompt/agent versioning, synthetic evaluations, and trajectory observability tools (LangSmith, Arize, Phoenix, CloudWatch).
  • Cloud Architecture & Systems Engineering: Expert-level knowledge of distributed systems, cloud-native AWS architectures (ECS Fargate, Lambda, Step Functions, EventBridge, DynamoDB, S3), Infrastructure as Code (AWS CDK in Python/TypeScript), and container execution (Docker, ECR).
  • AI Development: Direct experience building custom AI dev tooling, internal CLI agents, or custom LLM-based developer productivity tools such as Claude Code/Codex.
  • Exceptional Communication Skills: Proven ability to articulate complex technical architectures, risks, and trade-offs clearly to executive stakeholders, cross-functional partners, and external technical teams.

Nice To Haves

  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Electrical Engineering, or a related quantitative technical field.
  • Experience implementing human-in-the-loop (HITL) intervention protocols, self-correction/reflection loops, and safety guardrails for autonomous agents in mission-critical production environments.

Responsibilities

  • Partner directly with Product Management and executive leadership to define product scope, multi-agent capabilities, and technical roadmaps.
  • Translate broad, complex business and product requirements into clear technical designs, multi-agent orchestration architectures, and actionable engineering milestones.
  • Drive high-fidelity problem-solving to evaluate build-vs-buy decisions, foundational model selections, tool-calling strategies, and infrastructure investments.
  • Oversee the end-to-end implementation of multi-agent coordination frameworks (e.g., Bedrock Agents, LangGraph, AutoGen), agentic state persistence layers, and tool-use pipelines.
  • Provide architectural oversight across real-time and event-driven data pipelines, REST/gRPC microservices, AWS CDK constructs, and healthcare data integration standards (FHIR R4, HL7 v2, USCDI).
  • Champion rigorous engineering standards across the team for code quality, system integration, agent execution trajectories, unit/integration testing, and deployment hygiene.
  • Oversee the operational lifecycle of AI/ML models and autonomous agents, including automated retraining pipelines, prompt/agent versioning, vector index versioning, and continuous evaluation (LLM-as-a-Judge, held-out test sets).
  • Enforce robust AgentOps and observability practices using trajectory tracking, token/cost monitoring, execution latency tracking, and guardrail performance metrics.
  • Ensure all AI/ML architectures and agent runtimes comply with strict regulatory frameworks, including HIPAA, PHI/PII redaction, least-privilege tool execution controls, and zero-leakage data governance.
  • Lead, mentor, and expand a high-performing engineering organization, fostering a culture of technical excellence, accountability, and continuous learning.
  • Seamlessly operate across multiple technical contexts—from low-level container engineering, AWS infrastructure, and model fine-tuning to high-level multi-agent workflow engines and business strategy.
  • Serve as the primary technical interface and bridge between AI Engineers, MLOps Engineers, Platform Infrastructure, Product, Security, and Enterprise Leadership.

Benefits

  • Medical, Dental, and Vision benefits
  • Employer-paid Life and LTD
  • 401k w/ matching
  • Work/life balance
  • Paid Volunteer Program
  • Flexible working hours
  • Generous FTO
  • Remote work options
  • Employee Discounts
  • Parental Leave
  • Gym membership / Exercise class stipends
  • On site in HQ Free daily lunches
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service