Senior AI Engineer

NextDeavorNew York, NY
Remote

About The Position

You will help build and advance autonomous offensive security agents and the models that power them, delivering capabilities that improve exploit chaining, multimodal vision, mobile coverage, and customer-facing accuracy. You will collaborate with engineering and platform stakeholders to ship production-grade systems and run in EST hours as a remote contributor.

Requirements

  • 5+ years building production ML/AI systems, including at least 2 years working directly on LLMs or LLM-powered agents
  • Deep Python and strong production engineering practices (testing, code review, observability)
  • Hands-on fine-tuning experience: SFT, preference optimization (DPO, GRPO, RLHF/RLAIF), data curation, and synthetic data generation
  • Strong grasp of transformer architectures and training stack (PyTorch, Hugging Face, DeepSpeed or FSDP, accelerate)
  • Experience designing and shipping multi-agent or tool-using LLM systems in production
  • Rigorous evaluation design experience: building harnesses, tracking experiments, and data-driven decision making
  • Inference optimization experience (vLLM, TensorRT-LLM, quantization, throughput/latency tradeoffs)
  • Familiarity with retrieval pipelines, vector stores, and structured memory for agents
  • Kubernetes and containerized deployment fluency
  • Genuine interest in offensive security and the ability to ramp quickly on OWASP Top 10, API/web/mobile pentesting concepts

Nice To Haves

  • Offensive security certifications or experience (OSCP/OSWE/OSWA, CTF, bug bounty, red team)
  • Research publications at top ML/security venues or open source contributions to agent/LLM tooling
  • Experience with adversarial ML or red-teaming AI systems
  • Familiarity with mobile app reverse engineering or binary analysis

Responsibilities

  • Design, implement, and iterate on named agents, including orchestration patterns, hand-offs, planning loops, tool use, and shared memory
  • Contribute to model training and fine-tuning across data curation, supervised fine-tuning (SFT), preference optimization (DPO/GRPO/RLHF-style), and evaluation
  • Extend the co-evolutionary self-training (Javelin) loop so the system improves from its engagements
  • Build self-improvement systems (false-positive detection, tiered skill learning, agent directives, code-patch proposals) and pipelines for human approval
  • Design security-specific evaluations covering OWASP Top 10, exploit chaining, finding accuracy, and agent reliability; track performance over model and agent changes
  • Contribute to multimodal (vision) and mobile (iOS/Android) coverage and BYOK support efforts
  • Own production reliability: latency, cost, observability, failure-mode analysis, and Kubernetes-based deployment
  • Improve customer-facing accuracy surfaces and live accuracy gauges exposed to customers
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service