AI Engineer

NetBrain•Burlington, MA
•$150,000 - $180,000

About The Position

NetBrain is seeking a Senior AI Engineer to design and build production-grade agent and RAG systems that power intelligent, reliable automation across its platform. This role involves hands-on engineering and system-level thinking, covering architecture, evaluation, scalability, observability, and production reliability. The ideal candidate should be comfortable with ambiguity, capable of moving quickly from prototype to production, and possess a strong focus on quality, safety, and real-world impact.

Requirements

  • Bachelor's degree or higher in Computer Science, Artificial Intelligence, Electrical Engineering, or a related technical field; equivalent practical experience will also be considered.
  • 3+ years of experience in software engineering, machine learning, or applied AI.
  • 2+ years building, deploying, and operating production-grade LLM or Agent applications.
  • Must have delivered at least one LLM-powered feature end-to-end and owned its ongoing operation and improvement after production launch.
  • Deep understanding of Agent architectures and LLM behavioral characteristics.
  • Hands-on experience building multi-step workflows involving reasoning, tool execution, state management, structured outputs, validation, and error recovery.
  • Proven ability to diagnose and resolve production LLM/Agent failures.
  • Strong Python and distributed backend engineering skills.
  • Hands-on experience designing evaluation systems for LLM applications.
  • Strong understanding of security risks associated with LLM and Agent applications.
  • Ability to independently design, implement, debug, deploy, and operate complex production systems.

Nice To Haves

  • Master's or Ph.D. preferred.
  • Strong experience with RAG and advanced retrieval systems.
  • Familiarity with Agent frameworks such as LangGraph, LangChain, AutoGen, and LlamaIndex.
  • Experience designing Agent runtime mechanisms, including Human-in-the-Loop workflows.
  • Experience with LLM fine-tuning, including LoRA or other parameter-efficient fine-tuning techniques.
  • Experience applying LLM technologies to networking, infrastructure, cybersecurity, observability, or other complex technical domains.
  • Fluent in both English and Chinese, with strong cross-regional communication and collaboration skills.

Responsibilities

  • Design and implement core capabilities for an enterprise-grade Agent platform, including orchestration patterns, tool execution, context and memory management, and safety guardrails.
  • Design enterprise-grade Agent execution and governance mechanisms, including Human-in-the-Loop approval workflows, multi-tenant permission isolation, policy enforcement, and secure execution controls.
  • Build reusable Agent Skills, standardized tool interfaces, and a scalable tool ecosystem integrated with NetBrain platform capabilities and business workflows.
  • Design and implement LLM post-training strategies, including domain-specific Supervised Fine-Tuning (SFT), DPO/RLHF-based preference alignment, and parameter-efficient fine-tuning techniques like LoRA.
  • Build an Agent self-learning feedback loop that converts production execution traces, user feedback, and evaluation results into high-quality datasets for continuous improvement.
  • Analyze and optimize LLM behavior across areas such as instruction following, tool calling, structured output generation, contextual understanding, reasoning stability, and hallucination mitigation.
  • Build production-grade LLM and Agent evaluation frameworks and automated regression pipelines.
  • Establish release quality gates and hallucination-detection mechanisms for AI features.
  • Build comprehensive AI system observability capabilities, including distributed tracing, structured logging, metrics, dashboards, and alerting.
  • Rapidly diagnose and resolve production AI failures.
  • Design and implement highly reliable backend services for production AI and Agent workloads.
  • Continuously optimize latency, throughput, token consumption, and infrastructure cost.
  • Independently diagnose and resolve complex AI system issues.
  • Lead technical design for critical modules and system-level capabilities.
  • Drive technical improvements based on production data, evaluation results, and benchmarks.
  • Prototype, benchmark, and productionize emerging technologies such as GraphRAG, Knowledge Graphs, MCP, LLM Post-Training, and Agent Self-Learning.
  • Continuously evaluate Agent frameworks and supporting infrastructure, and provide technical recommendations.
  • Stay current with developments in LLM and Agent technologies and rapidly translate promising technologies into production-ready capabilities.

Benefits

  • 401k
  • medical/dental coverage
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service