Applied AI Engineer Intern - Advancing Agent Quality

Brightstar.AI•Miami, FL
•Onsite

About The Position

Brightstar.AI is seeking a full-time or recently graduated Masters or PhD-level Applied AI Engineer Intern – Advancing Agent Quality with a strong research mindset, affiliated with a leading (South Florida) university, to contribute to its AI Think Tank and to the hands-on build-out of Brightstar’s proprietary AI systems. The intern will pair academic rigor with applied engineering: building the evaluation environments, automated raters, and diagnostic tooling that establish — with evidence rather than impression — whether an agent is reliable enough to put in front of operators, executives, and investors. The first concrete assignment is the AI Twin Board, Brightstar’s proprietary board-simulation product, where the intern will help build the evaluation harness end to end: realistic task suites, calibrated LLM judge, failure-mode diagnostics, and the data flywheel that feeds improvements back into the agents themselves. The AI Twin Board is the starting point, not the full scope — as Brightstar’s AI portfolio expands, the same quality discipline will be applied to other agentic systems, internal platforms, and portfolio-company deployments. Working in coordination with university professors and research leaders, the intern will also bring up-to-date knowledge of the fast-evolving AI landscape — especially large language models, reasoning systems, emerging model capabilities, and the societal implications of advanced AI — and help translate what current and next-generation systems may enable over the next six months to three years into Brightstar’s evaluation standards and Think Tank point of view.

Requirements

  • Currently enrolled in a Masters or PhD program — or serving as a postdoctoral or early-career academic researcher — in Artificial Intelligence, Machine Learning, Computer Science, Computational Linguistics, or a related field.
  • 1 years of research, project, or professional experience with LLMs or Agents.
  • 2 year of experience with machine learning and deep learning.
  • Experience in software engineering and cloud-based development.
  • Experience with agent builders, e.g. OpenSDK, OpenAI Agent Builder, etc.
  • Able to communicate complex technical concepts clearly to non-technical executive audiences.

Nice To Haves

  • In-depth knowledge of machine learning algorithms, including supervised learning and reinforcement learning.
  • Hands-on experience building LLM or agent evaluation systems — autoraters, human-evaluation pipelines, benchmark design, or evaluation infrastructure.
  • Experience with multi-agent orchestration, retrieval-augmented generation, and tool-use / function-calling reliability.
  • Active research or publications in LLMs, reasoning systems, model behavior, AI safety, or the societal impact of advanced AI systems.
  • Strong research discipline, intellectual curiosity and the ability to separate meaningful signal from AI hype.
  • Affiliated with a (South Florida) university and working under the supervision of a faculty member or professor, or recent graduate.

Responsibilities

  • Design, build, and scale realistic agent environments and task suites that reflect how Brightstar’s agents are actually used — board-level deliberation, research synthesis, diligence workflows, and multi-step tool use.
  • Develop agentic LLM judge, calibrate them against human expert evaluation, and benchmark Brightstar agents against leading frontier models.
  • Build trajectory-analysis frameworks and diagnostic tooling that root-cause agent failure modes — passivity, hallucination, persona drift, brittle tool execution — and feed fixes back into agent prompts, harnesses, and architecture.
  • Enable agents to generate and iterate on their own verifiers (automated test cases, checklists, ground-truth sets) to support effective exploration and iterative problem-solving.
  • Harvest multi-turn interaction trajectories into high-quality datasets and reward signals that support fine-tuning and post-training of the models Brightstar relies on.
  • Contribute to the AI Twin Board build-out as the initial focus, and extend the same evaluation approach to other Brightstar AI initiatives as the roadmap grows.
  • Contribute frontier AI research perspectives to the Brightstar.AI Think Tank as part of weekly review sessions and select strategic discussions.
  • Review relevant research papers, conference findings, and publications from leading AI institutions and labs as inputs to the Think Tank knowledge base and to Brightstar’s evaluation methodology.
  • Help interpret academic and technical developments for a business and investment audience, including implications for industry transformation and applied AI opportunities.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service