Applied AI Engineer — Global Health & AI (REMOTE ROLE)

ICF•Rockville, MD
•$98,614 - $167,644•Remote

About The Position

ICF's Global Health and Development (GHD) line of business, part of the Health, People, and Human Services (HPHS) Group, is seeking an Applied AI Engineer to build frontier-model systems for various global health and survey programs. This role will support the development of AI-enabled tools that draft survey reports, let researchers and policymakers analyze data in natural language, translate questionnaires into low-resource languages, and support interviewers and data quality in the field. No tool is adopted unless it matches or beats the expert-led process it replaces. You will build the prompting, agent, and pipeline layer across the program. You will also build the evaluation harnesses that show whether each tool actually works, judged against ground truth produced by some of the most experienced survey methodologists in the world. This position can be remote within the United States, with occasional travel. It is contingent upon award of program funding.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related field (or equivalent years experience)
  • Minimum 4 years of professional software development experience, including building LLM-powered systems that real users relied on in production
  • U.S. citizenship is required by federal government contract
  • Proficiency in Python or another modern programming language (e.g., JavaScript/TypeScript, Go) for back-end and AI application development
  • Hands-on experience building with frontier-model APIs (from any provider, or through any cloud platform), including tool use and retrieval-augmented generation
  • Demonstrated experience designing LLM evaluations, including test-set construction, automated and human grading, regression testing, and ship/no-ship decisions.
  • Working knowledge of LLM cost, latency, and model-selection trade-offs
  • Experience building and deploying on a major cloud platform (e.g., AWS, Azure, or Google Cloud)
  • Demonstrated experience using agentic coding tools, such as Claude Code, OpenAI Codex, Cursor, Antigravity, or GitHub Copilot agent mode, to specify, review, test, and take ownership of production code.
  • Experience with agent frameworks and AI evaluation tooling.

Nice To Haves

  • Experience with tasks where ground truth came from subject matter experts (for example, statisticians, clinicians, scientists, or methodologists)
  • Experience with grounded generation over structured or tabular data, or with LLM-assisted statistical analysis
  • Experience with multilingual or low-resource-language NLP, translation, or speech
  • Experience deploying open-weight models
  • Experience with sensitive or regulated data, such as health data or human-subjects research
  • Familiarity with model governance, monitoring, and responsible AI practice
  • Experience in global health, international development, or low- and middle-income country settings
  • Public work such as open-source contributions, published evaluations, or technical writing about systems you built

Responsibilities

  • Design and build LLM pipelines and agents for survey workflows, including: an AI-assisted data-to-report pipeline, natural-language analysis of survey data, multilingual questionnaire translation, field support tools for interviewers and supervisors
  • Build evaluation harnesses as first-class deliverables. Work with the Technical Lead and survey experts to define acceptance criteria before the work is judged. Build test sets from expert-produced ground truth and run the non-inferiority comparisons that decide whether a tool moves forward.
  • Route work across models based on reasoning complexity, volume, cost, latency, and data-residency requirements.
  • Manage inference cost by tracking token usage, context limits, prompt caching, and batch processing, and account for these operating costs in model-selection decisions.
  • Connect models to data services and to the applications built by full stack developers.
  • Use agentic coding tools to produce software from clear specifications, then review, test, and own the resulting code.
  • Monitor production systems, conduct disciplined error analysis, and report positive and negative results.
  • Document methods and results for public release, including negative results. Help transfer tools to partner-country institutions so they can run them on their own.

Benefits

  • Reasonable Accommodations are available, including, but not limited to, for disabled veterans, individuals with disabilities, and individuals with sincerely held religious beliefs, in all phases of the application and employment process.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service