Senior Machine Learning Engineer

Pantheon DataReston, VA
Remote

About The Position

Pantheon Data is seeking a Senior Machine Learning Engineer to design, build, and operate production AI systems for federal clients - including hybrid retrieval-augmented generation (RAG) applications, Intelligent Document Processing (IDP) pipelines, and LLM-backed decision-support tools running in AWS GovCloud. This is a hands-on, production-focused engineering role. You will work with GovCloud-hosted foundation models on Amazon Bedrock, build hybrid retrieval over relational and vector data stores, design rigorous evaluation harnesses, and ship secure, traceable, evidence-grounded systems in FedRAMP High / DoD IL4–IL5 environments. This is not a notebook-only, prompt-only, or research-only role: successful candidates can walk through real systems they have built - the data flow, retrieval and inference architecture, deployment approach, evaluation strategy, failure modes, and what they personally implemented.

Requirements

  • Bachelor's degree in Computer Science , Engineering, Mathematics, Data Science, or a related technical field, or equivalent professional experience.
  • 8+ years of professional software engineering experience, including 5+ years hands-on in machine learning / AI engineering or a closely related role.
  • Demonstrated experience shipping AI/ML systems beyond notebooks and demos: production pipelines, services, APIs, inference endpoints, evaluation harnesses, or customer-facing tools you built and supported.
  • Hands-on experience building LLM-backed applications with managed foundation-model services - Amazon Bedrock strongly preferred (Claude, Titan, or similar models), ideally in AWS GovCloud or another regulated/isolated environment.
  • Deep experience with RAG architectures: embeddings, vector search ( pgvector or comparable), hybrid retrieval (semantic + lexical + structured), chunking strategies, re-ranking or rank fusion, and citation/evidence grounding.
  • Experience designing and running LLM/RAG evaluations: retrieval and answer-quality metrics, groundedness and hallucination measurement, golden datasets, automated regression evals, and human-in-the-loop validation.
  • Strong Python engineering skills: readable, maintainable, tested code; debugging; packaging; and integration with other systems.
  • Experience with AWS services relevant to this stack: Lambda, Step Functions, S3, IAM, CloudWatch, and relational databases (PostgreSQL preferred).
  • Experience working with unstructured or semi-structured data: PDFs, scanned documents, forms, tables, technical manuals, drawings, or engineering documentation.
  • Working knowledge of AI security practices: guardrails, prompt-injection defenses, output validation, data redaction, and least-privilege access patterns.
  • Ability to design and reason about end-to-end data flow: source data, preprocessing, retrieval, model/inference step, persistence, API/service boundary, evaluation, and user-facing output.
  • Experience with Git-based workflows, code review, CI/CD, and team-based software delivery.
  • Clear written and verbal communication, including explaining tradeoffs, limitations, and failure modes.
  • Ability to work effectively remotely in cross-functional teams.
  • Ability to meet deadlines and produce quality work.
  • Proficient in Microsoft Suite software including Outlook, Word, Excel, SharePoint, and PowerPoint.

Nice To Haves

  • Direct experience operating LLM workloads in AWS GovCloud under FedRAMP High or DoD IL4/IL5, including private VPC endpoints, zero- or restricted-egress boundaries, and ATO-supporting documentation.
  • Experience with Bedrock features beyond basic inference: Guardrails, Knowledge Bases, Agents, Model Evaluation, provisioned throughput, and cross-model routing.
  • Intelligent Document Processing depth: OCR pipelines (e.g., Textract ), layout-aware models, table/form extraction, and document quality scoring.
  • Familiarity with federal compliance frameworks and their engineering implications: NIST 800-53, CMMC, CUI handling, SBOM/supply-chain controls, and STIG-hardened environments.
  • Experience with infrastructure as code (Terraform), containerized deployment (Docker/ECS), and CI/CD pipelines (GitLab CI or GitHub Actions) in regulated environments.
  • Experience with observability and LLMOps tooling: structured logging, tracing, model/prompt version management, drift detection, and cost/latency dashboards.
  • Experience building internal validation tools, review interfaces, dashboards, or lightweight full-stack applications (e.g., Next.js/React front ends over Python services).
  • Familiarity with common ML frameworks and tooling: PyTorch , Hugging Face, scikit-learn, MLflow , or similar.
  • Advanced degree in a technical discipline.
  • Demonstrated mentorship of junior engineers or contribution to team technical direction.

Responsibilities

  • Design, implement, and maintain ML/AI software components for RAG, IDP, and generative AI systems serving federal customers.
  • Build and operate data pipelines for unstructured and semi-structured data: ingestion, extraction, cleaning, enrichment, validation, quality scoring, quarantine/review, and storage.
  • Develop and evaluate NLP, OCR, computer vision, embedding, retrieval, and LLM-based approaches for document understanding and question answering.
  • Contribute production-quality Python with clear structure, tests, logging, error handling, and documentation; participate in code review and team-based delivery.
  • Deploy and support AI services in cloud and containerized environments, including Bedrock model integration, batch processing, workflow orchestration, and observability (CloudWatch, tracing, structured logs).
  • Define evaluation methodology and metrics; build eval datasets and harnesses; analyze errors; and drive iterative improvement of retrieval and generation quality.
  • Implement responsible-AI and security controls: Bedrock Guardrails or equivalent, PII/sensitive-data redaction, content filtering, prompt-injection mitigation, and output traceability to source evidence.
  • Troubleshoot system behavior across model output, data quality, retrieval, schema design, infrastructure, latency, cost, and user workflow.
  • Mentor other engineers, help establish standards for evaluation, reproducibility, and responsible AI use, and raise the technical quality of the team.
  • Communicate clearly with technical and non-technical stakeholders, including project managers, customers, and executive leadership.

Benefits

  • SmartBenefits through the Washington Metro Area Transportation Authority
  • Tuition assistance may be available for continuing education expenses and certifications
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service