Senior Machine Learning Engineer - Intelligent Document Processing / Production AI Systems

Pantheon DataReston, VA
$140,000 - $200,000Remote

About The Position

Pantheon Data is seeking a Senior Machine Learning Engineer to help build production-oriented AI and Intelligent Document Processing (IDP) systems. This role requires a hands-on engineer capable of moving beyond experiments to build working software that ingests, processes, analyzes, retrieves, and explains information from complex unstructured and semi-structured sources. The ideal candidate possesses deep expertise in machine learning, NLP, OCR, computer vision, LLMs, and retrieval-based systems, coupled with broader engineering judgment regarding data pipelines, APIs, databases, cloud infrastructure, containers, testing, evaluation, observability, and production failure modes. This is a practical, hands-on role, not solely focused on notebooks, prompts, or research. Successful candidates should be prepared to discuss specific systems they have built, detailing the data flow, model or inference architecture, deployment approach, evaluation strategy, failure modes, and their personal implementation contributions.

Requirements

  • Bachelor's degree in Computer Science, Engineering, or a related technical field from an ABET accredited university.
  • 5+ years of professional hands-on experience in machine learning engineering, AI engineering, data science engineering, or a closely related software engineering role. Plus an additional 5 years of experience in technology and software engineering.
  • Demonstrated experience building AI/ML systems beyond notebooks, prototypes, or demos. Candidates should have shipped or supported pipelines, services, APIs, inference endpoints, evaluation harnesses, or production-facing tools.
  • Strong Python engineering experience, including readable, maintainable code; debugging; testing; packaging; and integration with other systems.
  • Hands-on experience with NLP, OCR, computer vision, LLMs, embeddings, semantic search, RAG, or other document-understanding techniques.
  • Experience working with unstructured or semi-structured data such as PDFs, scanned documents, forms, tables, images, logs, contracts, technical manuals, or engineering documentation.
  • Ability to design and reason about end-to-end data flow: source data, preprocessing, transformation, model/inference step, persistence, API/service boundary, evaluation, and user-facing output.
  • Familiarity with common ML frameworks and tooling such as PyTorch, TensorFlow, scikit-learn, Hugging Face, MLflow, or similar technologies.
  • Experience with databases and data stores, including SQL and at least one relational or non-relational data platform.
  • Experience using Git-based development workflows, code review, issue tracking, and team-based software delivery practices.
  • Clear written and verbal communication skills, including the ability to explain technical tradeoffs, limitations, and failure modes.
  • Ability to work effectively remotely in cross-functional teams.
  • Ability to meet deadlines and produce quality work.
  • Proficient in Microsoft Suite software including Outlook, Word, Excel, SharePoint, and PowerPoint.

Nice To Haves

  • Direct experience with Intelligent Document Processing, document AI, OCR pipelines, table extraction, form extraction, layout-aware processing, or evidence-grounded retrieval.
  • Experience building, deploying, or operating LLM-backed systems, including inference serving, prompt/version management, model evaluation, retrieval, observability, or cost/latency management.
  • Experience with cloud platforms such as AWS or Azure, including storage, compute, serverless, networking basics, IAM, monitoring, or managed ML/AI services.
  • Experience with containers and deployment workflows, including Docker, Kubernetes, CI/CD pipelines, automated tests, and environment promotion.
  • Experience building user-facing or internal tools such as validation interfaces, review workflows, dashboards, admin tools, or lightweight full-stack applications.
  • Experience with data engineering tools such as pandas, NumPy, Spark/PySpark, Databricks, Airflow, or similar workflow/data platforms.
  • Experience with observability, performance profiling, or debugging tools such as Grafana, CloudWatch, TensorBoard, tracing tools, GPU profiling tools, or application logs.
  • Experience with evaluation design, benchmarking, reproducibility, statistical analysis, error analysis, or human-in-the-loop validation.
  • Bachelor's or advanced degree in Computer Science, Engineering, Mathematics, Physics, Statistics, Data Science, or another technical discipline. Equivalent professional experience will also be considered.
  • Demonstrated ability to mentor junior developers or contribute to team technical direction.

Responsibilities

  • Design and build AI/ML capabilities for Intelligent Document Processing, including OCR post-processing, document parsing, NLP/LLM extraction, semantic search, retrieval, evidence grounding, and structured data generation.
  • Develop production-quality Python services, pipelines, and tooling that turn messy source documents into reliable, traceable, usable information.
  • Work across the full lifecycle of AI systems: data ingestion, preprocessing, model or LLM integration, evaluation, deployment, monitoring, and iterative improvement.
  • Build and improve systems that process PDFs, scanned documents, tables, forms, drawings, images, technical manuals, and other complex document types.
  • Collaborate with software engineers, data engineers, cloud engineers, product leads, customers, and leadership to turn ambiguous technical problems into working solutions.
  • Make practical engineering decisions about when to use deterministic logic, classical NLP, OCR, embeddings, LLMs, fine-tuned models, or human review workflows.
  • Help establish engineering standards for evaluation, reproducibility, model behavior, data quality, traceability, and responsible use of AI in customer-facing systems.
  • Design, implement, and maintain ML/AI software components for IDP and Generative AI systems.
  • Build data pipelines for unstructured and semi-structured data, including document ingestion, extraction, cleaning, enrichment, validation, and storage.
  • Develop and evaluate NLP, OCR, computer vision, embedding, retrieval, and LLM-based approaches for document understanding use cases.
  • Create APIs, internal tools, review interfaces, dashboards, or validation workflows that allow engineers and users to inspect, correct, and trust system output.
  • Contribute production-quality code with clear structure, tests, logging, error handling, and documentation.
  • Deploy and support ML/AI services in cloud or containerized environments, including model serving, batch processing, and workflow automation.
  • Design evaluation approaches for extraction quality, retrieval quality, model behavior, hallucination risk, and end-to-end system performance.
  • Troubleshoot system behavior across model output, data quality, retrieval, schema design, infrastructure, latency, cost, and user workflow issues.
  • Mentor other engineers and help raise the technical quality of the team.
  • Communicate clearly with both technical and non-technical stakeholders, including project managers, customers, and executive leadership.

Benefits

  • Competitive salaries and benefits
  • SmartBenefits through the Washington Metro Area Transportation Authority
  • Tuition assistance may be available for continuing education expenses and certifications
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service