Foundation Model Engineer

Bright Vision TechnologiesRichardson, TX
$200,000 - $230,000Remote

About The Position

We are looking for an Foundation Model Engineer to design, execute, and operationalize fine-tuning workflows for large language models across supervised, preference-based, and reinforcement learning approaches. The role requires deep practical experience with modern training stacks, careful dataset construction, rigorous evaluation methodology, and the engineering discipline to operate complex training pipelines reliably. The ideal candidate combines strong ML intuition with production-grade engineering practices, and is comfortable navigating the trade-offs between data quality, compute budget, evaluation rigor, and shipping velocity. In this role you will work closely with cross-functional partners — product, design, engineering, operations, and business stakeholders — to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production.

Requirements

  • Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, or a related field. Equivalent industry experience will also be considered.
  • 6+ years of experience in Machine Learning or AI engineering, including at least 2 years working with Large Language Models (LLMs) or Generative AI.
  • Strong programming skills in Python.
  • Hands-on experience with PyTorch and Hugging Face Transformers.
  • Experience fine-tuning open-source LLMs such as Llama, Mistral, Falcon, Gemma, or similar transformer-based models.
  • Knowledge of parameter-efficient fine-tuning techniques including LoRA, QLoRA, or PEFT.
  • Experience building data preprocessing and model training pipelines.
  • Familiarity with vector databases, embeddings, and Retrieval-Augmented Generation (RAG) concepts.
  • Experience with cloud platforms such as AWS, Azure, or Google Cloud Platform.
  • Working knowledge of Docker, Kubernetes, Git, and CI/CD practices.
  • Strong analytical, problem-solving, and debugging skills.
  • Excellent communication and collaboration skills.

Nice To Haves

  • Experience with LangChain, LlamaIndex, DSPy, or similar LLM orchestration frameworks.
  • Familiarity with distributed model training technologies such as DeepSpeed or FSDP.
  • Experience deploying LLMs using vLLM, TensorRT-LLM, or ONNX Runtime.
  • Knowledge of ML lifecycle and MLOps tools such as MLflow, Kubeflow, or Weights & Biases.
  • Experience building enterprise AI chatbots, copilots, document intelligence, or RAG-based applications.
  • Exposure to prompt engineering, AI evaluation frameworks, and LLM safety best practices.
  • Contributions to open-source AI or machine learning projects are a plus.
  • Experience working in Agile software development environments.

Responsibilities

  • Develop, fine-tune, and optimize Large Language Models (LLMs) for enterprise AI applications.
  • Design and implement end-to-end model training and fine-tuning pipelines using PyTorch, Hugging Face Transformers, and related frameworks.
  • Prepare, clean, and curate training datasets for supervised fine-tuning and instruction tuning.
  • Implement parameter-efficient fine-tuning techniques such as LoRA, QLoRA, and PEFT to improve training efficiency.
  • Evaluate model performance using standard NLP benchmarks, automated metrics, and task-specific evaluations.
  • Collaborate with data scientists, software engineers, and product teams to integrate LLM solutions into production applications.
  • Optimize model inference performance, latency, and resource utilization for scalable deployment.
  • Develop data preprocessing, feature engineering, and automation scripts using Python.
  • Deploy and monitor machine learning models on cloud platforms using MLOps best practices.
  • Troubleshoot model training issues, optimize hyperparameters, and improve overall model accuracy.
  • Document technical designs, experiments, and implementation details for knowledge sharing.
  • Stay current with advancements in Generative AI, transformer architectures, and open-source LLM technologies.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service