This role focuses on designing and executing fine-tuning pipelines for Vision-Language Models (VLMs) on domain-specific imagery datasets. The engineer will be responsible for data preprocessing, training orchestration, and hyperparameter optimization. They will also develop and implement evaluation frameworks for multimodal model performance, including task-specific metrics for image understanding, visual question answering, and spatial reasoning. A key aspect of the role involves building scalable training infrastructure on AWS (SageMaker, EC2 GPU instances) for distributed fine-tuning of large multimodal models. Additionally, the engineer will create data pipelines for curating, annotating, and transforming geospatial imagery datasets into model-ready formats for supervised and instruction-tuning workflows. Collaboration with applied scientists and solutions architects is essential for iterating on model architectures, adapter strategies (LoRA/QLoRA), and inference optimization techniques.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed