AI Platform Engineer

Bright Vision TechnologiesApex, MO
$130,000 - $180,000Remote

About The Position

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States. This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential. Bright Vision Technologies is seeking a highly experienced AI Platform Engineer with 10+ years of experience in distributed systems, cloud-native infrastructure, and AI platform engineering to design, build, and operate enterprise-scale AI inference and machine learning platforms. The ideal candidate will possess deep expertise in LLM serving, GPU optimization, Kubernetes, cloud infrastructure, distributed systems, and MLOps, with a proven ability to deliver highly scalable, reliable, secure, and cost-efficient AI platforms supporting production machine learning workloads.

Requirements

  • 6+ years of experience in distributed systems, infrastructure, or ML platform engineering.
  • Strong proficiency in Python and Go, Rust, or C++.
  • Experience with LLM inference frameworks (vLLM, TensorRT-LLM), Kubernetes, cloud platforms, and GPU optimization.
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Artificial Intelligence, or a related technical discipline.
  • 10+ years of professional experience in distributed systems, infrastructure engineering, cloud platforms, or machine learning platform engineering.
  • Strong programming skills in Python and at least one systems programming language such as Go, Rust, or C++.
  • Extensive experience with Large Language Model (LLM) serving, model inference optimization, and production AI infrastructure.
  • Hands-on experience with vLLM, TensorRT-LLM, Triton Inference Server, Ray Serve, or similar AI serving frameworks.
  • Strong expertise in Kubernetes, container orchestration, Docker, and cloud-native application architectures.
  • Experience optimizing GPU workloads using CUDA, NVIDIA GPU technologies, distributed inference, and high-performance AI infrastructure.
  • Experience with cloud platforms including AWS, Microsoft Azure, or Google Cloud Platform (GCP).
  • Strong understanding of distributed systems, networking, scalability, observability, and security best practices.
  • Excellent analytical, communication, collaboration, and technical leadership skills.

Nice To Haves

  • Experience designing and operating multi-region AI platforms and globally distributed inference services.
  • Knowledge of model optimization techniques such as quantization, pruning, compression, speculative decoding, KV cache optimization, and mixed-precision inference.
  • Experience with MLOps, GitOps, Infrastructure as Code (Terraform, Bicep, CloudFormation), and CI/CD automation.
  • Familiarity with service mesh technologies such as Istio or Linkerd, API gateways, and event-driven architectures.
  • Contributions to open-source AI infrastructure projects, technical publications, patents, or conference presentations.
  • Experience implementing FinOps strategies, cloud cost optimization, and enterprise AI governance.
  • Experience with multi-region AI deployments and AI infrastructure.
  • Familiarity with model optimization techniques such as quantization or compression.
  • Open-source contributions or experience supporting large-scale AI APIs.

Responsibilities

  • Design and maintain scalable AI model serving platforms.
  • Optimize inference performance, GPU utilization, and request routing.
  • Build autoscaling, deployment, and monitoring solutions.
  • Implement caching, security, and high-availability strategies.
  • Collaborate with ML teams to deploy and support production AI models.
  • Design, build, and maintain scalable AI inference and model-serving platforms for enterprise production environments.
  • Architect highly available, cloud-native infrastructure supporting Large Language Models (LLMs), foundation models, and machine learning services.
  • Optimize inference latency, throughput, GPU utilization, memory management, and request scheduling across distributed AI workloads.
  • Design autoscaling, workload orchestration, traffic management, and intelligent request routing strategies for AI services.
  • Implement model deployment, versioning, rollback, and lifecycle management using modern MLOps practices.
  • Develop monitoring, observability, logging, distributed tracing, and alerting solutions to ensure platform reliability and performance.
  • Implement caching strategies, API gateways, security controls, authentication, authorization, and high-availability architectures.
  • Collaborate with AI researchers, ML engineers, DevOps teams, and software engineers to deploy and support production AI models.
  • Drive cloud infrastructure optimization, resource utilization, FinOps initiatives, and operational excellence.
  • Mentor engineering teams, conduct architecture reviews, and establish best practices for AI platform engineering and cloud-native development.
  • Evaluate emerging AI infrastructure technologies, model-serving frameworks, and GPU acceleration techniques to drive continuous innovation.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service