Forward Deployed Engineer (Generative AI)

Tiger Analytics Inc.
Onsite

About The Position

Tiger Analytics is seeking an experienced Forward Deployed Engineer (Generative AI) with Gen AI experience to join their advanced analytics consulting firm. This role involves the on-site deployment, integration, and scaling of enterprise Generative AI solutions, embedding directly within customer engineering teams to operationalize Large Language Models (LLMs) and retrieval systems across Google Cloud Platform (GCP). The FDE will bridge the gap between AI research and production-grade cloud infrastructure, collaborating with cross-functional teams and business partners to drive strategy and ensure business value. This position offers significant career development opportunities in a fast-growing, entrepreneurial environment with a high degree of individual responsibility.

Requirements

  • Advanced knowledge of Vertex AI primitives, including Vertex AI Studio, Model Registry, Endpoint deployment, Vertex AI Pipelines (Kubeflow), and Vertex AI Vector Search.
  • Hands-on experience with LLM orchestration tools (LangChain, LlamaIndex, AutoGen) and deep learning frameworks (PyTorch, Hugging Face) optimized for GCP infrastructure.
  • Production experience setting up, optimizing, and querying Vertex AI Vector Search, or managed vector stores like Milvus, Pinecone, and pgvector (Cloud SQL/Spanner).
  • Proficiency in model serving frameworks (vLLM, TGI, Triton Inference Server) deployed via Vertex AI or GKE, alongside robust automated model evaluation pipelines.
  • Deep expertise in Google Kubernetes Engine (GKE) for managing GPU/TPU workloads, autoscaling, and scheduling.
  • Mastery of Terraform to provision secure, complex GCP environments, IAM roles, and Vertex AI resources.
  • Strong coding skills in Python (preferred) or Go, with an emphasis on writing clean, concurrent code and utilizing the Google Cloud SDK.
  • Ability to manage customer expectations around LLM non-determinism, hallucinations, and performance trade-offs.
  • Passion for keeping pace with the weekly advancements in the Generative AI landscape.
  • Exceptional skill in isolating errors across complex software layers, from GPU drivers up to prompt engineering logic.
  • Willingness to travel to client sites to lead high-stakes, on-site deployment sprints.

Responsibilities

  • Deploy, fine-tune, and optimize large-scale Gen AI models and LLM orchestration frameworks within customer cloud environments.
  • Architect scalable infrastructure for AI workloads utilizing GPU/TPU orchestration, high-performance storage, and low-latency networking.
  • Design and implement high-throughput data ingestion pipelines and Vector Database architectures for Retrieval-Augmented Generation (RAG).
  • Act as the primary technical consultant, guiding enterprise clients through AI safety, prompt engineering patterns, and inference cost optimization.
  • Feed edge-case deployment insights back to core AI research and platform engineering teams to improve product robustness.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service