AI Inference Engineer

F5San Jose, CA
$176,600 - $265,000

About The Position

The AI Inference Engineer plays a critical role in the AI lifecycle by bridging the gap between high-performance model development and optimized deployment environments. This position focuses on optimizing Large Language Models (LLMs) for inference, serving diverse environments—from GPU-rich data centers to resource-constrained edge devices—with a strong emphasis on maximizing throughput, minimizing latency, and maintaining model accuracy. This role is pivotal in advancing F5’s AI capabilities, ensuring enterprise-grade reliability by leveraging hardware acceleration, designing scalable infrastructure, and monitoring system performance.

Requirements

  • Proficiency in programming languages such as Python, C++, Rust, or Golang specifically for high-performance AI workflows.
  • Proven hands-on experience with tools like vLLM, TensorRT, Llama.cpp, and Ollama for inference development and optimization.
  • Strong familiarity with infrastructure technologies, including Docker, Kubernetes, and cloud platforms such as AWS, GCP, and Azure.
  • Comprehensive understanding of GPU and AI hardware, including techniques for profiling and optimizing performance for accelerators like NVIDIA GPUs and TPUs.

Nice To Haves

  • Prior experience deploying Large Language Models (LLMs) with advanced techniques like Speculative Decoding or PagedAttention.
  • Contributions to open-source inference libraries or hardware-level kernel development (e.g., CUDA, Triton kernels).
  • Background in MLOps or SRE roles focused on high-performance AI endpoints and reliability during demand surges.
  • Proficiency in designing scalable solutions for high-throughput inference environments optimized for traffic bursts.

Responsibilities

  • Build and maintain robust inference engines using tools like vLLM, TGI (Text Generation Inference), and NVIDIA Triton, ensuring high performance at scale.
  • Handle deployment optimizations to deliver low-latency AI serving solutions for multiple business applications.
  • Profile and optimize models for specialized hardware backends, including NVIDIA GPUs (CUDA/TensorRT), Apple Silicon (CoreML), and AI accelerators like TPUs and LPUs.
  • Collaborate with hardware teams to maximize utilization and performance across various computational environments.
  • Design and implement auto-scaling architectures for online (real-time) and batch inference pipelines, leveraging Kubernetes for inference routing and orchestration.
  • Ensure software solutions are optimized for peak performance during traffic spikes, maintaining reliability and scalability.
  • Establish robust observability frameworks to monitor Time to First Token (TTFT), tokens per second, and memory bandwidth utilization against service-level agreements (SLAs).
  • Build and execute performance and load testing suites to identify bottlenecks and ensure consistent reliability at scale.

Benefits

  • Incentive compensation
  • Bonus
  • Restricted stock units
  • Benefits
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service