Applied Intuition is seeking a software engineer with deep experience in optimizing ML models and deploying them on production-grade embedded runtime environments. The role involves working across the entire ML framework stack (e.g. PyTorch, JAX, ONNX, TensorRT, CUDA, XLA, Triton). The engineer will drive ML performance optimization on multiple technologies for on-road and off-road ADAS / AD stacks targeting deployment on a variety of embedded compute platforms. This includes developing compute usage strategies to optimize efficiency and latency of model inference, working on model pruning and quantization, and supporting deployment on memory constrained platforms. The role also involves collaborating closely with ML engineers and software developers to find and optimize efficient model architecture solutions, and setting up methodologies to profile model performance on target embedded compute platforms to identify performance bottlenecks during stack integration.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level