We are building our level 2 technical support team from the ground up, and we are looking for a Technical Support Engineer to join it. Our customers are not filing tickets about forgotten passwords; they are AI companies running massive GPU training jobs and production inference at scale, and when they contact us, something real is wrong. A node is reporting 7 GPUs instead of 8. A training run that has been going for 4 days is suddenly crawling, and nobody can say why. A Slurm partition is draining, and the queue is backing up. Our global operations center catches and handles what the runbooks cover, around the clock, everything past that comes to you. This is a hands-on diagnostic role for someone who genuinely enjoys the hunt. You will work Linux systems at depth, live inside Kubernetes and Slurm, read logs and metrics until the story makes sense, and talk directly to customer engineers who are every bit as technical as you are. When you solve something new, you will write it down so the global operations center can solve it next time without you. Because the tier is new, you will help define it: the escalation paths, the runbooks, the diagnostic tooling, the standards. If you have ever looked at a support organization and thought I could build this properly if someone would let me, this is that opening: Reporting to the technical support manager within customer experience.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level
Education Level
No Education Listed