Deepgram is seeking a Senior Software Engineer - Model Evaluation & AI Systems to join a team focused on validating the quality of their speech, audio, and multilingual models before customer release. This team is responsible for the evaluation and quality assurance processes that ensure Deepgram's models, including speech-to-text, text-to-speech, and increasingly LLM- and multimodal-powered systems, meet performance targets in both batch and streaming environments. The role involves building pipelines, harnesses, canaries, and test frameworks to detect regressions, hallucinations, and quality issues, and collaborating with Research to translate model expectations into automated, reproducible checks. The engineer will define evaluation methodology, build infrastructure for large-scale model quality measurement, create evaluation pipelines, establish pass/fail criteria based on Research benchmarks, and develop monitoring systems to ensure model integrity in production. This position directly impacts customer experience by providing trusted signals for release and optimization decisions. The ideal candidate is a strong engineer comfortable with building test infrastructure and analyzing model behavior, with hands-on experience evaluating modern AI systems being a significant advantage.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Senior
Education Level
No Education Listed