Software Test Engineer

DeepgramRemote, CA
$150,000 - $220,000

About The Position

Deepgram is seeking a Software Test Engineer to design, build, and maintain automated test frameworks and exploratory test suites across their products, models, APIs, and data platforms. The role involves identifying system weaknesses, testing edge cases and adversarial inputs, and automating validation to quickly detect regressions. The engineer will translate product requirements and model metrics into automated regression tests, evaluation pipelines, data-quality gates, load tests, and release criteria. Collaboration with QA, Research, Product, Data, and Engineering teams is key for planning testing, executing evaluations, supporting user acceptance testing, and communicating risks. The position requires providing precise reproduction steps for any issues found. The core excitement of this role lies in building scalable automation that ensures the reliability of Deepgram's products, models, and data workflows for customers.

Requirements

  • BS, MS, or PhD in Computer Science, AI, Applied Math, or a related field, or equivalent experience.
  • 5+ years of professional software or QA engineering experience, with a track record of shipping test infrastructure or evaluation systems (senior candidates with significantly deeper experience welcome).
  • Solid backend/scripting experience in a language such as Python, Rust, Go, or similar.
  • Experience designing and building automated test pipelines, evaluation frameworks, or data-processing systems.
  • Strong analytical skills and comfort reasoning about metrics, thresholds, and statistical variation in results — able to distinguish real regressions from noise.
  • Ability to take charge of ambiguous technical challenges and communicate effectively across research, engineering, and product teams.

Nice To Haves

  • Hands-on experience testing or evaluating modern AI systems such as LLMs, RAG pipelines, agents, or multimodal models, including analyzing model behavior and failure modes.
  • Experience with voice, audio, speech recognition, or real-time systems, and familiarity with metrics such as WER, MOS, latency, and time-to-first-byte.
  • Experience building or improving test, evaluation, benchmarking, or ML infrastructure used by multiple teams or external users.
  • A strong appreciation for test and evaluation quality, including correctness, reproducibility, determinism, and consistency across environments.
  • Experience building test tooling for React Native, mobile applications, or other cross-platform environments that extends validation beyond the desktop.
  • Familiarity with cloud infrastructure, containers, ephemeral test environments, CI/CD systems, and monitoring tools such as Grafana, canaries, and anomaly detection.
  • Experience serving as a technical bridge across teams or platforms—including product, QA, evaluation, training, inference, data, or agent frameworks—with the communication skills to build alignment and influence decisions.
  • Prior involvement in open-source projects through contributions, reviews, maintenance, or community engagement.
  • Experience with voice, audio, speech recognition, or real-time systems, and familiarity with metrics like WER, MOS, or latency/TTFB.
  • Prior involvement in open-source projects, through contributions, reviews, maintenance, or community engagement.
  • Experience acting as a technical bridge across teams or platforms (evaluation, training, inference, agent frameworks), combining architectural understanding with clear communication and influence.
  • Familiarity with cloud infrastructure, containerized/ephemeral environments, and monitoring tooling (e.g. Grafana, canaries, anomaly detection).

Responsibilities

  • Define and execute well-designed test plans across Deepgram's products, APIs, SDKs, model-powered features, and data platforms, ensuring production software is robust, reliable, and performs well.
  • Design, build, and maintain automated test suites and frameworks for functional, integration, end-to-end, regression, API, browser, and service-level testing across batch and streaming workflows.
  • Translate product requirements and customer acceptance criteria into clear test strategies, repeatable test cases, and enforceable release gates.
  • Build and maintain representative, customer-focused, and adversarial test datasets, fixtures, and test environments that exercise real-world inputs, edge cases, failure modes, and system limits.
  • Validate model-powered behavior—including speech-to-text, text-to-speech, and other AI features—using appropriate metrics, expected outputs, human review, and regression coverage, while partnering with Research and model-evaluation specialists as needed.
  • Build testing infrastructure, including test harnesses, reusable scripts, test-data tooling, result-aggregation pipelines, dashboards, and visualizations that make quality signals easy to understand and act on.
  • Integrate automated tests, quality checks, canaries, and release validation into CI/CD so regressions are detected continuously rather than through manual testing alone.
  • Partner with Engineering, Product, Research, Data, Infrastructure, and DevOps to understand system behavior, dependencies, variations, performance limits, and deployment risks, and to establish appropriate test coverage.
  • Test data ingestion, processing, annotation, and quality-control workflows, validating data integrity, completeness, representativeness, deduplication, leakage, and downstream readiness.
  • Execute staging and production validation, load and reliability testing, cross-browser and customer-workflow testing, and user acceptance testing in partnership with internal stakeholders and customer QA teams.
  • Maintain and improve the test-case repository, automation coverage, test documentation, and release-readiness reporting so teams have a clear view of what was tested, what passed, and what remains risky.
  • Write precise, actionable bug reports with reproducible steps, inputs, parameters, expected and actual results, logs or artifacts, and clear severity; participate in triage and escalate issues when necessary.
  • Help raise the bar through code reviews, test-design reviews, technical discussions, and strong engineering, automation, and QA practices.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service