Data Scientist

SonatypeUS - Remote, GA

About The Position

Sonatype is a leader in the software supply chain management industry, known for inventing componentized software development and pioneering the software supply chain category. They operate the world's largest repository of Java open-source components, Maven Central. Their platform helps organizations rapidly create, deploy, and maintain innovative software at scale, ensuring business alignment. Sonatype is trusted by over 2,000 organizations, including 70% of the Fortune 100, and serves over 15 million software developers. Their tools and guidance focus on delivering exceptional, secure software. They utilize AI/ML to provide clients, developers, and the industry with confidence in software quality, automation, and security. The company is committed to constant innovation, as demonstrated by their Nexus Repository for artifact management and their solutions for halting malicious open-source malware. This role is for a Data Scientist to join the AI & Data Science team, acting as an internal AI consultant and technical lead. The position involves helping multiple Sonatype teams apply machine learning and generative AI to real-world problems, including malicious-behavior and anomaly detection in security data, and developing GenAI experiences for developers and analysts. The role requires exploring complex datasets, designing experiments, building and validating models, and collaborating with product, engineering, and security experts to transform research ideas into practical, scalable solutions. The company has a mature data engineering team, allowing the Data Scientist to focus on model development and deployment. This role is suited for individuals who value autonomy, enjoy translating ambiguous ideas into working systems, and prefer working across different product areas.

Requirements

  • 5+ years of hands-on experience in applied data science, machine learning, AI engineering, or AI research.
  • Computer Science or equivalent technical degree strongly preferred
  • Strong Python skills and practical experience with data and AI libraries/platforms such as Databricks, and LLM APIs, scikit-learn
  • Experience building and shipping ML or GenAI applications—from early prototype through usable internal or customer-facing workflows.
  • Deep familiarity with modern LLM ecosystems, including OpenAI, Anthropic/Claude, Hugging Face, and open-weight models.
  • Ability to select models and design effective LLM applications using prompting, context management, structured outputs, retrieval, and tool use.
  • Experience building agentic or multi-step AI workflows with LangGraph, LangChain, Semantic Kernel, or similar orchestration frameworks.
  • Strong evaluation mindset: defining useful quality metrics, building representative evaluation datasets, assessing reliability, and making data-driven tradeoffs.
  • Comfortable working with large, messy, structured, and unstructured data to produce features, insights, and clear visualizations.
  • Proficiency with Git, testing, code review, and collaborative software-development practices.
  • Practical, balanced judgment: comfortable exploring emerging AI capabilities while building maintainable, secure, dependable systems.
  • Proactive and accountable, with strong written and verbal communication skills across technical and non-technical partners.

Nice To Haves

  • Strong MLOps experience, including MLflow or comparable tooling, experiment tracking, reproducible pipelines, model/application versioning, CI/CD, serving, and production monitoring.
  • Experience operating ML or GenAI systems at scale, including observability, tracing, incident response, and data or model-drift detection.
  • Experience with Databricks ML, AWS SageMaker, Azure ML, or similar managed ML platforms.
  • Familiarity with MCP, agent-tool integrations, LLM guardrails, and production safety practices.
  • Experience with AI-assisted development tools such as Copilot, Claude Code, or Codex.
  • Exposure to cybersecurity, fraud detection, anomaly detection, code analysis, or software supply-chain security.
  • Experience with PySpark and production data pipelines.
  • Experience working within a software product company or SaaS.

Responsibilities

  • Lead applied AI projects from concept to impact — prototype, validate, and help teams deploy practical ML and GenAI solutions.
  • Act as an internal consultant across product, engineering, security, and research teams: scope problems, evaluate approaches, and advise on ML/AI best practices and productive use of generative technologies.
  • Lead the research, development, and deployment of models for use cases such as malicious behavior detection, anomaly detection, and fraud analysis — using techniques ranging from classical ML to LLMs, embeddings, retrieval-augmented generation, and agentic workflows.
  • Design robust experiments and establish evaluation pipelines for model reliability, accuracy, and business impact (cross-validation, drift monitoring, ground-truth evaluation).
  • Bridge research and production: translate research insights into scalable APIs, tools, or workflows that enable other teams to adopt AI effectively.
  • Explore new techniques (LLMs, embeddings models, RAG, agentic workflows) to enhance developer and security experiences.
  • Communicate technical concepts, tradeoffs, and recommendations clearly to both technical and non-technical stakeholders through presentations, documentation, and collaboration; mentor peers and help elevate the organization's AI literacy and capabilities.
  • Partner with our data governance team to ensure compliance with data-privacy regulations and ethical considerations when working with customer data.

Benefits

  • Parental Leave Policy
  • Paid Volunteer Time Off (VTO)
  • Diversity and Inclusion Working Groups
  • Flexible working practices
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service