Data Scientist (ML / NLP / LLMs)

Sown To Grow
$120,000 - $135,000Hybrid

About The Position

Sown To Grow (STG) is a K12 education technology platform that empowers schools to improve student social, emotional, and academic health through an easy and engaging check-in and reflection process. In a short weekly routine, students check in on how they are feeling and reflect on the strategies that are working best for them (or new ones to try), and teachers respond with support and coaching. School leaders and student support staff use real-time reporting on students' emotions and reflections to proactively intervene to support student needs. The system also includes built-in screeners, supporting curriculum, and powerful artifacts of growth. At the heart of STG is a natural language and machine learning system that turns unstructured, deeply human reflection text into structured, trustworthy insights - reading the quality of student reflection and the strength of the student–educator relationship, and surfacing insights that help educators prioritize the kids who need support most. Built on research-validated models and extended with modern AI, our work serves one guiding philosophy: make this routine lasting and sustainable for all users by making it more engaging, meaningful, and efficient. We've spent years turning these insights into real-time, user-facing features and hardening the technical backbone to serve millions of students at low latency. Now we're pushing that foundation further with the latest in AI - building richer, more contextual, AI- and expert-guided support that helps teachers respond to students more effectively, and continually expanding the range of problems we take on. We're looking for a Data Scientist who can move fluently across this whole surface - from feature engineering and model validation to production ML systems, and from classical NLP into the responsible application of modern LLMs. You'll join a small data science / machine learning team - small enough that you'll know everyone's name and see your work ship in weeks, not quarters - working closely with product and engineering. We're deliberate about our toolkit: classical, feature-driven ML is the validated core of what we do today and isn't going anywhere, while modern LLMs open new frontiers we're actively investing in. We don't reach for the trendiest technique or cling to the familiar one - we choose the right tool for each problem, and we're looking for someone who enjoys making that call with us. Your work will directly shape how educators understand and support students, at a scale that reaches historically underserved communities.

Requirements

  • Bachelor's or higher degree in Computer Science, Data Science, Machine Learning, Math, Statistics, or a related field.
  • 2+ years of experience as a Data Scientist, ML Engineer, or Data Engineer, solving real-world problems with machine learning.
  • Strong proficiency in Python and the ML stack (pandas, numpy, scikit-learn; PyTorch or TensorFlow; Spark a plus).
  • Experience building and deploying ML solutions that involve natural language processing of text data.
  • Working knowledge of core ML techniques such as classification, clustering, prediction, recommender systems, and anomaly detection.
  • Working knowledge of the complete machine learning lifecycle - data, training, validation, deployment, monitoring, and retraining.
  • Solid understanding of how modern LLMs work under the hood - transformer architecture, training and fine-tuning, tokenization, embeddings. We care more about curiosity than credentials here: you enjoy digging into why a model behaves the way it does, not just what it returns. Hands-on with at least one of prompting/evaluation, fine-tuning, retrieval-augmented generation (RAG), or agentic/tool-use patterns, with a thoughtful view of when LLMs are and aren't the right approach.
  • Experience writing and maintaining high-quality production code, and comfort with Git-based workflows.

Nice To Haves

  • Strong interest in working in education technology in an impact-driven, mission-first role.
  • Experience productionizing ML for real-time, low-latency inference (e.g., AWS SageMaker or comparable), including containerization and CI/CD.
  • Experience building data-drift detection, model monitoring, and automated retraining systems in partnership with ML engineering teams.
  • Experience building, training, or fine-tuning language models from the ground up - e.g., implementing transformer components, training or adapting models on domain-specific data, or working with open-weight models beyond off-the-shelf APIs. This is a longer-term direction for us, and we value candidates who can grow into it.
  • Experience with responsible / trustworthy AI: fairness and bias evaluation, privacy-conscious handling of sensitive data, and building guardrails for user-facing generative features.
  • Experience designing human-in-the-loop evaluation and running online experiments (A/B testing, feature flagging).
  • Familiarity with the practical, ethical, and legal considerations of working with student data.

Responsibilities

  • Build and refine NLP/ML models over large volumes of unstructured student and educator text, using both classical approaches (feature engineering, classification, CNNs, ensemble methods) and modern LLM-based techniques where they're the right tool.
  • Partner with data science, product, and engineering to identify, define, and test opportunities to improve the product through ML/NLP - from turning existing models into user-facing features to prototyping entirely new capabilities.
  • Extend our models beyond "proof of concept" scope: new reflection prompts, younger students etc, widening accessibility while protecting accuracy.
  • Design and apply LLMs responsibly for generative and assistive features - for example, contextualized teacher-response suggestions, and resource recommendations - with careful attention to prompting, retrieval, grounding, evaluation, and guardrails.
  • Build validation, monitoring, and retraining pipelines in partnership with the ML engineering team - including data-drift detection - so models keep performing as code and data change.
  • Undertake preprocessing of structured and unstructured data, and build reliable, reproducible feature and evaluation workflows.

Benefits

  • Competitive compensation with performance-based incentives and meaningful equity
  • Comprehensive health and wellness benefits for you and your family
  • Flexible work arrangements and a genuine commitment to work-life balance
  • Real pathways for growth - as the platform and the data team expand, so does the scope of this role
  • A collaborative, mission-driven community where every voice is heard
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service