About The Position

We are looking for a senior data scientist to streamline marketing operations at team.blue, by building agentic systems to run them. You would report into the Applied AI team and work on marketing automation projects. Marketing here runs across many brands, markets and languages, on a stack that differs brand by brand. The work spans competitive and pricing monitoring, performance reporting and diagnosis, SEO and AI-answer visibility, content refresh, localisation and lifecycle production, paid search and social account hygiene, and tracking and consent QA. Each of these is a multi-step process across several systems, repeated per brand. The method you will follow matters more than the domain: map processes, quantify the time and resources they consume, determine the ROI impact of agentic automation, build a proof of concept, take it to production and measure the impact of your work. This is closer to building autonomous, business-impact systems than to building pure single purpose models.

Requirements

  • 7+ years building data and ML systems in industry, spanning both sides of the LLM shift.
  • Experience debugging systems before you could ask a model what was wrong.
  • Experience being the only person who did a job end to end (First or only data hire, the single ML person in a small company, or a one-person function inside a large one).
  • Shipped something with permission to act on live systems affecting real customers.
  • Expert in Python and ML.
  • Ability to write Python code that someone else can still read in six months.
  • Proficiency with a current toolchain (uv, Docker or an equivalent).
  • Experience building your own container and your own instrumentation.
  • Production experience with multi-step, tool-calling LLM workflows: orchestration, retries, idempotency, timeouts, partial-failure recovery.
  • Experience with State-machine design, not only train/serve pipelines.
  • Experience with integration against third-party APIs you do not control.
  • Experience with Auth flows, rate limits, pagination, sandbox behaviour that differs from production, and schema changes shipped without notice.
  • Experience with Cost and latency engineering as a first-class concern: model routing, caching, batching, and the instinct to know what a flow costs per run before Finance asks.
  • A safety instinct for systems that take actions: staging modes, approval gates, least-privilege scoping, a way back.
  • Experience with Evaluation design for generative and agentic output: LLM-as-judge, golden-transcript regression suites, red-teaming.
  • Experience with calibration: an agent that reports high confidence needs to be right at that rate, and you can show whether it is.
  • Process mapping and quantification: ability to sit with a domain expert, capture what actually happens rather than what the policy says, and attach hours to it.
  • Technical vendor evaluation: judging a martech vendor on API surface, data model, extensibility and true integration cost, not on the sales deck.

Nice To Haves

  • Master’s or PhD in Computer Science, AI, Machine Learning or a related field.
  • PromptOps at scale: versioning, testing and rollback of prompts and models as production artefacts.
  • Prior exposure to martech, ad-tech or SEO tooling and their APIs, or to automation in any domain where output is customer-facing.
  • Experience evaluating generated output across multiple languages.

Responsibilities

  • Map processes, quantify the time and resources they consume, and determine the ROI impact of agentic automation.
  • Build a proof of concept for agentic automation systems.
  • Take agentic automation systems to production and measure their impact.
  • Sit with an SEO or paid search owner and map how a traffic-drop investigation actually runs today across brands, then attach hours per week to each step of it.
  • Extract requirements live from people who do not think in data models.
  • Design the state transitions: what triggers, what branches, which APIs get called, where it waits for a human, and what happens when step 4 of 9 fails or a vendor rate-limits you mid-run.
  • Build the guardrails before the capability: dry-run mode, an approval gate ahead of anything that writes to a live account or publishes externally, least-privilege API scopes, a documented undo.
  • Decide where a human stays in the loop, at what confidence threshold, and design a review queue marketers will open a second time.
  • Write evals for output that precision and recall do not capture: is the diagnosis correct, is the cited source real and does it say what the agent claims, does a generated brief hold brand voice in Greek and Dutch as well as in English.
  • Check whether the tracking data an agent depends on is sound before building on it, and quantify the error when it is not.
  • Wire an agent to a webhook or a scheduled trigger, and make the handler idempotent so a retry does not double-post a recommendation or apply the same keyword exclusion twice.
  • Decide which steps in a flow warrant a frontier model and which run on something cheap, then prove that routing decision with numbers, because these flows run daily across many brands and the bill compounds.
  • Sit in a vendor demo asking what their API actually exposes, what the rate limits and quotas are, what their data model looks like, and what integration really costs us.

Benefits

  • Diversity and inclusion are at our core.
  • Respect, openness, and trusted collaboration are valued.
  • Commitment to caring for the environment and each other.
  • Ongoing ESG efforts and ambitious sustainability goals.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service