About The Position

The role Own the cost and AI-leverage layer of the analytics platform behind Capital One Shopping — the systems between petabyte-scale data and the humans and tools that query it. This is an own-and-build role, not a maintenance seat : you own the production platforms below, and in your first six months you ship three net-new systems on top of them. You'll report to the Engineering Director for the Shopping data platform as one of two senior IC pillars of the analytics org. What you own You own, support, and evolve three production platforms — ideation through implementation to production support — and you're the SME and mentor for the analysts, BAs, and engineers who use them: A data warehousing platform serving ~250K queries/day over ~20PB . An event ingestion pipeline taking in 6–7 billion events/day . The Airflow orchestration platform . You own the technology choices and the strategic backlog and priorities for this surface — and you carry ongoing production support and an on-call rotation for these platforms and the models on them. You build the three net-new systems below on top of that operational base. What you'll build A cost-attribution pipeline that parses Trino query logs at production traffic and attributes real AWS dollars to every report, query, user, and dbt model — reconciled against a ~$750K/month cloud bill. A forecast-driven autoscaling control loop for the shared Trino cluster and dbt worker pool — turning today's event-only Nomad autoscaler (Prime Day, Cyber Week) into steady-state, forecast-driven capacity. The single largest lever on the analytics AWS bill. A production Gen AI system — natural-language-to-SQL or RAG over the data catalog — with real LLM tool-use, grounding, and cost guardrails, adopted by internal teams. The stack Kafka streaming backbone into an S3 lakehouse (Hive + Iceberg), Cassandra, Postgres, DynamoDB, ElasticSearch, Aurora MySQL. Queried through Trino/Presto and Spark SQL, modeled in dbt, orchestrated on Airflow and Nomad and containers (Docker/Kubernetes), on a deep AWS footprint. SQL and Python daily; Go, Java, and TypeScript/JavaScript across the surrounding platform. The day-to-day "Manager" is the level, not the job — this is an individual-contributor role, and you'll spend most of your day hands-on in the editor. Roughly 70% building : writing the log-parsing and cost-attribution logic and its dbt models, building and tuning the forecast-driven autoscaler control loop, and building the RAG / natural-language-to-SQL system yourself. The other ~30% is technical coordination — reconciling your cost numbers with Finance, the R&D memo, aligning report owners — not status decks or people-management. Daily rhythm is multi-terminal Claude Code: query-log analysis in one, dbt work in another, AI iteration in a third. No direct reports — you build.

Requirements

  • 10+ years engineering experience owning and supporting mission-critical applications and platforms in production.
  • Deep experience with Kafka and streaming technologies and platforms.
  • Experience with enterprise data technologies and platforms.
  • Track record working with extremely large traffic and data volumes.
  • Fluency across a real stack: JavaScript, Java, HTML/CSS, TypeScript, SQL, Python, and Go, open-source RDBMS and NoSQL databases, container orchestration (Docker and Kubernetes), and a broad range of AWS tools and services.
  • Experience with forecast- or workload-driven infrastructure scaling on a shared platform.
  • Experience with Production Gen AI: RAG design, vector/graph stores, LLM tool-use.
  • Experience with a longer ML/AI arc (ranking, recommendation, or comparable production ML predating the 2023 Gen AI boom).
  • Experience with 0-to-1 delivery of a product that booked measurable revenue or adoption in its first weeks.
  • Experience as a lead-engineer modernizing an ingest pipeline off a legacy provider onto a cloud-native stack.
  • Experience directing a cross-team migration or deprecation at senior-executive scope to completion.
  • Comfortable working with large teams on large-scale systems, and cross-functionally (weekly Finance/BizOps and Analytics cadences).
  • Claude Code fluency — daily use, skill authoring, PR-level deliverables (hard requirement).
  • Bachelor’s Degree.
  • At least 6 years of experience in software engineering (Internship experience does not apply).
  • At least 1 year experience with cloud computing (AWS, Microsoft Azure, Google Cloud).

Nice To Haves

  • Master’s Degree.
  • 9+ years of experience in at least one of the following: JavaScript, Java, TypeScript, SQL, Python, or Go.
  • 4+ years of experience with AWS, GCP, Microsoft Azure, or another cloud service.
  • 4+ years of experience in open source frameworks.
  • 1+ years of people management experience.
  • 2+ years of experience in Agile practices.

Responsibilities

  • Own, support, and evolve three production platforms: a data warehousing platform serving ~250K queries/day over ~20PB, an event ingestion pipeline taking in 6–7 billion events/day, and the Airflow orchestration platform.
  • Serve as the SME and mentor for analysts, BAs, and engineers using these platforms.
  • Own technology choices, strategic backlog, and priorities for the analytics platform surface.
  • Carry ongoing production support and an on-call rotation for these platforms and their models.
  • Build a cost-attribution pipeline that parses Trino query logs and attributes AWS dollars to reports, queries, users, and dbt models, reconciled against a ~$750K/month cloud bill.
  • Build a forecast-driven autoscaling control loop for the shared Trino cluster and dbt worker pool.
  • Build a production Gen AI system (natural-language-to-SQL or RAG over the data catalog) with LLM tool-use, grounding, and cost guardrails.
  • Write log-parsing and cost-attribution logic and its dbt models.
  • Build and tune the forecast-driven autoscaler control loop.
  • Build the RAG / natural-language-to-SQL system.
  • Perform technical coordination, including reconciling cost numbers with Finance, R&D memo, and aligning report owners.

Benefits

  • Comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service