Staff Machine Learning Engineer, AI Insights

CoreWeaveNew York, NY
$188,000 - $250,000

About The Position

As a Staff Machine Learning Engineer on the AI Insights team, you will be a technical leader responsible for defining and implementing machine learning systems that power anomaly detection, signal correlation, incident understanding, recommendations, and other intelligence capabilities across CoreWeave’s observability and cloud platforms. You will manage the full lifecycle of these systems, from problem framing and model development to evaluation, data and inference service building, production integration, and scaled operation. This role involves close collaboration with software engineers, product managers, researchers, and domain experts to transform innovative ideas into reliable customer experiences. The focus is on building underlying ML systems, services, evaluation frameworks, and product capabilities for reliable AI-powered troubleshooting and optimization in real-world environments, rather than generic chatbot development.

Requirements

  • Significant experience designing and shipping machine learning systems that operate in production.
  • Strong software engineering skills in Python; experience with Go or another systems-oriented language is a plus.
  • Deep understanding of machine learning fundamentals, including model selection, feature engineering, experimentation, evaluation, and failure analysis.
  • Experience working with time-series, event, log, metric, trace, or other operational data at meaningful scale.
  • Demonstrated ability to define evaluation methodology for ambiguous or domain-specific ML problems.
  • Strong systems thinking and the ability to reason about distributed systems, data quality, latency, observability, and operational failure modes.
  • Excellent communication and collaboration skills, including the ability to work effectively with research, product, infrastructure, and customer-facing teams.
  • A track record of technical leadership, mentorship, and influence across organizational boundaries.

Nice To Haves

  • Experience with observability platforms or technologies such as Grafana, Prometheus, VictoriaMetrics, ClickHouse, Loki, or Kafka.
  • Experience with Kubernetes and cloud infrastructure, especially for telemetry, logging, or application observability.
  • Experience with anomaly detection, incident intelligence, search, recommendations, or ranking systems.
  • Experience with large language model evaluation, post-training, retrieval, tool use, or grounded generation—particularly when combined with structured telemetry and deterministic systems.
  • Experience building ML products for infrastructure, developer tools, reliability engineering, or other technical users.
  • Familiarity with human-in-the-loop workflows, access control, auditability, and safety requirements for operational systems.

Responsibilities

  • Lead the technical design and delivery of production machine learning systems for infrastructure observability, troubleshooting, and optimization.
  • Develop approaches for anomaly detection, time-series analysis, event correlation, ranking, recommendation, classification, and root-cause inference across high-volume telemetry.
  • Build evaluation frameworks and datasets that measure accuracy, usefulness, robustness, and safety in real operational scenarios.
  • Translate research and prototypes into maintainable, scalable services with clear operational ownership.
  • Design data pipelines, feature-generation workflows, model-serving paths, and feedback loops for continuous improvement.
  • Integrate ML capabilities with telemetry platforms and customer-facing experiences, including Mission Control, Grafana, and related observability services.
  • Establish practical standards for experimentation, offline and online evaluation, monitoring, reproducibility, and model lifecycle management.
  • Make thoughtful trade-offs across model quality, latency, cost, interpretability, reliability, and ease of operation.
  • Work with teams responsible for metrics, logs, traces, telemetry enrichment, and platform APIs to create coherent cross-system intelligence.
  • Mentor engineers and raise the technical bar for machine learning and production engineering across the organization.
  • Communicate technical decisions clearly and influence roadmaps across teams without relying on formal authority.

Benefits

  • Medical, dental, and vision insurance - 100% paid for by CoreWeave
  • Company-paid Life Insurance
  • Voluntary supplemental life insurance
  • Short and long-term disability insurance
  • Flexible Spending Account
  • Health Savings Account
  • Tuition Reimbursement
  • Ability to Participate in Employee Stock Purchase Program (ESPP)
  • Mental Wellness Benefits through Spring Health
  • Family-Forming support provided by Carrot
  • Paid Parental Leave
  • Flexible, full-service childcare support with Kinside
  • 401(k) with a generous employer match
  • Flexible PTO
  • Catered lunch each day in our office and data center locations
  • A casual work environment
  • A work culture focused on innovative disruption
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service