Software Engineer, Inference AI/ML

Core Weave•Bellevue, WA

48d•$92,000 - $135,000•Hybrid

About The Position

CoreWeave is the AI Hyperscaler, delivering a cloud platform of cutting edge services powering the next wave of AI. Our technology provides enterprises and leading AI labs with the most performant, efficient and resilient solutions for accelerated computing. Since 2017, CoreWeave has operated a growing footprint of data centers covering every region of the US and across Europe. CoreWeave was ranked as one of the TIME100 most influential companies of 2024. As the leader in the industry, we thrive in an environment where adaptability and resilience are key. Our culture offers career-defining opportunities for those who excel amid change and challenge. If you're someone who thrives in a dynamic environment, enjoys solving complex problems, and is eager to make a significant impact, CoreWeave is the place for you. Join us, and be part of a team solving some of the most exciting challenges in the industry. CoreWeave powers the creation and delivery of the intelligence that drives innovation. What You'll Do: Join the Inference team to ship production features that improve latency, reliability, and cost for model serving on our GPU platform. As an IC1, you'll implement well-scoped changes, learn our operational practices, and grow quickly with mentorship from experienced engineers.

Requirements

BS/MS in CS, EE, or related field, or equivalent practical experience.
Foundations in data structures, algorithms, and networked services.
Experience with Python or Go (C++ a plus) and Linux fundamentals; Git/CI basics.
Exposure to containers and Kubernetes (coursework or projects welcome).

Nice To Haves

Internship or project that deployed a microservice or ML inference demo.
Coursework/research with PyTorch or TensorFlow; simple CUDA projects a plus.
Familiarity with Grafana/Prometheus/OpenTelemetry or similar tooling.
Curiosity about GPU inference concepts (micro-batching, KV cache, streaming).

Responsibilities

Implement well-scoped features and fixes in Python/Go/C++ for model-serving services (e.g., Triton, vLLM, TensorRT-LLM, Ray Serve).
Write tests, code comments, and short design docs; participate in code reviews.
Add basic metrics and dashboards; assist with alarms and runbooks.
Follow on-call runbooks and learn incident response in a guided rotation.
Contribute to performance experiments (e.g., request batching, concurrency, caching) with guidance.