Hedra is the platform, models, and infrastructure for visual intelligence. We build models and systems that push the frontier of visual intelligence, along with the infrastructure required to make those models fast, efficient, reliable, and accessible at scale. We’re a team in San Francisco, backed by a16z and other leading investors. Researchers and engineers at Hedra work closely across boundaries, own problems end to end, and have significant influence over both what we build and how we build it. We’re looking for an Inference Optimization Engineer to work alongside our research team on making state-of-the-art visual models fast and efficient at inference time. You’ll work at the boundary between research and systems, taking new model architectures and figuring out how to run them efficiently on modern hardware. That means understanding where time and memory are being spent, identifying opportunities for algorithmic and systems-level improvements, and implementing optimizations across model architecture, inference algorithms, runtimes, kernels, and distributed execution. The problems rarely live neatly within one layer of the stack. Depending on what you find, you might modify how a model executes, develop a new inference technique, write a custom GPU kernel, rethink memory movement, or change how work is distributed across accelerators. We care more about technical depth, curiosity, and demonstrated ability than years of experience. We’re open to experienced ML systems engineers as well as exceptional early-career engineers or researchers who have already gone unusually deep on model performance, GPU systems, or efficient inference.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Entry Level
Education Level
No Education Listed