Staff Software Engineer AI/ML

Samsung SemiconductorSan Jose, CA
$163,000 - $253,000

About The Position

We are seeking a Senior Staff Engineer to build and optimize the EDA design environment and large-scale compute infrastructure that supports our semiconductor design organizations, and to establish company-wide standards and processes for EDA licensing and R&D software. In this role, you will design and standardize the shared design environments and infrastructure used across multiple design organizations, driving engineering productivity and cost efficiency at scale. AI/PI Group is an internal AI and Process Innovation organization within Samsung DSA, dedicated to transforming how our company works through artificial intelligence. We accelerate AI adoption across the entire organization — from reshaping day-to-day workflows and automating core internal processes to empowering our workforce with practical, intelligent tools. The Applied AI Engineering team is the driving force behind this vision, leading innovation at the intersection of machine learning and system engineering to develop and operate our next-generation AI frameworks. We aim to enable frontier AI models to autonomously plan, retrieve information, coordinate with tools, and execute multi-step workflows across our internal knowledge ecosystem. We are actively seeking talented Machine Learning Engineers specializing in building next-generation AI/ML solutions.

Requirements

  • BS with 10+ years, MS with 8+ years, or PhD with 5+ years of experience in Computer Science, Electrical Engineering, or a related field
  • Demonstrated track record of deploying and operating LLM- or vision-powered systems in production, with proven experience in evaluation and safe rollout practices (A/B testing, canary releases, offline/online evaluation) — including defining quality metrics, benchmarking models, and driving quality/cost/latency trade-off decisions
  • Hands-on expertise with agentic AI frameworks (LangGraph, CrewAI, ADK) — multi-agent orchestration and patterns, tool routing, state management, and/or interoperability protocols (MCP, A2A)
  • Experience with LLM inference optimization and serving (e.g., vLLM, TensorRT-LLM), covering quantization, KV-cache management, and hardware-aware deployment
  • Strong Python and C/C++ skills with deep experience in PyTorch/TensorFlow, and experience designing, building, and securing large-scale distributed systems

Nice To Haves

  • PhD of software engineering experience
  • MLOps exposure: prompt/version management, monitoring, observability
  • Experience with ML, graphics or computer vision accelerator
  • Understanding of PPA (performance, power, and area) trade-offs, memory controller architecture and/or general computer architecture is beneficial
  • Familiarity with state-of-the-art AI workloads and their compute and memory requirements
  • Experience with performance modeling of heterogenous systems is beneficial
  • Ability to meet aggressive project deadlines in a team environment

Responsibilities

  • Design, build, and productize agentic AI applications that combine vision and language models — owning the architecture end-to-end from prototype through deployment and ongoing operation.
  • Architect multi-agent systems using frameworks such as LangGraph, LangChain, AutoGen, or CrewAI, including agent orchestration, tool and function routing, state and memory management, error recovery, and human-in-the-loop workflows.
  • Establish evaluation frameworks and quality bars for agentic and multimodal systems: define reference and non-reference metrics, build benchmark suites and regression harnesses, and instrument agent trajectories for offline and online evaluation.
  • Drive quality, cost, and latency trade-off decisions with data — selecting models, routing strategies, and inference configurations that meet product requirements within compute and memory budgets.
  • Optimize inference performance on accelerated hardware through quantization, batching, caching, kernel- and runtime-level tuning, and model/hardware co-design; profile workloads to identify and eliminate bottlenecks.
  • Identify and solve multi-discipline AI acceleration problems, especially with memory bottlenecks, involving algorithms, network design, hardware architecture
  • Work with researchers and application developers to enable the latest machine learning work to optimize performance.

Benefits

  • Medical/Dental/Vision/401k
  • Charitable giving match
  • 4+ weeks of paid time off a year
  • Holidays
  • Sick leave
  • Stipend for fertility care or adoption
  • Medical travel support
  • Virtual vet care
  • On-demand apps for emotional wellness
  • Free confidential therapy sessions
  • Onsite Café
  • Gym
  • Virtual classes
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service