Staff Software Engineer AI/ML

Samsung SemiconductorSan Jose, CA
Onsite

About The Position

Our technology solutions power the tools you use every day--including smartphones, electric vehicles, hyperscale data centers, IoT devices, and so much more. Here, you’ll have an opportunity to be part of a global leader whose innovative designs are pushing the boundaries of what’s possible and powering the future. We believe innovation and growth are driven by an inclusive culture and a diverse workforce. We’re dedicated to empowering people to be their true selves. Together, we’re building a better tomorrow for our employees, customers, partners, and communities. About the Role We are seeking a Senior Staff Engineer to build and optimize the EDA design environment and large-scale compute infrastructure that supports our semiconductor design organizations, and to establish company-wide standards and processes for EDA licensing and R&D software. In this role, you will design and standardize the shared design environments and infrastructure used across multiple design organizations, driving engineering productivity and cost efficiency at scale. AI/PI Group is an internal AI and Process Innovation organization within Samsung DSA, dedicated to transforming how our company works through artificial intelligence. We accelerate AI adoption across the entire organization — from reshaping day-to-day workflows and automating core internal processes to empowering our workforce with practical, intelligent tools. The Applied AI Engineering team is the driving force behind this vision, leading innovation at the intersection of machine learning and system engineering to develop and operate our next-generation AI frameworks. We aim to enable frontier AI models to autonomously plan, retrieve information, coordinate with tools, and execute multi-step workflows across our internal knowledge ecosystem. We are actively seeking talented Machine Learning Engineers specializing in building next-generation AI/ML solutions.

Requirements

  • BS with 10+ years, MS with 8+ years, or PhD with 5+ years of experience in Computer Science, Electrical Engineering, or a related field
  • Demonstrated track record of deploying and operating LLM- or vision-powered systems in production, with proven experience in evaluation and safe rollout practices (A/B testing, canary releases, offline/online evaluation) — including defining quality metrics, benchmarking models, and driving quality/cost/latency trade-off decisions
  • Hands-on expertise with agentic AI frameworks (LangGraph, CrewAI, ADK) — multi-agent orchestration and patterns, tool routing, state management, and/or interoperability protocols (MCP, A2A)
  • Experience with LLM inference optimization and serving (e.g., vLLM, TensorRT-LLM), covering quantization, KV-cache management, and hardware-aware deployment
  • Strong Python and C/C++ skills with deep experience in PyTorch/TensorFlow, and experience designing, building, and securing large-scale distributed systems

Nice To Haves

  • PhD of software engineering experience
  • MLOps exposure: prompt/version management, monitoring, observability
  • Experience with ML, graphics or computer vision accelerator
  • Understanding of PPA (performance, power, and area) trade-offs, memory controller architecture and/or general computer architecture is beneficial
  • Familiarity with state-of-the-art AI workloads and their compute and memory requirements
  • Experience with performance modeling of heterogenous systems is beneficial

Responsibilities

  • Design, build, and productize agentic AI applications that combine vision and language models — owning the architecture end-to-end from prototype through deployment and ongoing operation.
  • Architect multi-agent systems using frameworks such as LangGraph, LangChain, AutoGen, or CrewAI, including agent orchestration, tool and function routing, state and memory management, error recovery, and human-in-the-loop workflows.
  • Establish evaluation frameworks and quality bars for agentic and multimodal systems: define reference and non-reference metrics, build benchmark suites and regression harnesses, and instrument agent trajectories for offline and online evaluation.
  • Drive quality, cost, and latency trade-off decisions with data — selecting models, routing strategies, and inference configurations that meet product requirements within compute and memory budgets.
  • Optimize inference performance on accelerated hardware through quantization, batching, caching, kernel- and runtime-level tuning, and model/hardware co-design; profile workloads to identify and eliminate bottlenecks.
  • Identify and solve multi-discipline AI acceleration problems, especially with memory bottlenecks, involving algorithms, network design, hardware architecture
  • Work with researchers and application developers to enable the latest machine learning work to optimize performance.

Benefits

  • Medical/Dental/Vision/401k
  • Charitable giving match
  • 4+ weeks of paid time off a year, plus holidays and sick leave
  • Stipend for fertility care or adoption
  • Medical travel support
  • Virtual vet care
  • On-demand apps and free confidential therapy sessions
  • Onsite Café and gym, plus virtual classes
  • Flexible environment
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service