As a Software Engineer on the Foundational Model Service (FMS) team, you will be responsible for maintaining and improving the platform that serves foundational models (small, large language, and embedding models) to engineering teams across Cisco. This involves serving models on Kubernetes clusters using runtimes like Nim and Vllm, as well as benchmarking, evaluating, and monitoring models for releases. The platform operates under a 99.9% uptime SLA, requiring a focus on reliability and operational excellence. Your role will include supporting model onboarding and releases, participating in an on-call rotation, automating validation processes, and enhancing monitoring capabilities. Additionally, you will apply Agentic solutions to proactively identify and resolve operational issues. You will develop software aligned with Cisco Design Thinking Principles, emphasizing simplification and user experience, while adhering to secure coding practices, protecting user privacy, and following software development best practices. Collaboration with design, product management, and other engineering teams will be key to building effective customer solutions. You will be responsible for creating technical design documentation, contributing to end-user documentation, and debugging platform issues in both development and production environments. This role is ideal for a systems or infrastructure engineer with experience in LLMs, understanding of inference serving, and a proactive approach to learning and problem-solving. You will work alongside experienced engineers and gain deep insights into AI infrastructure development, model runtimes, and large-scale product delivery. On-call duties are shared among the team.
Stand Out From the Crowd
Upload your resume and get instant feedback on how well it matches this job.
Job Type
Full-time
Career Level
Mid Level