Software Engineer

Cisco•Richardson, NC
•Onsite

About The Position

We run the platform that serves foundational models to Cisco IT. Our Foundational Model Service (FMS) gives engineering teams across the company access to small, large language, and embedding models. We serve those models on Kubernetes clusters, with Nim, Vllm and other runtimes. Beyond serving, we benchmark, evaluate, monitor, and release new models as improvements and demand warrant. Our customers depend on the platform under a 99.9% uptime SLA, and we build and operate accordingly. As a Software Engineer on the FMS team, you'll keep that platform running and make it easier to run. You'll support model onboarding and releases, take your turn on call, help automate the validation that gates every deployment, and contribute to the monitoring that shows us and our customers how the service is behaving. You'll also apply Agentic solutions to our own operations to catch problems earlier and remediate them automatically. Your Impact: You'll develop software consistent with Cisco Design Thinking Principles, with simplification and user experience at its core, using secure coding practices, protecting user privacy, and following software development best practices. You'll partner with design, product management, and other engineering teams to build the right solution for our customers. You'll create technical design documentation for the team, contribute to the documentation end users rely on, and debug platform issues both in development and in production. This is a strong role for a systems or infrastructure engineer who has worked with LLMs, whether through academic coursework or by building and running them on the job, and who understands how inference serving works. We're looking for a self-starter who takes on unfamiliar work and grows into it. You'll work with engineers who have built these systems from the ground up. New runtimes and new models arrive frequently, so the platform is always evolving, and you'll develop a deep understanding of developing AI infrastructure, model runtimes, and delivering products at scale. On call is shared across the team.

Requirements

  • Bachelor’s degree in Computer Science, Information Systems, or a related field with 3 years of related experience; master’s degree with 1 year of related experience; PhD; or equivalent practical work experience.
  • Experience developing and operating production services on Kubernetes and Linux, including exposure to GPU based AI infrastructure.
  • Hands on experience serving models with vLLM, NVIDIA NIM, Triton, or a comparable inference runtime, including building supporting services.
  • Experience leading technical work, mentoring engineers, documenting systems, and supporting critical services through an on call rotation.

Nice To Haves

  • Experience evaluating and benchmarking models to support production release decisions.
  • Experience operating GPU workloads with NVIDIA CUDA or AMD ROCm.
  • Familiarity with distributed inference, disaggregated serving, KV cache aware routing, capacity planning, or cost optimization.
  • Experience fine tuning transformer models or evaluating RAG and agent systems using relevant frameworks.

Responsibilities

  • Keep the platform running and make it easier to run.
  • Support model onboarding and releases.
  • Take turns on call.
  • Help automate the validation that gates every deployment.
  • Contribute to the monitoring that shows us and our customers how the service is behaving.
  • Apply Agentic solutions to our own operations to catch problems earlier and remediate them automatically.
  • Develop software consistent with Cisco Design Thinking Principles, with simplification and user experience at its core, using secure coding practices, protecting user privacy, and following software development best practices.
  • Partner with design, product management, and other engineering teams to build the right solution for our customers.
  • Create technical design documentation for the team.
  • Contribute to the documentation end users rely on.
  • Debug platform issues both in development and in production.

Benefits

  • medical, dental and vision insurance
  • a 401(k) plan with a Cisco matching contribution
  • paid parental leave
  • short and long-term disability coverage
  • basic life insurance
  • grants of Cisco restricted stock units
  • 10 paid holidays per full calendar year
  • plus 1 floating holiday for non-exempt employees
  • 1 paid day off for employee’s birthday
  • paid year-end holiday shutdown
  • 4 paid days off for personal wellness
  • 16 days of paid vacation time per full calendar year (for non-exempt employees)
  • flexible vacation time off program (for exempt employees)
  • 80 hours of sick time off provided on hire date and each January 1st thereafter
  • up to 80 hours of unused sick time carried forward from one calendar year to the next
  • Additional paid time away may be requested to deal with critical or emergency issues for family members
  • Optional 10 paid days per full calendar year to volunteer
  • annual bonuses (for non-sales roles)
  • performance-based incentive pay (for sales roles)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service