About The Position

This is a Lead Infrastructure Production Management & Reliability Engineering position at the Vice President level, which is part of the job family responsible for maintaining the stability and reliability of the organization's infrastructure systems, ensuring optimal performance and availability to support business operations. Since 1935, Morgan Stanley is known as a global leader in financial services, continuously evolving and innovating to better serve our clients and our communities in more than 40 countries around the world. The ideal candidate will have extensive experience planning and executing critical network security and infrastructure patch rollouts in fast-paced, highly regulated, and uptime-sensitive environments. The role requires deep expertise in routing, switching, data center networking, internet connectivity, network security, and production change management. The successful candidate will demonstrate a strong track record of executing complex infrastructure changes with minimal risk, maintaining network resiliency, and supporting business-critical services operating in 24x7 global environments.

Requirements

  • 8+ years of enterprise network engineering experience.
  • Strong hands-on experience with: Cloud observability skill set and cloud native tools like CloudWatch, APIs, VPC reachability analyzer, Network Watcher etc.
  • Leaf-Spine Architectures across Cisco and Arista Platforms
  • Ansible, Python, Terraform, Github and Jira expertise along with network knowledge on various protocols like BGP/OSPF etc.
  • Proven experience executing critical security patching and production network upgrades in large-scale environments.
  • Strong experience supporting both internal infrastructure and internet-facing production services.
  • Demonstrated ability to perform complex maintenance and migration activities during scheduled change windows.
  • Strong understanding of network resiliency, high availability, and disaster recovery principles.

Nice To Haves

  • CCNP, CCIE, or equivalent industry certifications.
  • Knowledge of network automation frameworks.
  • Experience with large-scale global data center environments.
  • Familiarity with agile delivery and DevOps methodologies.
  • Financial services domain knowledge is a plus
  • Strong ownership and accountability.
  • Calm and decisive under pressure.
  • Excellent troubleshooting and analytical skills.
  • Strong communication and stakeholder management capabilities.
  • Ability to operate effectively in fast-paced, highly dynamic production environments.

Responsibilities

  • Provide L1/L2 teams with escalation support for day-to-day operational issues and network-related incidents, ensuring timely troubleshooting and resolution.
  • Provide weekend technical and escalation support for network break-fix activities and planned maintenance, helping mitigate operational risk and maintain infrastructure stability.
  • Implement, and support enterprise-scale network infrastructure across data centers, cloud, campus and branch environments.
  • Provide Tier-3 support and technical leadership during major incidents and escalations.
  • Develop and maintain network standards, update documentation, and engineering procedures.
  • Coordinate maintenance windows, rollback plans, validation testing, and post-implementation reviews.
  • Coordinate end-to-end remediation activities including maintenance windows and pre/post validations across multiple stakeholders, including business application owners, multiple infrastructure teams (Unix, Load Balancer, Firewall etc), security teams, and all relevant third-party vendors (Cisco/Arista etc).
  • Monitor network performance and proactively identify risks before they impact production.
  • Develop automation solutions using Python, Ansible, or similar tools to improve operational and deployment efficiency.
  • Standardize deployment procedures and reduce manual intervention through automation.
  • Drive operational excellence through continuous process improvement and infrastructure optimization.
  • Ensure compliance with security, audit, and regulatory requirements through timely remediation of vulnerabilities.

Benefits

  • Comprehensive employee benefits and perks in the industry.
  • Ample opportunity to move about the business for those who show passion and grit in their work.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service