Lead Software Engineer – Systems Management Tools

TEKsystemsRocklin, CA
$70 - $75Onsite

About The Position

The Technical Infrastructure organization designs, builds, and operates the enterprise platforms that power Williams-Sonoma, Inc. across Ecommerce, Supply Chain, Corporate, Retail, and Distribution Center environments. Within this organization, the Systems Management Tools team develops and supports the enterprise observability, logging, monitoring, and automation platforms used across the company. The team is responsible for the software engineering, architecture, and operational excellence of OpenSearch, enterprise logging services, Kubernetes-hosted tooling, Kafka ingestion pipelines, Logstash processing, monitoring, and internally developed operational tools that enable engineering teams to build, operate, and troubleshoot enterprise applications at scale. The team partners closely with application engineering, platform engineering, infrastructure, cloud, security, network, DevOps, Site Reliability Engineering (SRE), and operations teams to deliver highly available, scalable, and secure observability platforms supporting both Ecommerce and Supply Chain business units. This reflects the team's documented ownership of OpenSearch clusters, Rancher/Kubernetes environments, Kafka ingestion, Filebeat onboarding, Logstash, and operational tooling. As Lead Software Engineer – Systems Management Tools, you will provide technical leadership for the design, development, architecture, and lifecycle management of Williams-Sonoma's enterprise observability platform and internally developed operational tools. You will lead software engineering initiatives centered around OpenSearch, enterprise logging, Kubernetes-based platform services, monitoring, automation, and operational tooling while helping define the technical roadmap for reliability, scalability, security, and operational excellence across the observability ecosystem. This role is ideal for a highly experienced software engineer who combines deep application development expertise with distributed systems knowledge, platform engineering experience, and operational excellence. You will lead complex engineering initiatives, mentor engineers, influence architecture, and partner across Infrastructure, Platform Engineering, Security, Cloud, and Application Development organizations to deliver enterprise-scale systems supporting mission-critical Ecommerce and Supply Chain platforms. The role aligns well with WSI's Lead Software Engineer expectations of authoring technical designs, technical ownership of delivery, reviewing architecture, influencing engineering practices, mentoring engineers, and driving cross-functional initiatives.

Requirements

  • Bachelor's degree in Computer Science, Software Engineering, Information Systems, or related field, or equivalent practical experience.
  • 10+ years of professional software engineering experience designing and developing enterprise applications.
  • Demonstrated experience leading complex software engineering initiatives across multiple engineering teams.
  • Strong programming experience in one or more modern languages such as Java, Python, Go, C#, or similar.
  • Expertise with OpenSearch and/or Elasticsearch administration, APIs, cluster architecture, monitoring, index management, and performance tuning.
  • Experience with Kubernetes, Rancher, Docker, Helm, container orchestration, and cloud-native application development.
  • Experience developing and maintaining enterprise logging platforms using Kafka, Filebeat, Logstash, and distributed messaging technologies.
  • Experience designing REST APIs, microservices, distributed applications, and automation platforms.
  • Strong understanding of CI/CD pipelines, Git-based development, automated testing, Infrastructure-as-Code, and DevSecOps practices.
  • Experience with observability platforms, monitoring systems, metrics collection, alerting, and operational analytics.
  • Experience implementing enterprise authentication and security integrations such as LDAP, SAML, RBAC, or OAuth is highly desirable.
  • Experience with AI-assisted operational tooling, automation, or AIOps solutions is preferred, reflecting the team's current direction toward AI-enabled OpenSearch and operational support tools.
  • Strong understanding of distributed systems, software architecture, resiliency patterns, and performance optimization.
  • Strong written and verbal communication skills, including executive-ready technical documentation and presentations.

Nice To Haves

  • You possess deep expertise in modern software engineering principles, distributed systems, and enterprise platform development.
  • You have extensive experience developing enterprise applications using Java, Python, Go, C#, or similar object-oriented programming languages.
  • You have significant experience designing and supporting OpenSearch or Elasticsearch platforms in enterprise environments.
  • You understand distributed search architectures, indexing strategies, cluster design, shard allocation, replication, index lifecycle management, and performance optimization.
  • You have experience developing software that integrates with Kubernetes, Rancher, Docker, REST APIs, and cloud-native platforms.
  • You have strong experience building enterprise automation using APIs, scripting, Infrastructure-as-Code, and CI/CD pipelines.
  • You understand enterprise logging architectures including Filebeat, Kafka, Logstash, ingestion pipelines, structured logging, and monitoring platforms.
  • You have experience developing operational tooling, administrative applications, dashboards, and self-service platforms for engineering organizations, consistent with the team's documented OpenSearch utilities and AI-enabled operational tools.
  • You possess strong software architecture skills with experience designing scalable, maintainable, highly available enterprise applications.
  • You can lead technical discussions across software engineering, infrastructure, cloud, security, and operations organizations.
  • You have demonstrated success mentoring engineers and raising engineering quality through code reviews, architecture reviews, and technical leadership.
  • You possess excellent communication, documentation, organizational, and stakeholder management skills.

Responsibilities

  • Lead the technical architecture, software development, and long-term evolution of the Systems Management Tools platform supporting enterprise logging, monitoring, observability, and operational automation.
  • Design, develop, and maintain enterprise software solutions supporting OpenSearch, Logstash, Kafka, Filebeat, Kubernetes, Rancher, and internally developed operational tooling.
  • Define engineering standards for software architecture, API design, coding standards, testing, deployment automation, documentation, and operational readiness.
  • Drive the technical roadmap for OpenSearch platform modernization, logging architecture, enterprise search, monitoring, alerting, and automation capabilities.
  • Lead development of internally developed applications and utilities supporting OpenSearch administration, diagnostics, health monitoring, patch automation, cluster lifecycle management, AI-assisted operations, and operational self-service capabilities, consistent with the team's documented OpenSearch Tool Kit, MCP platform, AI assistants, and related tooling.
  • Provide technical leadership for enterprise logging architecture, including Filebeat onboarding, Kafka messaging, Logstash pipelines, index lifecycle management, ingestion optimization, and OpenSearch cluster architecture.
  • Lead software engineering efforts that improve operational efficiency through automation, self-service tooling, infrastructure APIs, CI/CD integration, and Infrastructure-as-Code practices.
  • Design scalable solutions supporting high-volume log ingestion, distributed search, enterprise monitoring, and operational analytics across Ecommerce and Supply Chain platforms.
  • Lead cross-functional software engineering initiatives involving OpenSearch platform upgrades, Kubernetes platform enhancements, logging modernization, observability improvements, and enterprise automation projects.
  • Partner with Infrastructure Engineering, Cloud Engineering, Platform Engineering, Security, Application Development, and Site Reliability Engineering teams to deliver resilient enterprise platform services.
  • Guide incident response, production troubleshooting, root cause analysis, post-incident reviews, and long-term engineering improvements for enterprise observability services.
  • Establish engineering metrics, operational dashboards, performance benchmarks, and platform health reporting supporting service reliability and continuous improvement.
  • Evaluate emerging technologies related to AI-assisted operations, observability, distributed search, platform engineering, automation, and software development.
  • Mentor engineers through architecture reviews, code reviews, design discussions, technical coaching, and software engineering best practices.
  • Drive documentation excellence across architecture decisions, API specifications, operational runbooks, deployment standards, software lifecycle documentation, and engineering standards.
  • Represent the Systems Management Tools team in enterprise architecture reviews, technology governance, engineering strategy discussions, and cross-functional planning.
  • Participate in critical production support, after-hours maintenance, and major incident response as required.

Benefits

  • Medical, dental & vision
  • Critical Illness, Accident, and Hospital
  • 401(k) Retirement Plan – Pre-tax and Roth post-tax contributions available
  • Life Insurance (Voluntary Life & AD&D for the employee and dependents)
  • Short and long-term disability
  • Health Spending Account (HSA)
  • Transportation benefits
  • Employee Assistance Program
  • Time Off/Leave (PTO, Vacation or Sick Leave)
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service