Windows Site Reliability Engineer

Cloudious•Jersey City, NJ
•Onsite

About The Position

We are seeking a Windows Site Reliability Engineer to join our team. This role involves taking ownership of vendor applications, ensuring their smooth operation, and managing their lifecycle. You will be responsible for learning and documenting each vendor's application stack, acting as the technical point of contact with vendors, and assessing risks associated with releases and patches. The role also includes application packaging, building and maintaining CI/CD pipelines, and managing deployments across different environments. Post-production operations, including technology lifecycle management, defining SLIs/SLOs, monitoring, incident response, and automating repetitive tasks, are also key responsibilities. You will also be responsible for creating runbooks and knowledge articles to support operations teams.

Requirements

  • 5+ years supporting enterprise Windows applications in an SRE, application support, or systems engineering role.
  • Hands-on application packaging experience: MSI/MST, MSIX or App-V, silent install parameters, and packaging tools such as Orca, InstallShield, Advanced Installer, or PSADT.
  • Software distribution with SCCM/MECM and/or Intune.
  • Strong Windows Server and Windows client administration, including services, registry, permissions, Group Policy, and Active Directory.
  • Experience supporting thin-client (Citrix, RDS, or AVD), thick-client, and IIS-hosted web applications.
  • Proven experience supporting third-party vendor software and working directly with vendor support and engineering teams.
  • CI/CD pipeline experience with tools such as Azure DevOps, Jenkins, or GitHub Actions, plus Git source control.
  • Strong PowerShell scripting for automation and toil reduction.
  • Experience deploying across lower environments and production under formal change management (ITIL or similar).
  • Working knowledge of SQL Server sufficient to support application databases and troubleshoot connectivity and performance.
  • Experience with monitoring and observability tools such as SCOM, Splunk, Dynatrace, Datadog, or Azure Monitor.
  • Clear written communication for runbooks, RCAs, and vendor escalations.

Nice To Haves

  • Infrastructure as code and configuration management (Terraform, Ansible, PowerShell DSC).
  • Cloud hosting experience for Windows workloads in Azure or AWS.
  • Containerizing Windows or .NET applications, or MSIX app attach for AVD.
  • Experience with vulnerability remediation and patch compliance for third-party software.
  • Familiarity with SRE practices such as error budgets, toil budgets, and capacity planning.
  • Certifications such as Microsoft Certified: Azure Administrator, Endpoint Administrator, ITIL Foundation, or Citrix CCA.

Responsibilities

  • Vendor application ownership: Learn and document each vendor's application stack, including client type, dependencies, databases, middleware, licensing, and integration points.
  • Act as the technical point of contact with vendors for releases, patches, defects, and escalations, holding them to agreed SLAs.
  • Review vendor release notes and installers to assess risk and impact before any change reaches production.
  • Application packaging: Package, repackage, and customize Windows thick-client applications for silent, repeatable deployment (MSI, MST transforms, MSIX, App-V, PSADT wrappers).
  • Resolve packaging issues such as dependencies, prerequisites, file and registry conflicts, and per-user versus per-machine installs.
  • Maintain a versioned package library and distribute packages through SCCM/MECM or Intune.
  • Release and deployment: Build and maintain CI/CD pipelines for vendor applications and their configuration (e.g., Azure DevOps, Jenkins, or GitHub Actions).
  • Promote releases through Dev, Test, UAT, and Production with consistent configuration and documented rollback plans.
  • Plan and run production deployments, including change requests, maintenance windows, smoke tests, and post-deployment validation.
  • Deploy and configure web applications on IIS, and thin-client apps on Citrix or RDS/AVD.
  • Post-production operations: Own technology lifecycle management (TLM): track vendor support dates, plan upgrades, and retire end-of-life versions before support lapses.
  • Define SLIs and SLOs, and build monitoring, alerting, and dashboards for application health.
  • Lead incident response and blameless root-cause analysis, and drive fixes with vendors and internal teams.
  • Measure and reduce toil by automating repetitive work such as installs, configuration, health checks, certificate renewals, and patching.
  • Write runbooks and knowledge articles so support teams can operate each application independently.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service