Site Reliability Engineer (SRE)

Rocket.netJupiter, FL
Remote

About The Position

At Rocket.net, reliability, performance, and customer experience are at the center of everything we build. We are looking for a Site Reliability Engineer to help maintain the health, stability, and performance of our hosting platform while providing advanced technical support to our customers. The Platform Operations team acts as a critical escalation layer between WordPress Support and Engineering. This role combines infrastructure operations, platform monitoring, troubleshooting, and advanced customer support. As a Site Reliability Engineer, you will help ensure Rocket.net's servers, services, and customer environments are operating at the highest standards. You will assist WordPress Support Engineers with complex issues, support VIP customers with advanced technical requests, investigate platform-level problems, and work with internal teams to deliver fast and effective solutions.

Requirements

  • 3+ years of experience in SRE, DevOps, Platform Engineering, or similar roles.
  • Strong experience troubleshooting Linux production environments.
  • Experience supporting customer-facing technical environments.
  • Strong understanding of web hosting technologies including NGINX, Apache, PHP-FPM, MySQL/MariaDB, and Redis.
  • Advanced troubleshooting skills across WordPress, servers, DNS, networking, and performance issues.
  • Experience with Linux command line (SSH).
  • Strong understanding of DNS, HTTP/HTTPS, SSL/TLS, CDN, and caching technologies.
  • Experience with Cloudflare, WAF, and web performance optimization.
  • Experience with monitoring tools and incident response processes.
  • Ability to troubleshoot complex issues independently and communicate technical solutions clearly.
  • Excellent written and verbal communication skills (English).
  • Ability to work under pressure during customer-impacting incidents.

Nice To Haves

  • Experience supporting managed WordPress hosting platforms.
  • Experience handling VIP customers or enterprise-level support.
  • Experience with high-traffic websites and performance optimization.
  • Knowledge of Bash scripting, yum/dnf, or automation tools.
  • Familiarity with observability tools such as Nedata or Datadog or similar.
  • Experience with incident management and postmortems.

Responsibilities

  • Monitor the health, availability, and performance of Rocket.net servers, services, and customer environments.
  • Proactively identify infrastructure issues, performance degradation, and potential service disruptions.
  • Investigate alerts and operational events to maintain platform stability.
  • Perform regular platform health checks and ensure critical systems are operating correctly.
  • Participate in incident response and coordinate troubleshooting during customer-impacting events.
  • Communicate platform issues, updates, and resolutions to relevant internal teams.
  • Provide advanced technical support for VIP customers and customers with complex hosting-related issues.
  • Act as a senior escalation point for WordPress Support Engineers when issues require deeper technical investigation.
  • Troubleshoot complex issues involving servers, websites, networking, DNS, performance, caching, and hosting infrastructure.
  • Assist customers with advanced technical problems beyond standard WordPress troubleshooting.
  • Investigate and resolve issues involving server resources, application performance, connectivity, and platform behavior.
  • Work directly with customers when required to provide expert-level technical assistance.
  • Ensure escalated customer issues are handled with urgency, ownership, and clear communication.
  • Troubleshoot and maintain Linux-based production environments.
  • Investigate issues related to NGINX, Apache, PHP-FPM, MySQL/MariaDB, Redis, and other platform services.
  • Assist with server maintenance, configuration changes, and operational improvements.
  • Support security updates, system hardening, and infrastructure best practices.
  • Monitor resource usage and identify capacity or performance concerns.
  • Help improve monitoring, automation, and operational workflows.
  • Work closely with WordPress Support Engineers, Shift Leads, Site Reliability Engineers, and Engineering teams.
  • Provide technical guidance and knowledge sharing to Support teams.
  • Help create internal documentation, troubleshooting guides, and knowledge base articles.
  • Identify recurring issues and recommend improvements to reduce future incidents.
  • Participate in incident reviews and root cause analysis.

Benefits

  • Ability to work from anywhere in the world.
  • You can travel and work in a new city every month.
  • Work and travel without ever using your vacation time.
  • This is a remote position.
  • Flexible Vacation.
  • Paid Education.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service