SRE DevOps Engineer

RedolentSunnyvale, CA
Onsite

About The Position

One of our direct client is urgently looking for a SRE DevOps Engineer @ Sunnyvale, CA. This role involves supporting Java full stack backend application system components in a massively scalable, high performance, multi-tenant, international eCommerce platform with multiple micro-services deployed in a cloud environment. The engineer will be responsible for root-causing reactive and proactive production issues, leading and participating in complex, cross-functional projects, and partnering with architects and development leads to create high-level designs that accelerate the omni-customer experience. The role also involves proactively identifying areas for automation, speed, and innovation, troubleshooting business and production issues, and providing support to the business by responding to user questions and concerns. Additionally, the engineer will assist in guiding small groups of engineers, including offshore associates, and demonstrate up-to-date expertise in applying best practices. The role also requires modeling compliance with company policies and procedures and providing support for business solutions by building partnerships with stakeholders, identifying business needs, and monitoring progress.

Requirements

  • Hands on experience debugging 5xx and 4xx Java/Spring and Node/Python.
  • Experience creating database objects (tables, views, indexes).
  • CI/CD experience automation and implementation experience.
  • Experience with event streaming platforms like Kafka is a plus.
  • Experience with analytics & monitoring platform like Grafana/graphite/MMS/Splunk is a plus.
  • Splunk
  • Grafana
  • SRE
  • Cloud
  • DevOps
  • Azure
  • Docker
  • Kubernetes
  • Java (Basic)
  • Python (Scripting)
  • Experience in support and triage production incidents.
  • Experience with Application development and root cause analysis.
  • Experience in driving high availability across multiple organizations.
  • Experience in putting together architecture diagrams.
  • Infrastructure experience that involves, setup, scale, and decommissioning.
  • Prior cloud experience, planning and driving efficiencies.
  • Automation and CI/CD experience.
  • Application container experience using Kubernetes.
  • Implement the database structure such as tables, indexes.
  • Reviewing and tuning the SQL scripts.
  • Reviewing database structure changes that provided by application developers and data modelers.
  • Working with application developers to tune the performance of the database.
  • Experience creating best in class application availability metrics and dashboards.
  • Managing infrastructure scale, setup and decommissioning.
  • Public cloud experience, Azure, GCP and Private Data Centers.
  • Driving P1 production incident calls, communicating up to the point & summarizing action plans for each owner and follow-up until closure.
  • Ability to take right priority decision and run the operational excellence with innovative ideas, without much guidance/supervision.
  • Ability to build and run tools necessary for operational success.
  • Documenting SOPs for repetitive issues, building knowledge base articles for team’s benefit.

Nice To Haves

  • Experience with event streaming platforms like Kafka is a plus.
  • Experience with analytics & monitoring platform like Grafana/graphite/MMS/Splunk is a plus.
  • Kubernetes and Docker experience is a plus.

Responsibilities

  • Dig into issues on our eCommerce site and identify root cause.
  • Support and triage production incidents.
  • Create Dashboards, Alerting and Monitoring.
  • Act as a Subject Matter Expert.
  • Develops Innovation strategies, processes, automation, and failover experience.
  • Drives the execution of multiple business plans and projects.
  • Drive high availability across multiple organizations.
  • Put together architecture diagrams.
  • Manage workloads in private and public data centers.
  • Handle infrastructure experience that involves setup, scale, and decommissioning.
  • Provide prior cloud experience, planning and driving efficiencies.
  • Implement automation and CI/CD experience.
  • Manage application container experience using Kubernetes.
  • Support Java full stack backend application system components in a massively scalable, high performance, multi-tenant, international eCommerce platform with multiple micro-services deployed in cloud environment, root causing every reactive/proactive production issues.
  • Lead and participate in medium- to large-scale, complex, cross-functional projects.
  • Partner with architects and development leads to come up with high level design to accelerate omni-customer experience, recommending out-of-box engineering best practices.
  • Pro-Actively identify areas to drive automation/speed/innovation.
  • Troubleshoot business and production issues by gathering information (for example, issue, impact criticality, possible root cause); performing root cause analysis to reduce future issues; engaging support teams to assist in the resolution of issues; developing solutions; driving the development of an action plan; performing actions as designated in the plan; interpreting the results to determine further action; and completing online documentation.
  • Provide support to the business by responding to user questions, concerns, and issues (for example, technical feasibility, implementation strategies); researching and identifying needed solutions determining implementation designs; providing guidance regarding implications of new and enhanced systems; identifying short and long term solutions; and directing users to appropriate contacts for issues outside of associate's domain.
  • Assist in providing guidance to small groups of 5 to 6 engineers, including offshore associates, for assigned Engineering projects by proving pertinent documents, directions, examples, and timeline.
  • Demonstrate up-to-date expertise and apply this to the development, execution, and improvement of action plans by providing expert advice and guidance to others in the application of information and best practices; supporting and aligning efforts to meet customer and business needs; and building commitment for perspectives and rationales.
  • Model compliance with company policies and procedures and support company mission, values, and standards of ethics and integrity by incorporating these into the development and implementation/Support of business plans; using the Open Door Policy; and demonstrating and assisting others with how to apply these in executing business processes and practices.
  • Provide and support the implementation of business solutions by building relationships and partnerships with key stakeholders; identifying business needs; determining and carrying out necessary processes and practices; monitoring progress and results; recognizing and capitalizing on improvement opportunities; and adapting to competing demands, organizational changes, and new responsibilities.
  • Drive P1 production incident calls, communicating up to the point & summarizing action plans for each owner and follow-up until closure.
  • Take right priority decisions and run operational excellence with innovative ideas, without much guidance/supervision.
  • Build and run tools necessary for operational success.
  • Document SOPs for repetitive issues, building knowledge base articles for team’s benefit.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service