Senior Platform Software Engineer

Oracle•Nashville, TN
•$92,500 - $209,500

About The Position

This role involves owning and evolving platform components, performing deep debugging across distributed services, and implementing resilience and capacity optimizations. The engineer will also be responsible for authoring documentation, samples, and adoption guidance. Key responsibilities include designing software solutions, adhering to the software development lifecycle, developing new features, leading code reviews, conducting debugging and troubleshooting, implementing software testing and quality assurance processes, conducting performance profiling and optimization, and troubleshooting API functionality and integration. The role also involves designing and developing software systems aligned with pre-defined architecture, collaborating with team leads, implementing performance optimization and scalability strategies, collaborating with teams to understand customer issues, providing technical guidance to customers, coaching and mentoring others, influencing the team to ensure customer satisfaction, and investigating complex product maintenance issues. Additionally, the role requires establishing and following development practices and coding standards, ensuring code quality, keeping up-to-date with industry best practices, implementing secure coding practices, and performing periodic maintenance and testing operations for systems. The engineer will independently manage work, monitor timelines, prioritize work, collaborate across teams, identify and address issues, analyze data for troubleshooting, contribute to knowledge sharing, develop ideas for process improvements, and embrace continuous learning.

Requirements

  • 4+ years of experience supporting commercial software, cloud infrastructure, or storage services in a distributed production environment.
  • Strong knowledge of Linux, operating systems, networking fundamentals, and production operations.
  • Knowledge of storage systems, file systems, and distributed-system fundamentals.
  • Experience with automation or infrastructure tooling such as Python, Terraform, Bash.
  • Strong troubleshooting, debugging, incident-management, and ticket-resolution skills.
  • Ability to coordinate effectively with development, operations, networking, security, and other partner teams.
  • Clear written and verbal communication, sound task management, and the ability to produce high-quality technical documentation.
  • Ability to interpret service metrics, logs, alarms, and operational trends during troubleshooting and reliability analysis.
  • Familiarity with Agile practices, change management, and production-readiness principles.

Nice To Haves

  • Hands-on experience with NetApp storage platforms, administration, monitoring, troubleshooting, or performance analysis is an added advantage.
  • Experience operating enterprise file-storage services or other large-scale storage platforms.
  • Experience with cloud infrastructure and services, particularly Oracle Cloud Infrastructure.
  • Experience designing or improving highly available, resilient, and scalable operational systems.
  • Experience developing automation for provisioning, deployment, monitoring, remediation, or operational workflows.
  • Understanding of storage protocols, data protection, backup, replication, and disaster-recovery concepts.
  • Demonstrated automation mindset: identifying repetitive operations and operational toil, automating them safely, and measuring the resulting improvement.
  • Creative and analytical problem solving, including structured brainstorming and root-cause analysis to produce scalable, durable solutions for recurring problems.

Responsibilities

  • Perform routine service operations, including host and storage-system management, security and compliance activities, automated and manual ticket resolution, and customer notifications.
  • Monitor and maintain the availability, reliability, performance, and capacity of OCI File Storage Servicess across pre-production and production environments.
  • Participate in customer-impacting incidents, lead timely mitigation, coordinate with partner teams, and drive long-term corrective actions.
  • Review and approve the operational readiness of new service features and changes.
  • Execute and own software and infrastructure deployments using approved automated tools and change-management practices.
  • Proactively identify recurring incidents, ticket noise, operational gaps, and sources of toil; create automation, bugs, runbooks, metrics, and alarms to address them.
  • Investigate performance, capacity, networking, host, and storage-system issues, escalating effectively when deeper product or vendor expertise is required.
  • Identify growth trends and provide input to capacity planning and service scaling.
  • Create and improve alarms, dashboards, operational procedures, troubleshooting guides, and service documentation.
  • Participate in large-scale events involving service dependencies and coordinate recovery activities.
  • Contribute to problem management by identifying root causes, tracking corrective actions, and preventing recurrence.
  • Continuously improve service operability, observability, resilience, and customer experience.
  • Owns moderately complex components within platform services or SDKs; leads team-level improvements to integration frameworks and developer tooling.
  • Performs deep debugging across a bounded set of distributed services, driving fixes that protect downstream consumers and upgrade paths.
  • Analyzes usage, performance, and error budgets for specific platform surfaces; implements targeted resilience and capacity optimizations.
  • Authors and curates team-scoped documentation, samples, and adoption guidance.
  • Own a bounded platform component (service module, SDK area) and evolve its contracts for multi-tenant use.
  • Perform deep debugging across a limited-service graph; drive compatibility-safe remediation plans.
  • Implement targeted resilience/capacity patterns and document adoption guidance for team consumers.
  • Designs software solutions and analyzes and helps identify requirements to achieve business and operational goals, independently.
  • Adheres to and suggests improvements to all phases of the software development lifecycle.
  • Utilizes working knowledge to develop new software features and enhancements following design specifications and develops documents to clarify software design and code.
  • Leads code reviews in designated areas to help drive improvements.
  • Conducts debugging and troubleshooting to identify and fix moderately complex software issues. Develops fixes for identified issues.
  • Implements software testing (e.g., functional and non-functional testing), quality assurance processes, software error logging, monitoring, and observability for effective debugging, ensuring review by manager and/or lead throughout the process.
  • Exercises judgment and discretion to conduct performance profiling and optimization of coding.
  • Troubleshoots and resolves moderately complex issues related to application programming interface (API) functionality and integration.
  • Implements moderately complex API versioning, lifecycle, and interoperability strategies.
  • Designs and develops software, systems, and services aligned to pre-defined system architecture.
  • Develops working knowledge of software architecture decisions and best practices.
  • Collaborates with team leads to review work to ensure alignment with software architecture.
  • Implements moderately complex performance optimization and scalability strategies in software design.
  • Collaborates within and beyond immediate team to understand customer issues and align solutions.
  • May provide technical guidance and support to customers regarding customer-reported issues, independently.
  • Coaches, mentors, and guides others to advocate for customers' interests and suggests product enhancements based on feedback.
  • Influences team to ensure customer satisfaction through timely resolution of issues and effective communication.
  • Implements customer issue and/or defect handling and training processes, independently.
  • Investigates and troubleshoots complex product maintenance issues to ensure customer agreement on short- and long-term solutions (e.g., future enhancements).
  • Collaborates with the team to establish and follow development practices and coding standards.
  • Exercises judgment and discretion to ensure code quality and adherence to broad acceptance criteria during development.
  • Keeps up to date with industry best practices and applies them to software development processes.
  • Implements secure coding practices to prevent security vulnerabilities.
  • Performs periodic maintenance and testing operations for systems that require upgrading or patching (e.g., for critical vulnerabilities).
  • Exercises judgment and discretion to drive improvements, ensure automation, testing, and debugging of systems to ensure service/product availability, health, support, and reliability.
  • Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.
  • Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.
  • Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices.
  • Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.
  • Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.

Benefits

  • Medical, dental, and vision insurance, including expert medical opinion
  • Short term disability and long term disability
  • Life insurance and AD&D
  • Supplemental life insurance (Employee/Spouse/Child)
  • Health care and dependent care Flexible Spending Accounts
  • Pre-tax commuter and parking benefits
  • 401(k) Savings and Investment Plan with company match
  • Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
  • 11 paid holidays
  • Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
  • Paid parental leave
  • Adoption assistance
  • Employee Stock Purchase Plan
  • Financial planning and group legal
  • Voluntary benefits including auto, homeowner and pet insurance
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service