Staff Software Engineer - Reporting, Data Platform & Observability

OneTrustAtlanta, GA
$139,725 - $209,588Hybrid

About The Position

OneTrust is seeking a Staff Software Engineer to join the Reporting and Data Platform team. This is a hands-on individual contributor role focused on designing, building, operating, and improving distributed backend services and data-processing platforms. You will work across Java microservices and Python/PySpark data pipelines, with a strong focus on reliability, scalability, performance, and observability. You will take complex or ambiguous problems from investigation through production delivery and help improve the systems that power reporting and data-driven experiences.

Requirements

  • Strong professional experience building and operating production software systems as a highly autonomous individual contributor.
  • Strong proficiency in Java and Spring Boot, with experience designing and operating distributed systems and microservices.
  • Production experience with asynchronous or event-driven systems, preferably Apache Kafka.
  • Strong experience with Python, PySpark, Apache Spark, and Delta Lake, plus production experience with Azure Databricks or a comparable managed Spark platform.
  • Hands-on experience with test-driven development, automated testing strategies, and quality gates that support fast, reliable delivery.
  • Strong understanding of metrics, logs, distributed tracing, dashboards, monitoring, and alerting, including hands-on experience with Datadog and Grafana.
  • Experience creating or responding to PagerDuty incidents, Datadog alerts, or equivalent production alerting workflows, and willingness to participate in an on-call rotation.
  • Experience using AI engineering tools such as Devin, Claude, or similar systems to produce production-ready code, tests, documentation, and operational improvements.
  • Strong design-thinking skills and the ability to reduce delivery-cycle time through clear architecture, smaller increments, reusable patterns, and pragmatic technical trade-offs.
  • Ability to independently diagnose complex performance and reliability problems and communicate implementation decisions and technical trade-offs clearly.

Nice To Haves

  • Experience with both batch and streaming data pipelines and with optimizing Spark or Databricks workloads for performance, reliability, and cost.
  • Experience with Databricks SQL, Databricks SDKs, Delta operations, and schema migrations.
  • Familiarity with Azure Blob Storage, Azure Identity, and Azure Key Vault.
  • Experience operating reporting, analytics, dashboard, or large-scale export systems.
  • Experience with Kubernetes, containers, CI/CD, and infrastructure as code.
  • Experience defining or applying service-level indicators, service-level objectives, and error budgets, and using incident and alert trends to prioritize engineering work.
  • Understanding of data governance, encryption, audit-ability, and tenant isolation.
  • Experience modernizing established production systems incrementally.

Responsibilities

  • Own complex features and technical improvements from discovery through production rollout, making substantial hands-on contributions across backend services and data-processing pipelines.
  • Investigate ambiguous problems, identify root causes, evaluate trade-offs, and implement pragmatic solutions that improve code quality, maintainability, automated testing, and operational readiness.
  • Review code and technical designs, document important implementation decisions and system behavior, and partner with product managers, engineers, and other teams to clarify requirements and deliver outcomes.
  • Apply AI-assisted engineering tools such as Devin, Claude, or similar systems to accelerate delivery while maintaining production-quality design, code, tests, security, and operational readiness.
  • Design and implement production services using Java, Spring Boot, and Maven, including APIs, asynchronous workflows, report generation, aggregation, export, and scheduling capabilities.
  • Develop event-driven functionality using Kafka and related messaging patterns, and work with caching technologies, relational storage, and service-to-service integrations.
  • Improve service performance, scalability, fault tolerance, and resource efficiency through appropriate patterns for retries, idempotency, caching, backpressure, concurrency, and failure recovery.
  • Diagnose issues across services, queues, databases, and downstream dependencies, and modernize established capabilities incrementally while maintaining production stability.
  • Build and maintain ingestion and transformation pipelines using Python, PySpark, Azure Databricks, and Delta Lake across batch and streaming workloads.
  • Implement schema evolution, checkpoint management, deduplication, replay, late-arriving-data handling, and standardized data-layer patterns.
  • Optimize Spark joins, partitioning, Delta operations, cluster utilization, and query performance while troubleshooting failed, delayed, or inefficient Databricks workloads.
  • Protect tenant boundaries across joins, aggregations, deduplication, and Delta operations; implement data-quality controls; and monitor data freshness, completeness, and correctness.
  • Work securely with Azure storage, identities, secrets, and encryption mechanisms.
  • Improve observability across backend services, event-driven workflows, and data pipelines using meaningful metrics, structured logs, traces, and business telemetry.
  • Build and maintain actionable dashboards, monitors, and alerts using Datadog and Grafana, applying OpenTelemetry, Prometheus, and Micrometer patterns where appropriate.
  • Participate in the on-call rotation and incident-response workflows, using PagerDuty, Datadog monitors, or equivalent platforms to diagnose production issues and drive sustainable resolution.
  • Reduce recurring alerts and operational toil by improving alert quality, eliminating noisy or non-actionable monitors, creating runbooks and diagnostic tools, and implementing corrective actions from blameless incident reviews.
  • Improve end-to-end correlation and monitor availability, error rates, latency, ingestion lag, data freshness, event throughput, consumer lag, job health, rejected records, checkpoint health, tenant-specific failures, data-quality violations, and Spark resource utilization.

Benefits

  • Comprehensive healthcare coverage
  • Flexible PTO
  • Equity RSUs
  • Annual performance bonus opportunities
  • Retirement account support
  • 14+ weeks of paid parental leave
  • Career development opportunities
  • Company-paid privacy certification exam fees
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service