Software Engineer | Observability

SingleStore•San Francisco, CA
•Remote

About The Position

We are seeking a Software Engineer to join the Observability Team and play a critical role in designing and delivering core capabilities for SingleStore's observability platform. This is a hands-on engineering role with end-to-end ownership of features and projects at the intersection of distributed systems, cloud infrastructure, database technology, and AI-powered observability. As a Software Engineer on this team, you will solve complex system-level problems, contribute meaningfully to technical direction, and partner with Product and customer-facing teams to ensure our platform meets the needs of enterprise customers and new adopters alike. The team provides comprehensive customer observability over their workloads — enabling customers to understand their usage patterns, identify bottlenecks, optimize performance, and leverage powerful alerting capabilities. Our platform unifies traces, logs, and metrics through an open-source-first approach. This is an ideal role for an engineer who thrives on deep technical challenges, takes pride in building durable systems, and is excited about bringing AI-powered observability experiences to customers.

Requirements

  • 2+ years of professional software development experience building distributed systems or backend services
  • Strong proficiency in Go (Golang) — experience with Rust, Python, or C++ is also valuable
  • Deep understanding of distributed systems concepts: scalability, consistency, high availability, concurrency, and failure modes
  • Familiarity with distributed systems managed via Kubernetes
  • Demonstrated ability to design and build reliable, high-performance system software
  • Experience working in environments where performance, scalability, and reliability are critical
  • Familiarity with observability concepts: traces, logs, metrics, APM, and monitoring patterns
  • Strong problem-solving and debugging skills with the ability to root-cause complex production issues
  • Excellent communication skills, both written and verbal, with ability to collaborate in multicultural, remote-first teams
  • Code quality mindset: you value simplicity, performance, maintainability, and thorough testing

Nice To Haves

  • Experience with time-series data and understanding of metrics cardinality challenges
  • Proficiency with SQL and experience working with relational or distributed databases
  • Experience building cloud-native SaaS platforms with multi-tenant architecture
  • Multi-cloud experience: working with AWS, GCP, Azure, or other cloud providers in a production setting
  • Kubernetes proficiency: operating, monitoring, or developing for Kubernetes clusters
  • Open-source observability tools: hands-on experience with Grafana, Alertmanager, Loki, Tempo, OpenTelemetry Collector, OTLP protocol
  • OpenTelemetry expertise: experience with, or active contributions to OTel projects
  • Time-series database experience: Prometheus TSDB, InfluxDB, Mimir, TimescaleDB, or SingleStore
  • Experience with data pipeline technologies: Apache Kafka, Parquet, Arrow, or stream processing frameworks (Flink, etc.)
  • Experience working with AI agents or LLM-powered applications, including agentic workflows for observability — enabling customers to query telemetry data in natural language
  • Experience in a SaaS or cloud-native company delivering managed services to customers

Responsibilities

  • Design and implement scalable observability features for traces, logs, and metrics — spanning ingestion, processing, storage, and visualization
  • Work across control plane and data plane components in a multi-cloud environment (AWS, GCP, Azure), ensuring reliable operation and data consistency at scale
  • Build high-throughput data pipelines that process telemetry data using OpenTelemetry Collector and related open-source tooling
  • Develop and maintain alerting capabilities with Alertmanager, enabling customers to define, tune, and manage alerts with routing, inhibition, and notification management
  • Optimize time-series data storage and query performance using SingleStore DB, handling high-cardinality data and complex analytical queries
  • Contribute to data visualization dashboards in Grafana, creating intuitive customer-facing experiences to explore telemetry data
  • Collaborate closely with Product Management to translate customer and business requirements into robust technical solutions
  • Investigate and resolve difficult issues in production and development environments, debugging data synchronization across distributed systems and cloud providers
  • Participate in on-call rotations to ensure system reliability and respond to incidents promptly

Benefits

  • Comprehensive customer observability over their workloads — enabling customers to understand their usage patterns, identify bottlenecks, optimize performance, and leverage powerful alerting capabilities.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service