About The Position

Come join tastytrade, part of IG Group, as we build the reliability practice behind the brokerage platform that active options, futures, and traders rely on every market day. As our first Senior Site Reliability Engineer, you'll harden the systems behind order execution and market data delivery — designing for fault tolerance, closing gaps in our telemetry, and making sure our HashiCorp Nomad-based service fabric scales cleanly as trading volume grows. You'll work embedded alongside our infrastructure and application engineering teams, contributing directly to our Ruby, Java, and Elixir services. This is a rare opportunity to shape a practice and a culture from day one, on a platform where every order, quote, and position has to be right, because real client capital is on the line.

Requirements

  • Hands-on experience designing fault-tolerant, self-healing distributed systems — not just describing the patterns, but having shipped them.
  • Deep understanding of one or more: distributed systems, Linux systems, cloud-native architectures, containerization.
  • Experience running gap analyses on observability/telemetry systems: identifying what's not instrumented, not alerted on, or not visible until it's too late.
  • A track record scaling systems under real production load, including capacity planning and architectural bottleneck identification.
  • Hands-on experience with OpenTelemetry, Prometheus, and Grafana, with the ability to instrument services directly.
  • Strong Linux internals and networking fundamentals, including TCP/IP, UDP/multicast, packet capture, and flow analysis.
  • On-call experience on production systems and comfort building a blameless post-incident review process.
  • Working knowledge of SLOs and error budgets as a tool, not the job description; HashiCorp Nomad, Consul, or Vault experience is a strong plus.
  • Strong programming skills in a language such as Python, Ruby, Java, or similar.

Responsibilities

  • Build self-healing, fault-tolerant infrastructure and internal tooling that automates repetitive operational work and reduces toil for Platform and Application teams.
  • Run gap analysis across our observability stack to find blind spots in telemetry, logging, and alerting coverage, then close them so failures surface before customers feel them.
  • Own scalability work across our HashiCorp Nomad service fabric: capacity planning, load testing, and identifying architectural bottlenecks before they become incidents.
  • Extend our observability stack (Prometheus, Honeycomb, OpenTelemetry) with the instrumentation needed to actually see the failure modes above.
  • Set SLOs and error budgets with multi-window burn-rate alerting for critical brokerage flows, once the fault-tolerance and telemetry foundation is in place.
  • Mentor engineers across teams to build a culture of site reliability champions so the practice outlives any one person.

Benefits

  • Performance Bonuses
  • Stock Purchase Options
  • Medical/Vision/Dental Benefits
  • 401k Plan
  • 20 Paid Vacation Days (plus an additional paid vacation day the month of your birthday!)
  • 10 Paid Sick Days
  • Gym Membership Reimbursement
  • Commuter Benefits
  • Pet Insurance
  • Wellness & Mental Health Programs
  • Charitable Donation Matching
  • Two Paid Volunteer Days Off
  • Daily catered lunch when in the office
  • Full kitchen with snacks and beverages
  • In-building gym
  • Shuttle to/from Metra
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service