Staff Network Reliability Engineer, RAN Operations

Skylo Technologies,
$150,000 - $162,000Remote

About The Position

The Staff Network Reliability Engineer, RAN Operations, in the Global Product Support & Customer Success organization, serves as the RAN domain authority within Skylo's production NTN network. This role is the escalation target for all RAN-domain Sev 1-2 events, responsible for diagnosing and resolving issues related to SINR anomalies, beam coverage gaps, preamble failures, timing drift, interference conditions, and gNodeB/eNodeB hardware faults at a signal and protocol level. The position owns the health of the RAN layer, which is the air interface between space and device, ensuring every attach, random-access attempt, and RRC connection depends on the NTN vRAN's integrity. Responsibilities include defining service-aligned KPIs, leading continuous RF optimization, writing and owning runbooks, setting detection logic and diagnostic standards, and identifying toil and failure patterns as engineering requirements. At the Staff NRE level, this role also acts as a force multiplier by mentoring Senior NREs, contributing to the automation backlog with operational requirements, and partnering with RAN Product Engineering to ensure new software releases and parameter changes meet operational readiness standards before production deployment.

Requirements

  • 8–10+ years of experience in RAN engineering and operations in a production 24x7 environment — carrier or vendor side, with direct ownership of live LTE/5G RAN infrastructure.
  • Deep 3GPP RAN expertise: TS 38.300, NR/LTE-NTN procedures, NB-IoT/LTE-M/5G-NR air interfaces, UE random-access procedures (PRACH/MSG1–MSG5), RRC state machine, and beam management.
  • Extensive understanding of vRAN/ORAN architecture: 7-2x split functions, RU/DU/CU-CP/CU-UP component roles, eCPRI/CPRI interfaces, and virtualized RAN software systems.
  • Production RAN troubleshooting: demonstrated ability to diagnose SINR anomalies, beam coverage gaps, preamble failures, timing drift, interference conditions (CW, narrowband, adjacent carrier), and RU/DU/CU hardware and software faults using RF telemetry and gNB logs.
  • RF and timing domain knowledge: GNSS, PTP, SyncE, Octoclock/reference clock behavior, and NTN-specific timing compensation mechanisms.
  • Production observability: Prometheus/Grafana/in house tools for RAN KPI dashboarding; OSS alarm integration (SNMP, Pub/Sub, or equivalent).
  • Packet capture and trace analysis: proficiency with Wireshark or equivalent for L2/L3 call flows, PCAP analysis, and gNB trace decoding.
  • Kubernetes operational literacy: pod health monitoring for DU/CU workloads, kubectl proficiency, log aggregation and correlation.
  • Runbook authorship: ability to write RAN diagnostic procedures at the level where a less-experienced engineer can execute them independently.
  • Strong written and verbal communication: capable of delivering RCA documents, engineering escalations, and MNO-facing technical summaries.

Nice To Haves

  • Direct experience with NTN or satellite RAN operations: NTN-specific random-access timing, satellite beam management, and coverage optimization over LEO/GEO constellations.
  • RF design, optimization, and performance improvement experience — parameter tuning, coverage analysis, interference hunting.
  • Experience with RAN vendors: Samsung, Mavenir, Fujitsu, Nokia — including vendor CLI tools, log formats, counter definitions, and support escalation processes.
  • Deep understanding of LTE, VoLTE, 5G VoNR L2/L3 call flows with ability to analyze traces.
  • Scripting ability in Python or Bash — sufficient to automate RF log parsing, KPI extraction, or build diagnostic utilities.
  • Experience contributing to closed-loop automation requirements or service assurance platform development for RAN event categories.
  • Knowledge of cloud-native environments (GKE, AWS) hosting virtualized RAN functions.

Responsibilities

  • Own 24x7 NTN RAN health across Skylo's production network: gNodeB/eNodeB status, node availability, attach and access success rates, RRC setup and abnormal release behavior, RF integrity (SINR, RSSI, noise floor stability), and RAN-layer SLA compliance.
  • Monitor and triage RAN alarms using OSS dashboards, Grafana/inhouse telemetry, and NTN-specific alarm patterns — distinguish transient RF fluctuations from systemic degradation before escalating or acting.
  • Execute and own RAN-domain runbooks for all fault categories: cell outage recovery, timing resynchronization, parameter rollback, restart procedures, and interference mitigation actions — without requiring engineering team involvement for covered fault classes.
  • Own timing and synchronization health: GNSS lock status, Octoclock/reference clock drift, PTP/SyncE alignment, and NTN-specific timing compensation — detect and resolve timing anomalies before they cascade into access failures.
  • Monitor beam-level and cell-level performance across the NTN coverage area; identify underperforming beams and initiate optimization or escalation based on defined criteria.
  • Serve as the L3 escalation authority for all RAN-domain incidents: take ownership from the Incident Manager, diagnose at the RF and protocol level using eNodeB/gNodeB logs, RF telemetry, PRACH/MSG1–MSG5 analysis, RRC traces, and KPI correlation, and deliver a resolution or a decision-grade root cause.
  • Lead RAN-domain troubleshooting bridges: command the technical investigation, direct vendor and engineering participants, correlate signals across RU, DU, CU-CP, CU-UP, and timing subsystems, and drive the bridge to a documented resolution or a clear engineering handoff.
  • Diagnose and resolve RAN failure modes: SINR anomalies, beam coverage gaps, preamble failures (MSG1–MSG5), RRC setup failures, abnormal RRC releases, CW/narrowband/adjacent-carrier interference, cell unavailability, eCPRI link failures, and DU/CU software faults.
  • Engage RAN vendors with technical specificity: reproduce failures with log evidence and RF traces, own the vendor ticket lifecycle, enforce SLA response commitments, and escalate vendor delays with full impact context.
  • Participate in the global 24x7 on-call rotation as the RAN domain escalation tier — reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker, not the first responder.
  • Own RAN-domain RCA end-to-end: lead the post-incident investigation, document the complete causal chain from triggering RF condition or hardware event through downstream subscriber impact, and deliver systemic action items with owners, timelines, and measurable success criteria.
  • Deliver Initial RCA documentation within defined SLA windows post-incident closure; own the final RCA through engineering review and sign-off.
  • Identify systemic RAN failure patterns — recurring interference sources, timing drift trends, vendor software regressions, hardware batch faults — and translate them into engineering requirements with clear impact, scope, and acceptance criteria.
  • Contribute to the weekly and monthly Network Performance Report: RAN availability by beam/cell, attach success rates, MTTR by fault category, top recurring issues, and SLA deviation analysis.
  • Define, implement, and continuously refine RAN KPIs: attach and access success rates, RRC setup and release behavior, SINR/RSSI distributions, interference metrics (CW, narrowband, adjacent carrier), MSG5 preamble success rates, timing health indicators, and chronic degradation tracking — calibrated to Skylo's NTN behavior, not vendor-default counters.
  • Lead continuous optimization of Skylo's live NTN RAN: parameter tuning for NTN-specific behavior, beam-level and cell-level performance optimization, coverage improvements, and interference mitigation.
  • Proactively track RAN availability and performance metrics against MNO SLA commitments — flag degradation trends before they breach thresholds and initiate preventive optimization before a subscriber impact occurs.
  • Partner with Change Management to ensure safe execution of all RAN configuration changes: auditable change records, peer review gates, and explicit rollback plans for every optimization action.
  • Author, own, and maintain all RAN-domain runbooks and SOPs — every procedure tested before production reliance; runbooks are living documents updated after every incident that reveals a gap.
  • Define the diagnostic decision tree for each known RAN fault class: entry condition, RF triage steps, isolation method, resolution action, and escalation criteria — written at the level where a Senior NRE can execute independently.
  • Translate RAN behavior into operational artifacts that scale: detection logic, alert threshold definitions, escalation criteria — ensuring the team can detect, triage, and escalate RAN events accurately without your direct involvement.
  • Identify runbook gaps from incident post-mortems and operational observations; prioritize gap closure based on incident frequency, subscriber impact, and MTTR.
  • Partner with RAN Product Engineering on software release readiness: define observability and operational acceptance criteria for new RAN builds before they reach production; flag missing alarm coverage, changed default parameters, and new failure modes.
  • Collaborate with Core NRE on end-to-end subscriber path issues where RAN and Core domains intersect: attach failures, handover interruptions, and signaling chain problems that span domains.
  • Collaborate with Cloud Infrastructure NRE on Kubernetes-layer issues affecting RAN NFs: DU/CU pod scheduling, PVC availability, network policy changes, and GKE upgrade impacts on RAN workloads.
  • Surface toil and automation opportunities to the Service Assurance & Automation team — document the procedure, frequency, and MTTR cost as structured input to the automation backlog; contribute to closed-loop automation requirements for P3/P4 RAN events.
  • Mentor Senior NREs in RAN domain depth: KPI interpretation, RF pattern recognition, trace analysis (PCAP, gNB traces), early degradation indicators, and escalation judgment.
  • Work closely with RAN vendors for incident resolution and RCA — manage vendor ticket lifecycle from opening through closure, ensure reproducible evidence including RF traces and gNB logs is provided, and escalate responsiveness failures.
  • Engage Skylo's RAN engineering team with full operational context when issues exceed operational resolution authority — deliver a structured problem statement, timeline, log bundle with RF telemetry, and a clear question rather than a vague escalation.
  • Support MNO partner technical discussions on RAN-layer SLA definitions, coverage performance, and NTN-specific RF behavior — provide operational evidence and data to support partner conversations.

Benefits

  • Competitive compensation packages including a stock option-based equity program
  • Comprehensive benefits including medical, dental, vision, and retirement plan
  • Monthly allowances for wellness and education reimbursement
  • A generous time-off policy, holidays, and the opportunity to temporarily work abroad
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service