Senior Linux Infrastructure Engineer

Northern LightSomerville, MA
Hybrid

About The Position

Northern Light provides the world's most sophisticated machine learning-powered competitive intelligence platform for market research. For over 25 years, we've been helping Fortune 1000 enterprises make smarter, faster, and more informed decisions through our award-winning SinglePoint knowledge management platform. Our clients include global leaders across technology, pharmaceuticals, telecommunications, and life sciences who depend on us to transform fragmented data into strategic clarity. We're a company that takes pride in our compulsive drive to provide exceptional client support. We wake up each day ready to tackle the challenges of knowledge management and we never stand still. Our recent innovations include generative AI capabilities, machine learning insights, and advanced competitive intelligence automation. The Opportunity Northern Light is seeking a Senior Linux Infrastructure Engineer to take hands-on ownership of the Linux infrastructure behind our platform's compute-heavy backend, which runs on our own hardware in a private colocation cage in Somerville, MA. You will be the primary owner of that datacenter environment: the person who keeps it reliable and secure today, and who leads its next phase, including a hardware refresh across the fleet and deeper automation of the environment. On our Platform team you will own Linux infrastructure, working closely with the engineering, security, and operations teams, including the team that runs our customer-facing frontend in AWS. Our philosophy is to buy our platforms rather than build them: where a supported, vendor-backed product exists, we run it and follow the vendor's best practices instead of maintaining our own substitute. Automation on top of those platforms is very much your work, and we want someone who writes it well and knows how to get the most out of a vendor relationship.

Requirements

  • Substantial production Linux systems engineering experience, typically 7+ years, including several years where you were accountable for the reliability and security of the environment, whether alone or as a senior member of a small team.
  • Deep RHEL-family Linux skills (RHEL, Oracle Linux, Rocky, Alma, CentOS): systemd, kernel and performance tuning, storage (LVM, RAID, NFS), and networking (bonding, VLANs, firewalld/iptables).
  • Production experience administering a virtualization platform (VMware vSphere, KVM/libvirt, OpenShift Virtualization, Proxmox, Hyper-V, or similar).
  • Hands-on Ansible authoring: you have written and maintained roles and playbooks, not only run them.
  • Experience running a patch and vulnerability management program in production: scheduled scanning with Tenable/Nessus, Qualys, Rapid7, or similar, interpreting results, and driving remediation across a fleet.
  • Experience installing and operating server applications and database servers at the system level (packaging, storage, TLS, access control, backups) from vendor documentation, including the judgment to plan a database upgrade that can be rolled back and to treat a backup as unproven until it has been restored.
  • Experience with enterprise server hardware (HPE ProLiant or equivalent): out-of-band management (iLO/IPMI), firmware, diagnostics, and component replacement, and comfort doing occasional physical work in a datacenter.
  • Experience designing or materially improving highly available, redundant infrastructure, and a track record of leading incident response and writing useful root-cause analyses.
  • Solid networking fundamentals: enough to make routine switch changes yourself, and to diagnose and scope switch, firewall, and load-balancer issues well enough to hand them to network engineers.
  • Strong documentation habits and clear written and spoken communication.
  • Must be authorized to work in the United States (unfortunately, we cannot sponsor visas).

Nice To Haves

  • A BS or MS in Computer Science, Computer Engineering, or Information Technology is a plus, but practical experience matters more to us.
  • Prior experience with the specific products named above is a plus but not required; we expect a strong engineer to pick them up here.

Responsibilities

  • Keep the fleet current and consistent: patch, harden, and upgrade Linux servers on a controlled cadence through managed repositories and staged rollouts, and install, upgrade, and maintain the applications and database servers to vendor guidance — largely through Ansible roles and playbooks you write and the Ansible Automation Platform you operate.
  • Run the vulnerability management cycle: scheduled Tenable scans, triage of findings, remediation prioritization, patching or documented mitigation, and compliance reporting that stands up to customer security reviews and audits.
  • Own infrastructure backups (policy, platform administration, restore testing) and partner with the application team on disaster recovery planning and exercises.
  • Keep monitoring and logging reliable and free of noise: complete coverage of the fleet, alerts that fire on real problems and not on everything else, and log collection you can trust when investigating an incident.
  • Lead incident response for infrastructure issues, run post-incident reviews, drive corrective actions to closure, and keep the resulting SOPs, runbooks, and infrastructure diagrams current and clear to technical and non-technical readers alike.
  • Run the datacenter as a remotely operated facility: accurate NetBox records, labeled cabling, iLO/out-of-band access to everything, spare parts on the shelf, and clear work orders for remote hands.
  • Work with vendors and our colo provider on hardware lifecycle and capacity, and coordinate scheduled maintenance windows and network changes with our network partner.
  • Participate in occasional scheduled after-hours maintenance (historically 3–5 times per year).

Benefits

  • Competitive salary
  • benefits
  • professional growth opportunities
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service