Cloud System Administrator 2

Wyetech•Laurel, MD
•Onsite

About The Position

Wyetech is seeking a highly skilled Cloud System Administrator Level 2 / Senior Cloud & Distributed Systems Engineer to support mission-critical cloud-based data repositories serving thousands of Intelligence Community users. This position operates in a dynamic, high-tempo operational environment where requirements can evolve rapidly in response to global events. The selected candidate will administer and engineer large-scale Hadoop and Accumulo clusters, maintain system reliability and security, and collaborate across infrastructure, networking, hardware, and security teams to ensure continuous mission availability. This is not a traditional system administration role. It is a reliability-focused distributed systems engineering position operating at scale, with a strong emphasis on automation, observability, troubleshooting, performance optimization, and continuous improvement.

Requirements

  • Seven (7)+ years of Linux Systems Administration experience.
  • Deep understanding of Linux operating systems and internals.
  • User and group account administration, including LDAP.
  • Configuration and administration of DHCP, DNS, and TFTP.
  • System patching, upgrades, and security hardening.
  • Performance tuning and resource optimization.
  • Minimum three (3) years of experience administering large distributed systems.
  • Experience supporting multiple clusters.
  • Experience supporting clusters spanning at least three racks.
  • Experience supporting environments with a minimum of 60 nodes per site.
  • Experience with Hadoop, including HDFS and YARN tuning.
  • Experience with Accumulo, including tablet balancing and performance optimization.
  • Experience with distributed storage technologies such as Cassandra, Scality, Swift, Gluster, Lustre, GPFS, Amazon S3, or comparable technologies.
  • Experience with Kubernetes orchestration services, with CKA-level knowledge preferred.
  • Docker containerization and image management.
  • Helm charts and cluster configuration.
  • StatefulSets and persistent volume management.
  • Cloud-based storage architectures.
  • Five (5)+ years of scripting experience using Bash, Python, or Perl.
  • Experience with configuration management technologies such as: Puppet Ansible Salt
  • Infrastructure as Code experience using Terraform or CloudFormation.
  • Experience integrating CI/CD pipelines and Git-based workflows.
  • Experience implementing and managing monitoring and observability solutions such as: Prometheus/Grafana ELK/OpenSearch Splunk Cloud-native monitoring platforms
  • Experience designing and tuning alerting frameworks.
  • Experience defining and supporting SLAs/SLOs.
  • Incident response participation and documentation.
  • Root Cause Analysis and post-incident review experience.
  • Capacity planning and performance analysis.
  • Understanding of VLANs, port-channel bonding, and Layer 2/Layer 3 interactions.
  • TCP/IP troubleshooting.
  • Experience with load-balancing technologies such as F5, HAProxy, and NGINX.
  • Firewall rule management.
  • Network performance analysis and troubleshooting.
  • RAID and storage architecture knowledge.
  • Object-storage optimization.
  • Data replication and backup strategies.
  • Multi-site failover and Disaster Recovery (DR) planning.
  • Understanding of RPO/RTO requirements and considerations.
  • Experience with Active/Active or Active/Passive cluster designs.
  • Experience with system hardening, with STIG implementation preferred.
  • Experience with vulnerability-scanning tools such as ACAS/Nessus.
  • Familiarity with the Risk Management Framework (RMF).
  • Security logging and audit compliance.
  • Experience operating within TS/SCI environments.
  • Experience with software development/engineering activities including requirements analysis, installation, integration, evaluation, enhancement, maintenance, testing, and problem diagnosis/resolution.
  • Experience with Open Source/NoSQL technologies supporting highly distributed and massively parallel computation, including HBase, Accumulo, and Bigtable.
  • Experience with the MapReduce programming model and technologies such as Hadoop, Hive, and Pig.
  • Experience with the Hadoop Distributed File System (HDFS).
  • Experience working with serialization formats such as JSON and/or BSON.
  • Experience developing or supporting RESTful services.
  • Experience working with UNIX/Linux operating environments.
  • Experience with Object-Oriented systems and requirements analysis.
  • Experience developing solutions that integrate and extend FOSS/COTS products.
  • Experience with software integration and testing, including test plans and test scripts.
  • Demonstrated technical-writing skills and experience producing technical documentation in support of engineering or software-development projects.
  • Minimum three (3) years of experience administering large distributed systems as described above.
  • Seven (7) years of Linux Systems Administration experience.
  • Five (5) years of scripting experience.
  • Security+ or other DoD 8570-compliant certification is required.
  • Candidate must also possess one (1) of the following certifications: AWS Certified SysOps Administrator – Associate AWS DevOps Engineer – Professional Certified Kubernetes Administrator (CKA)
  • Active TS/SCI security clearance with current polygraph is required.
  • Due to federal contract requirements, United States Citizenship and position appropriate security clearance is required.

Nice To Haves

  • CKA-level knowledge preferred for Kubernetes orchestration services.

Responsibilities

  • Monitor system health, performance, availability, and reliability across large distributed clusters.
  • Troubleshoot and resolve complex hardware, software, network, infrastructure, and cloud-platform issues.
  • Perform Root Cause Analysis (RCA) and contribute to post-incident reviews and corrective actions.
  • Maintain and optimize terabyte-scale Hadoop and Accumulo environments.
  • Administer and maintain distributed storage systems.
  • Engineer and improve monitoring, observability, and alerting frameworks.
  • Participate in system architecture and engineering design discussions.
  • Create and maintain automation scripts and Infrastructure as Code (IaC) deployments.
  • Patch, upgrade, configure, and harden systems in accordance with applicable security and compliance standards.
  • Administer LDAP-based user and group accounts.
  • Maintain hardware inventory and asset tracking.
  • Provide after-hours on-call support in a mission-driven operational environment.
  • Interface and collaborate with hardware, networking, infrastructure, software, and security teams.

Benefits

  • The company automatically contributes 20% of each employee's gross compensation to a Simplified Employee Pension (SEP) IRA, with no requirement for employee matching.
  • All contributions are fully vested from day one, ensuring immediate ownership of retirement funds.
  • Wyetech provides a generous PTO plan of up to 200 hours annually, aligned with applicable state leave regulations.
  • Employees have the flexibility to adjust their PTO allocation at the start of each calendar year, ensuring it meets their evolving needs.
  • A Choice of Medical Plan Options, some with Health Savings Account (HSA)
  • Vision and Dental
  • Life and AD&D Benefits
  • Short and Long-Term Disability
  • Hospital Indemnity, Accident, and Critical Illness Insurances
  • Optional Identity Theft and Legal Protection Services
  • Employee Referral Bonus Eligibility up to $10,000
  • Mobility Among Wyetech-supported Contracts
  • Various contract and work locations throughout Maryland, Virginia, Colorado, Texas, Utah, Alaska, Hawaii and OCONUS
  • Various team-building events throughout the year such as: monthly lunches, summer company picnic, and an annual holiday party.
  • Employees receive two complimentary branded clothing orders annually.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service