Mid Systems Engineer (w/ active Secret)

Critical SolutionsNorfolk, VA
$85,000 - $110,000Hybrid

About The Position

This role involves working closely with development and operations teams to ensure speedy and reliable software deployments, monitor systems, and enhance platform reliability. The engineer will be responsible for identifying and fixing system bugs, developing features using AI coding tools and script repositories for automation, scaling, testing, and securing cloud infrastructure and pipelines. A key aspect of the role is enhancing performance monitoring through tools like Splunk, identifying and optimizing performance bottlenecks, and contributing to the SRE journey by improving engineering build, maintenance, automation, and reliability using SRE/DevOps tools and Infrastructure-as-Code. The position requires developing and coding high-quality pipeline automation workflows, creating and executing test strategies for failure scenarios, and building automated systems for continuous performance, stress, and load testing. Collaboration with SREs, developers, and operations teams to define reliability goals and testing strategies is essential. The role also includes ensuring new services and features are thoroughly tested before production, validating monitoring, logging, and alerting mechanisms, and ensuring accurate measurement and tracking of SLIs and SLOs. The engineer will independently resolve most conflicts between timeline, budget, and scope, escalating complex issues to senior management.

Requirements

  • Currently possess and ability to maintain Secret clearance
  • Requires BS degree and 5-10 years of prior relevant experience or Master's with 4-8 years of prior relevant experience.
  • Minimum of DoD 8570.01 IAT Level II Certification required prior to onboarding and must maintain certification while supporting the Contract
  • Must be able to support program execution in classified environments and access SIPRNet from an Agency location on short notice (local travel).
  • Experience with automated script design, coding, debugging, and maintenance skills (using bash, python, etc.) preferred
  • Experience designing, deploying, and maintaining Microsoft Endpoint Configuration Manager (MECM) environments at enterprise scale.
  • Proven ability to optimize MECM infrastructure for performance, reliability, and high availability.
  • Hands-on experience creating and managing MECM device collections, task sequences, applications, packages, and operating system deployment images.
  • Expertise in automating MECM administrative tasks using PowerShell, APIs, and infrastructure-as-code tooling.
  • Strong background in Windows endpoint management, including patching strategies, compliance baselines, configuration items, and reporting.
  • Demonstrated ability to improve endpoint reliability through monitoring, observability tooling, and proactive remediation.
  • Experience integrating MECM with cloud-based services such as Intune, Azure AD, and Windows Update for Business.
  • Ability to troubleshoot complex MECM client and server issues using logs, diagnostics, and platform health data.
  • Experience supporting enterprise-wide operating system deployments and in-place upgrades using MECM automation pipelines.
  • Knowledge of SQL Server administration relevant to maintaining MECM databases and enabling performance tuning.
  • Experience contributing to system reliability through continuous improvement, infrastructure monitoring, and configuration drift reduction.
  • Strong understanding of networking concepts supporting MECM components such as DP distribution, boundary groups, and PXE services.
  • Ability to work in a highly collaborative, forward thinking, and innovation-driven environment.
  • Knowledge of Agile and DevSecOps/SRE concepts and best practices, with a desire to grow that knowledge.
  • Hand-on experience with Atlassian products (Jira, Confluence, Bitbucket, etc.).
  • Experience creating JIRA and/or Azure DevOps workflows, projects, custom configurations.
  • Experience administrating/maintaining SRE platform via Ansible playbooks (e.g. upgrading Jenkins).
  • Experience in automating tasks with scripting languages like PowerShell, or Python.
  • Integrating/maintaining with various 3rd party CI/CD tools like Jenkins and Gitlab.
  • Experience with commercial cloud infrastructure deployment environments such as AWS and Azure.
  • Experience with automated provisioning and configuration tools like Terraform, Cloud Formation, Chef, Puppet, Ansible, or similar technologies.
  • Working knowledge of the Risk Management Framework (RMF), DISA STIGs.

Nice To Haves

  • Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, or CloudFormation for automating test environments.
  • ITILv4, Scrum Master, or Agile SAFe certification(s) or applicable experience.

Responsibilities

  • Work alongside the development and operations teams to ensure speedy and reliable software deployments, monitor systems, and improve overall reliability of the platform.
  • Discover and document system bugs and fix them.
  • Develop features utilize the AI coding tool and repository of scripts to automate, scale, test, and secure the cloud infrastructure and the pipelines.
  • Enhance performance monitoring of the various systems via Splunk or other dashboard reporting tools.
  • Identify performance bottlenecks and optimize the performance of cloud infrastructure.
  • Contribute to continuing our SRE journey by suggesting ways to improve engineering build, maintenance, automation and reliability across the platform with SRE/DevOps tools and Infrastructure-as-Code.
  • Develop and code high-quality pipeline automation workflows to support inside and outside the cloud platform that are appropriate for business and technology strategies.
  • Develop and execute test strategies that simulate real-world failure scenarios, including network disruptions, hardware failures, and system overloads.
  • Create, script, and run performance tests to measure system behavior under varying levels of load and traffic.
  • Identify bottlenecks, performance degradation, and areas for optimization.
  • Design, implement, and maintain automated test suites for infrastructure and application components.
  • Ensure that testing is integrated into the CI/CD pipeline to validate system reliability with every release.
  • Build automated systems for continuous performance testing, stress testing, and load testing.
  • Work closely with SREs, developers, and operations teams to define reliability goals and develop appropriate testing strategies to validate those goals.
  • Ensure that new services and features undergo thorough testing for performance, reliability, and failure recovery before deployment to production.
  • Validate that monitoring, logging, and alerting mechanisms are functioning correctly by testing systems under failure conditions.
  • Ensure that Service Level Indicators (SLIs) and Service Level Objectives (SLOs) are accurately measured and tracked through automated testing frameworks.
  • Resolve most conflicts between timeline, budget, and scope independently but intuitively raise sophisticated or consequential issues to senior management.

Benefits

  • 100% premium coverage for Medical, Dental, Vision, and Life Insurance
  • Supplemental Insurance
  • 401K matching
  • Flexible Time Off (PTO/Holidays)
  • Higher Education/Training Reimbursement
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service