DevOps Engineer

Re:Build Manufacturing•Framingham, MA
•Onsite

About The Position

This engineer supports deployments and keeps production running in Cadonix’s controlled AWS environment, including AWS GovCloud (US), where only US Persons are permitted under ITAR. Product teams manage their own continuous integration and delivery workflows and publish artifacts. This role moves those artifacts into the restricted production environment. It verifies them and resolves infrastructure problems from new deployments or spontaneous issues in running systems. Because access to this environment is limited, the role is often the only person able to see what is actually happening in production.

Requirements

  • A bachelor’s degree or equivalent experience in Computer Science, Software Engineering, or a related field.
  • Minimum of 5 years’ experience in DevOps, cloud infrastructure or production operations, including at least 3 years supporting production workloads on AWS.
  • Practical, hands-on experience diagnosing real production problems on AWS: EC2, ECS or EKS, VPC network architecture and security configurations, ALB/NLB, Route 53, ACM, IAM policy evaluation, S3, RDS, and CloudWatch.
  • Capable of reading CloudTrail and CloudWatch logs to reconstruct what happened.
  • Understand container images, artifact versioning, environment configuration and secrets injection, and rollback.
  • Comfortable working with CI/CD pipelines owned by other teams and diagnosing where a handoff has gone wrong.
  • Able to read, modify, and safely apply existing Terraform, CDK, or CloudFormation. Understands state, drift, and why manual console changes cause problems later.
  • Solid Linux administration and fixing.
  • Working understanding of TCP/IP, DNS, TLS, routing, and firewall behaviour — enough to tell a network problem from an application problem.
  • Proficient in Bash and Python (or Go) for automation, diagnostics, and small tooling.
  • CloudWatch, Prometheus, Grafana, or similar. Able to build a dashboard or alert that answers a real operational question.
  • Comfortable supporting software they did not write and may not have source access to, working from telemetry, timing, and behaviour rather than code.
  • Persistent and methodical when the obvious answer is not available.
  • Good judgement about what constitutes sensitive or client-related information, including inside logs, stack traces, configuration, and crash artifacts.
  • Holds the boundary under production pressure and advances rather than improvising.
  • Keeps collaborators informed during an incident without being asked.
  • Works effectively alongside product groups across time zones, and with external vendors under support contracts.
  • Reliable and self-directed, since much of the environment cannot be observed by anyone else.

Nice To Haves

  • AWS GovCloud experience is a plus but not required; equivalent experience in isolated or regulated environments is equally relevant.
  • Does not need to have designed a large estate from scratch.
  • Familiarity with ITAR, CUI, FedRAMP, PCI-DSS, or HIPAA environments is an advantage.

Responsibilities

  • Assist with and carry out releases of product team materials into the regulated production environment.
  • Perform pre-deployment verification, run the deployment, check service stability afterward, and complete rollback when a release does not behave as encouraged.
  • Collaborate with product teams to address deployment failures, providing feedback on environmental needs so their pipelines generate artifacts that deploy efficiently.
  • Investigate and resolve production infrastructure issues across AWS platforms, including compute, container services, networking and security groups, load balancers, DNS and certificates, IAM permissions, storage, and managed databases.
  • Address problems that arise immediately after deployment. Also handle those that occur unexpectedly in a stable system.
  • Determine the root cause instead of just restoring service.
  • Coordinate the stability of production systems.
  • Coordinate the management of alerting, dashboards, and on-call rotation.
  • Write and maintain runbooks to make recurring issues standard procedure.
  • Preserve the environment’s integrity and timeliness: operating system and foundational image updates, certificate renewal, backup and restore verification, capacity headroom, and cost monitoring.
  • Implement modifications to existing infrastructure using infrastructure as code and version control instead of manual console edits.
  • Work in an environment where data remains secure and only authorized US Persons have entry.
  • Handle records, diagnostic information, and support artifacts in line with export-control and customer data-handling requirements.
  • When a problem requires help from an external software vendor, develop sanitized evidence and reproductions. This allows them to assist without accessing controlled or customer data.
  • Raise the issue through accurate channels instead of bypassing boundaries.
  • Record environment configuration, deployment procedures, and incident history to ensure operational knowledge is accessible beyond a single individual.

Benefits

  • Every employee of Re:Build will share ownership in the company and will share in the financial rewards of the success we achieve together, at all levels of the company!
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service