Senior Cloud Platform Engineer

CyberData TechnologiesBethesda, MD

About The Position

CyberData Technologies is seeking a Senior Cloud Platform Engineer to support a complex NHGRI environment spanning Amazon Web Services, Microsoft Azure, and Google Cloud Platform. This is a hands-on engineering role focused on multi-cloud FinOps, cloud automation, Infrastructure as Code, and secure AI-platform operations. The engineer will operate and optimize cloud services, implement financial-governance and compliance controls, maintain institutional AI platforms, support research billing arrangements, and translate usage, cost, security, and performance data into actionable recommendations for technical and executive stakeholders.

Requirements

  • Bachelor’s degree in computer science, information technology, engineering, data analytics, or a related technical field, or equivalent relevant experience.
  • Seven or more years of relevant cloud engineering, platform operations, DevOps, FinOps, or systems-engineering experience.
  • Hands-on experience operating production environments in at least one major cloud provider—AWS, Azure, or GCP—with practical engineering experience in at least one additional cloud platform.
  • Experience managing cloud costs, budgets, forecasts, tagging, chargeback, showback, or optimization activities.
  • Experience using a cloud financial-management, governance, or FinOps platform.
  • Hands-on experience with Terraform, Infrastructure as Code, and CI/CD pipelines.
  • Experience developing cloud automation using scripting, APIs, serverless functions, or workflow tools.
  • Experience operating an AI platform, API gateway, developer platform, or other multi-tenant cloud service.
  • Understanding of cloud identity, networking, private endpoints, logging, monitoring, and security controls.
  • Ability to analyze technical and financial data and present clear recommendations to technical, business, and executive stakeholders.
  • Experience preparing recurring cost, usage, performance, or risk reporting for leadership and security forums.
  • Strong troubleshooting, technical-writing, and communication skills.
  • Ability to obtain and maintain required HHS/NIH suitability and system access.

Nice To Haves

  • Prior NHGRI, NIH, HHS, federal health, or regulated federal-cloud experience.
  • Direct experience with Kion, CloudHealth, Apptio Cloudability, AWS Cost Explorer, Azure Cost Management, or GCP Cloud Billing.
  • Experience with LibreChat, LiteLLM, AWS Bedrock, Azure OpenAI, or GCP Vertex AI.
  • Experience managing model gateways, API keys, quotas, rate limits, model versions, endpoint availability, and provider routing.
  • Experience implementing token-level AI usage and cost attribution.
  • Experience with GitHub Enterprise Cloud and automated deployment pipelines.
  • Experience with AWS Lambda, Azure Functions, Cloud Custodian, or comparable automation technologies.
  • Experience with Microsoft Entra ID, Azure Privileged Identity Management, AWS PrivateLink, or Azure ExpressRoute.
  • Experience integrating cloud and AI telemetry with Splunk or OpenSearch.
  • Familiarity with the NIST AI Risk Management Framework, federal AI governance, FISMA, and cloud-authorization requirements.
  • Experience developing governed agentic AI workflows or LLM-based automation.
  • Experience supporting scientific, biomedical, genomic, or research-computing environments.
  • Familiarity with Terra billing projects, NIH STRIDES, AnVIL, NIH Data Commons, or the All of Us Research Program Workbench.

Responsibilities

  • Administer multi-cloud FinOps processes across AWS, Azure, and GCP.
  • Configure and maintain Kion or comparable cloud-governance and financial-management platforms.
  • Implement budgets, alerts, spend thresholds, enforcement actions, tagging, and cost-allocation methods, including tiered passive and active spend-enforcement mechanisms.
  • Implement configuration-compliance checks and automated remediation using cloud-governance platform features such as Cloud Custodian actions.
  • Analyze cloud consumption, identify optimization opportunities, and support forecasting and right-sizing.
  • Maintain a current inventory of cloud accounts, subscriptions, projects, funding sources, credits, and research billing arrangements across AWS, Azure, and GCP.
  • Support research investigators and laboratories in configuring Terra and other research billing projects across cloud credits, grant funding, and laboratory funding sources in coordination with NIH CIT and applicable NHGRI research program offices.
  • Deliver biweekly and monthly cost and usage reporting at the account and service levels.
  • Support recurring Cloud Optimization Assessments and executive-level cost reporting.
  • Develop and maintain Terraform configurations, GitHub pipelines, and cloud-automation workflows.
  • Operate and maintain AI platforms using LibreChat, LiteLLM, or comparable model-gateway technologies.
  • Integrate and manage services such as AWS Bedrock, Azure OpenAI, and GCP Vertex AI.
  • Manage model endpoints, versions, quotas, rate limits, provider routing, and service availability.
  • Implement identity, access, private connectivity, logging, monitoring, and security controls for AI services, including time-bound privileged-access patterns using technologies such as Microsoft Entra ID and Azure Privileged Identity Management.
  • Track AI usage and token-level costs by model, provider, user group, or organizational unit.
  • Perform token-cost analysis to support budget forecasting and cost allocation.
  • Develop and operationalize governed agentic AI workflows for approved IT and research use cases, including appropriate access, logging, cost, and human-review controls.
  • Integrate cloud and AI telemetry with Splunk, OpenSearch, Kion, or comparable monitoring and reporting platforms.
  • Coordinate with enterprise network, subscription, identity, and security teams on platform connectivity, subscription policies, and identity-provider requirements.
  • Troubleshoot cloud-platform, automation, integration, performance, and service-availability issues.
  • Develop operational runbooks, SOPs, dashboards, technical documentation, and knowledge articles.
  • Deliver knowledge-sharing sessions and direct stakeholder support on AI capabilities, security considerations, and lessons learned.
  • Provide limited after-hours support for maintenance, incidents, or platform-recovery activities.

Benefits

  • FinOps Certified Practitioner or FinOps Certified Professional
  • HashiCorp Certified: Terraform Associate
  • AWS Certified Solutions Architect – Associate or AWS Certified CloudOps Engineer – Associate
  • Microsoft Certified: Azure Administrator Associate
  • Google Associate Cloud Engineer
  • AWS Certified Machine Learning Engineer – Associate, Microsoft Certified: Azure AI Apps and Agents Developer Associate, or Google Professional Machine Learning Engineer
  • Splunk Core Certified Advanced Power User or Splunk Cloud Certified Admin
  • Other relevant FinOps, DevOps, AI-platform, cloud-operations, or cloud-security certification
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service