Infrastructure Engineer III

JPMorgan Chase & Co.Plano, TX

About The Position

As an Infrastructure Engineer III at JPMorganChase within the Corporate Sector – Infrastructure Platforms, you utilize strong knowledge of software, applications, and technical processes within the infrastructure engineering discipline. Apply your technical knowledge and problem-solving methodologies across multiple applications of moderate scope. You belong to the top echelon of talent in your field. At one of the world's most iconic financial institutions, where infrastructure is of paramount importance, you can play a pivotal role.

Requirements

  • Formal training or certification on infrastructure engineering concepts and 3+ years applied experience
  • Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support infrastructure engineering workflows with strong validation habits and awareness of data sensitivity.
  • Ability to review and validate AI-assisted recommendations before implementation, escalating when uncertain and ensuring outcomes align to resiliency, security, and auditability expectations.
  • Strong knowledge of one or more scripting languages
  • Experience with multiple cloud technologies (ability to operate in and migrate across public and private clouds)
  • Proficiency in scripting/automation and infrastructure-as-code (Python, PowerShell, Ansible, Terraform) and fluency in at least one programming language (e.g., Python, GO, Shell Scripting, .NET, etc.), including hands-on experience applying AI-assisted automation and agentic patterns to engineering/operational workflows.
  • Proven hands-on ability to leverage GitHub Copilot and coding assistants to accelerate skill development and build AI agents/workflows that support specific infrastructure operations tasks (e.g., triage, automation, runbook execution)
  • Knowledge of infrastructure engineering areas such as operating systems (Linux/Windows), networking terminology and protocols, databases, deployment practices, and automation.
  • Familiarity with Web products like Apache, Tomcat, IIS, WebSphere/IHS.
  • Deep proficiency in reliability, scalability, performance, security, enterprise system architecture, toil reduction, and similar best practices, with demonstrated ability to implement these practices within a platform.
  • Advanced knowledge and experience in observability, monitoring, alerting, and telemetry collection using tools such as Grafana, Dynatrace, Datadog, Prometheus, Cloudwatch, Splunk, etc., including designing and implementing monitoring dashboards using Splunk or Dynatrace with effective production support.

Nice To Haves

  • Implementation of CI/CD pipelines, code reviews using GitHub, and process automation with Python and scripting.
  • Hands-on experience and certifications in AWS, Azure, GCP, or other cloud environments; AWS/Azure exposure with understanding of resiliency, scalability, observability, and monitoring.
  • Experience utilizing Terraform or other infrastructure-as-code technologies for cloud resource management.
  • Experience supporting complex and mission critical applications involving a multitude of components of varying technical generations; Familiarity with modern front-end technologies and experience in designing, developing, and implementing software solutions.
  • Strong drive to expand infrastructure engineering knowledge across emerging technologies and domains; drive to self-educate and evaluate new technology.

Responsibilities

  • Applies technical knowledge and problem-solving methodologies to projects of moderate scope, with a focus on improving data and systems running at scale, and ensures end to end monitoring of applications
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate monitoring and capacity analysis and documentation, validating outputs and handling operational data according to sensitivity and security requirements.
  • Accountable for making significant decisions for a project consisting of multiple technologies and applications
  • Applies reuse-first, AI-assisted approaches to identify recurring capacity risks and improve remediation workflows, ensuring changes are validated and aligned to resiliency and security expectations.
  • Demonstrate an AI-first mindset by building hands-on agentic automation for operational workflows (e.g., triage, runbook execution, incident summarization, change validation) with appropriate controls and monitoring.
  • Demonstrate and champion site reliability culture and practices by exerting technical influence throughout your team; lead initiatives to improve reliability and stability using data-driven analytics to improve service levels.
  • Collaborate with team members, technical experts, key stakeholders, and customers to identify service level indicators and establish reasonable service level objectives and error budgets.
  • Diagnose and implement changes to resolve issues, modernize technology processes, and ensure end-to-end monitoring of applications.
  • Collect and analyze monitoring/telemetry data in test and production environments; design and implement dashboards; ensure system reliability, performance, and security.
  • Escalate issues appropriately with detailed technical write-ups; partner with application and infrastructure teams to identify and remediate capacity risks; understand platform interdependencies and limitations.

Benefits

  • comprehensive health care coverage
  • on-site health and wellness centers
  • a retirement savings plan
  • backup childcare
  • tuition reimbursement
  • mental health support
  • financial coaching
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service