Principal, Mainframer Performance and Capacity

Fidelity InvestmentsWestlake, TX
Onsite

About The Position

The Fidelity Mainframe Capacity and Performance team is responsible for capacity modeling, forecasting, performance analysis, system configuration tuning, monitoring, and L2/L3 problem determination supporting Fidelity’s large complex IBM z Systems infrastructure. The team supports critical online and batch applications and works closely with business partners, Mainframe Hosting, CICS, MQ, WAS, DB2, Storage, Network, application teams, IBM, and other strategic vendors. This Grade 6 role is focused on hands-on performance analysis, problem determination, monitoring, reporting, and automation across CICS, MQ, WebSphere Application Server (WAS), Liberty Java Batch, Compute Grid, Kafka MQ connectors, and Operational Decision Manager (ODM). The associate will apply working knowledge of how applications use these middleware platforms to identify bottlenecks, explain workload behavior, support tuning, and improve operational readiness. The Associate will perform detailed capacity and performance engineering for mainframe middleware services. The role requires a practical understanding of CICS internal functions and application usage; MQ queue managers, CHINIT address spaces, queues, channels, and off-mainframe conversation partners; and WAS/Compute Grid components supporting online and batch workloads. The Associate will gather and analyze metrics from SMF/RMF, SYSVIEW, WAS PMI, IntelliMagic, SAS/MXG, Prometheus, VictoriaMetrics, PostgreSQL, Grafana, and Datadog. The role will create repeatable reports and dashboards, support incident and change analysis, assist with stress-test and modernization measurement, and contribute to automation that reduces manual effort. This is a hands-on individual-contributor role. The Principal works under the direction of senior technical leaders while independently owning defined analyses, reports, dashboards, scripts, operational improvements, and problem investigations. The role is expected to document findings clearly, collaborate across teams, and build deeper expertise over time.

Requirements

  • Hands-on experience supporting or analyzing one or more IBM mainframe middleware domains such as CICS, MQ, WAS, Liberty, Compute Grid, Kafka integration, or ODM.
  • Working knowledge of z/OS, SMF/RMF data, GP/zIIP use, memory, Coupling Facility, zFS, log streams, and subsystem interactions.
  • Experience with performance monitoring, problem determination, capacity reporting, application support, systems programming, or production engineering in a complex enterprise environment.
  • Experience using one or more analysis platforms such as SYSVIEW, IntelliMagic, OMEGAMON, SAS/MXG, Grafana, Datadog, Prometheus, VictoriaMetrics, PostgreSQL, or equivalent tools.
  • Ability to read technical metrics, isolate likely causes, test assumptions, and communicate findings with senior engineers and application partners.
  • Programming or scripting experience in SAS, REXX, Python, Java, Assembler, Bash, or a comparable language is preferred.
  • Experience with dump analysis, JVM diagnostics, IPCS, or subsystem traces is preferred but may be developed in the role.
  • Analytical problem-solving skills and curiosity about how middleware products and application workloads behave under normal and stressed conditions.
  • Ability to connect CICS, MQ, WAS, Kafka, ODM, z/OS, network, storage, and application signals during problem determination.
  • Clear written and verbal communication, including the ability to summarize technical findings, decisions, risks, and next steps.
  • Discipline in building repeatable reports, scripts, dashboards, procedures, and technical documentation.
  • Collaborative working style with application, infrastructure, operations, architecture, and vendor teams.
  • Willingness to learn from senior SMEs, take ownership of defined work, and progressively deepen expertise across multiple middleware domains.
  • Practical engineering judgment with attention to stability, resiliency, performance, operational simplicity, and cost.

Nice To Haves

  • Assembler, Java, Bash, SYSVIEW Report Writer, IPCS, CICS/MQ and Java dump analysis, Fidelity C2C environment

Responsibilities

  • Analyze CICS performance and behavior, including internal functions, region activity, transactions, programs, files, queues, dispatching, thresholds, and how application programs use CICS services.
  • Analyze IBM MQ behavior across queue managers, CHINIT address spaces, queues, channels, and connections to distributed or off-mainframe conversation partners; identify tuning and problem-determination opportunities.
  • Support WAS on z/OS, Liberty, Java Batch (JSR 352), and Compute Grid workloads, including PJM, TopJobs/SubJobs, online and batch activity, and WAS PMI metric collection and interpretation.
  • Support performance analysis of Kafka MQ connectors, brokers, channels, message flow, and tunable settings on the MQ and Kafka sides.
  • Support ODM mainframe components by understanding business-rule invocation from WAS, CICS, and batch, and identifying what can be monitored and tuned.
  • Apply working knowledge of z/OS components used by middleware platforms, including Coupling Facility services, real memory, zFS file systems, log streams, GP engines, and zIIP engines.
  • Create and run SAS programs for performance reporting and use MXG fields associated with CICS, MQ, WAS, z/OS, and other IBM product areas.
  • Create REXX programs and scripts to gather performance metrics not available through standard SMF reporting.
  • Develop and maintain Grafana dashboards and support metric collection using Python and Prometheus with VictoriaMetrics or PostgreSQL data stores and Datadog reporting where applicable.
  • Use SYSVIEW capabilities including CTRANLOG, thresholds, alerting, administration, and Report Writer to investigate and report subsystem behavior.
  • Use IntelliMagic for navigation, analysis, trending, and visualization of available mainframe performance and capacity data.
  • Analyze Java and JVM dumps for WAS-related problems and, with guidance, use IPCS for CICS and MQ problem determination.
  • Document analysis methods, assumptions, findings, recommendations, scripts, dashboards, and operational procedures to improve team reuse and backup coverage.
  • Participate in production support, incident analysis, change validation, stress testing, and after-hours support as required for critical services.
© 2026 Teal Labs, Inc
Privacy PolicyTerms of Service