-
port scientific applications to maximize performance across CPU, GPU, memory, storage, and I/O. Contribute technical expertise to faculty projects through the RCC Consultant Partnership Program and other
-
solutions on modern HPC and GPU-accelerated systems. This role includes supporting climate and geophysical science applications, enabling large-scale AI training and inference workflows, and contributing
-
, and administers CPU/GPU HPC clusters, including management and compute nodes, storage infrastructure, interconnects such as InfiniBand, and physical infrastructure in the datacenter and related systems
-
managing GPU-enabled infrastructure (NVIDIA GPUs, CUDA, multi-GPU systems) in cloud and/or on-prem environments. Familiarity with GPU orchestration in Kubernetes (e.g., NVIDIA device plugin, GPU scheduling
-
responsible for AI/ML research infrastructure, including managing and optimizing on-premises GPU resources and AWS cloud services such as Bedrock and SageMaker. Responsibilities include deploying, monitoring