-
Engineering at Scale : Profile, model, and optimize distributed codes using MPI, OpenMP, CUDA, ROCm, and related HPC technologies while bridging theoretical AI models with real hardware constraints. Cross
-
-squares solvers, or scalable tensor kernels. GPU performance engineering — CUDA/HIP kernel design, communication–computation trade-offs, or performance modeling on heterogeneous systems. Theory of
-
stack, lifecycle-managed nodes, TF2, rosbag workflows, and distributed system diagnostics. Experience with CUDA and deployment on NVIDIA edge-GPU platforms (Jetson Orin, Thor, or Spark). Experience
-
(Chroma, CPS, MILC, QUDA, Grid, etc.) is highly desirable. Familiarity with C/C++ and accelerator application programming models such as CUDA, HIP, SYCL, OpenMP, Kokkos or Raja, vectorization; MPI
-
models, or accelerators, including CPUs, GPUs, FPGAs, and emerging computing technologies. Familiarity with quantum software frameworks such as Qiskit, CUDA-Q, PennyLane, Cirq, or equivalent platforms