-
of the group. Strong knowledge of performance modeling, simulation, and benchmarking of parallel and distributed computing systems and of the workflow systems that run on them. Familiarity with the FAIR data and
-
and cloud computing platforms. Formulating necessary solutions using various parallel computing paradigms and tools, HPC schedulers (such as slurm), Containers and Kubernetes, Python, Bash and other
-
for Science @ Scale: Pretraining, instruction tuning, continued pretraining, Mixture-of-Experts; distributed training/inference (FSDP, DeepSpeed, Megatron-LM, tensor/sequence parallelism); scalable evaluation
-
management of High-Performance Computing (HPC) systems within a classified environment. We are looking for candidates with experience in HPC architecture, cluster management, and parallel computing, with a
-
research teams to install, port, optimize, benchmark, and tune scientific applications and toolsets. Support software environments for a wide range of computational workflows, including parallel and
-
for Science @ Scale: Pretraining, instruction tuning, continued pretraining, Mixture-of-Experts; distributed training/inference (FSDP, DeepSpeed, Megatron-LM, tensor/sequence parallelism); scalable evaluation