Sort by
Refine Your Search
-
Category
-
Field
-
management of HPC systems within a classified environment. We are looking for candidates with experience in HPC architecture, cluster management, and parallel computing, with a proven ability to work within
-
discretization techniques for unstructured meshes and/or finite elements with an emphasis on highly scalable algorithms for exascale HPC environments Experience with parallel computing environments, HPC in a Linux
-
discretization techniques for unstructured meshes and/or finite elements with an emphasis on highly scalable algorithms for exascale HPC environments Experience with parallel computing environments, HPC in a Linux
-
discretization techniques for unstructured meshes and/or finite elements with an emphasis on highly scalable algorithms for exascale HPC environments Experience with parallel computing environments, HPC in a Linux
-
well as experience with HPC environments and parallel computing. Demonstrated hands-on experience and understanding of developing scientific data management, workflows and resource management problems. Strong problem
-
Language Models (LLMs). Distributed Machine Learning: Specialization in data parallelism, model-parallelism, and collective communication strategies in large-scale environments. Proficiency in frameworks
-
massively parallel algorithms and code performance profiling are a plus. Special Requirements: Applicants cannot have received their Ph.D. more than five years prior to the date of application and must
-
management of High-Performance Computing (HPC) systems within a classified environment. We are looking for candidates with experience in HPC architecture, cluster management, and parallel computing, with a
-
or model parallel training. Experience with multi-physics simulations on HPC and with ML models. Experience working in a multi-disciplinary research environment. Demonstrated written and oral
-
for Science @ Scale: Pretraining, instruction tuning, continued pretraining, Mixture-of-Experts; distributed training/inference (FSDP, DeepSpeed, Megatron-LM, tensor/sequence parallelism); scalable evaluation