-
and in computer architecture domains, our ability to analyze data still falls behind the unstoppable data collection rates. Data-intensive applications are increasingly more demanding in sophisticated
-
. Scalable AI training: distributed training, efficient fine-tuning, evaluation, and deployment on large-scale compute. Multimodal learning: representation learning and generative modeling across heterogeneous
-
dependencies for execution on Alps / CSCS infrastructure. Run and monitor Slurm-based training and evaluation jobs. Debug failures related to distributed execution, checkpointing, filesystem performance
Searches related to distributed computing
Enter an email to receive alerts for distributed-computing positions