-
compressed into lightweight student models using knowledge distillation, enabling efficient real-time inference on mobile devices. The distilled models will be deployed and optimized on mobile platforms, with
-
-time, on-device applications. This project focuses on exploring model optimization strategies, including compression, quantization, and efficient architecture design, to reduce the resource footprint
-
that both parameter estimation and model selection can be interpreted as problems of data compression. The principle is simple: if we can compress data, we have learned something about its underlying
-
to cloud-based machine learning services, on-device ML is privacy-friendly, of low latency, and can work offline. User data will remain at the mobile device for ML inference. Problems: In order to enable