Table of Contents
Fetching ...

Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud

Eric Ding

TL;DR

Propius tackles the scalability and heterogeneity challenges of collaborative ML across edge devices by introducing a two-plane platform: a control plane for scalable, multi-tenant resource sharing and a data plane for efficient plan distribution and result aggregation. The control plane uses a soft-state scheduler with online and small-batch modes to manage transient edge resources, while the data plane supports FedAvg-like algorithms and integrates with CDNs for scalable content delivery. Empirical results show improvements up to $1.88\times$ in resource utilization, $2.76\times$ in throughput, and $1.26\times$ faster job completion relative to baselines, demonstrating practical benefits for edge-to-cloud collaborative learning. Overall, Propius provides a scalable, adaptable infrastructure that abstracts heterogeneity from developers and enables efficient, multi-tenant collaborative ML at scale.

Abstract

Collaborative Machine Learning is a paradigm in the field of distributed machine learning, designed to address the challenges of data privacy, communication overhead, and model heterogeneity. There have been significant advancements in optimization and communication algorithm design and ML hardware that enables fair, efficient and secure collaborative ML training. However, less emphasis is put on collaborative ML infrastructure development. Developers and researchers often build server-client systems for a specific collaborative ML use case, which is not scalable and reusable. As the scale of collaborative ML grows, the need for a scalable, efficient, and ideally multi-tenant resource management system becomes more pressing. We propose a novel system, Propius, that can adapt to the heterogeneity of client machines, and efficiently manage and control the computation flow between ML jobs and edge resources in a scalable fashion. Propius is comprised of a control plane and a data plane. The control plane enables efficient resource sharing among multiple collaborative ML jobs and supports various resource sharing policies, while the data plane improves the scalability of collaborative ML model sharing and result collection. Evaluations show that Propius outperforms existing resource management techniques and frameworks in terms of resource utilization (up to $1.88\times$), throughput (up to $2.76$), and job completion time (up to $1.26\times$).

Propius: A Platform for Collaborative Machine Learning across the Edge and the Cloud

TL;DR

Propius tackles the scalability and heterogeneity challenges of collaborative ML across edge devices by introducing a two-plane platform: a control plane for scalable, multi-tenant resource sharing and a data plane for efficient plan distribution and result aggregation. The control plane uses a soft-state scheduler with online and small-batch modes to manage transient edge resources, while the data plane supports FedAvg-like algorithms and integrates with CDNs for scalable content delivery. Empirical results show improvements up to in resource utilization, in throughput, and faster job completion relative to baselines, demonstrating practical benefits for edge-to-cloud collaborative learning. Overall, Propius provides a scalable, adaptable infrastructure that abstracts heterogeneity from developers and enables efficient, multi-tenant collaborative ML at scale.

Abstract

Collaborative Machine Learning is a paradigm in the field of distributed machine learning, designed to address the challenges of data privacy, communication overhead, and model heterogeneity. There have been significant advancements in optimization and communication algorithm design and ML hardware that enables fair, efficient and secure collaborative ML training. However, less emphasis is put on collaborative ML infrastructure development. Developers and researchers often build server-client systems for a specific collaborative ML use case, which is not scalable and reusable. As the scale of collaborative ML grows, the need for a scalable, efficient, and ideally multi-tenant resource management system becomes more pressing. We propose a novel system, Propius, that can adapt to the heterogeneity of client machines, and efficiently manage and control the computation flow between ML jobs and edge resources in a scalable fashion. Propius is comprised of a control plane and a data plane. The control plane enables efficient resource sharing among multiple collaborative ML jobs and supports various resource sharing policies, while the data plane improves the scalability of collaborative ML model sharing and result collection. Evaluations show that Propius outperforms existing resource management techniques and frameworks in terms of resource utilization (up to ), throughput (up to ), and job completion time (up to ).
Paper Structure (14 sections, 4 figures, 3 tables)

This paper contains 14 sections, 4 figures, 3 tables.

Figures (4)

  • Figure 1: Propius Architecture
  • Figure 2: Resource utilization of different scheduling method. Resource utilization is defined as the average client partitipation time (contributing to a job) divided by the overall job completion time.
  • Figure 3: Binding throughput of different scheduling methods. Throughput is defined as the number of client-job binding per second.
  • Figure 4: Average job completion time of different scheduling methods