Bandwidth Allocation with Device Partitioning for Federated Learning over Industrial IoT networks


Abstract

We consider a federated learning (FL) system in which Industrial Internet-of-Things (IIoT) devices collaboratively train a global model over wireless channels without sharing local data. In such systems, communication time is a primary bottleneck that constrains overall training efficiency. Unlike conventional networks that prioritize individual quality-of-service requirements, FL systems collectively aim to converge to an optimal global model as efficiently as possible, which calls for a fundamentally different approach to bandwidth allocation. In this paper, we propose a novel bandwidth allocation policy that exploits the heterogeneity of device computing capabilities to minimize total training time. Rather than distributing bandwidth among all selected devices simultaneously, the proposed policy partitions the participating devices into ordered subsets and sequentially grants each subset exclusive access to the full bandwidth. We formally prove that this partitioning-based policy achieves a strictly lower training time than any bandwidth allocation scheme without partitioning, irrespective of the underlying scheduling algorithm. Furthermore, by reducing per-device transmission duration, the proposed policy also minimizes uplink energy consumption, which is particularly beneficial for battery-constrained IIoT devices. Extensive experiments on real-world datasets — including GC10-Det, an industrial surface defect benchmark, and CIFAR-10, a standard image classification benchmark — demonstrate that the proposed policy consistently reduces training time and energy consumption compared to existing bandwidth allocation schemes, approaching the theoretical lower bound on round time.

Bandwidth allocation, Device heterogeneity, Energy efficiency, Federated learning, Industrial Internet-of-Things

1 Introduction↩︎

The rapid evolution of Industry 4.0 and the proliferation of Industrial Internet-of-Things (IIoT) devices have generated an unprecedented volume of operational data. Leveraging artificial intelligence (AI), this data has become a critical resource for optimizing manufacturing processes, enabling predictive maintenance, and enhancing quality control [1]. Nevertheless, the sheer scale of locally generated data, combined with stringent requirements for protecting proprietary industrial information, poses fundamental challenges to conventional machine learning approaches that rely on centralizing raw data for model training.

Federated learning (FL) has emerged as a promising paradigm to address these challenges in industrial settings [2]. Rather than transferring raw data to a central server, FL enables IIoT devices to perform local model training and transmit only model updates. This collaborative framework allows multiple devices to jointly train a shared global model while preserving the data privacy of individual manufacturing units.

Despite its privacy-preserving merits, FL inherently requires frequent communication between participating devices and the central server. In industrial environments, which often span complex, multi-domain networks with severely constrained wireless resources, the repeated transmission of model updates constitutes a major deployment bottleneck [3]. A substantial body of research has addressed this challenge from an algorithmic perspective, proposing strategies such as model compression [4], reduced aggregation frequency [5], local update momentum [6], and adaptive aggregation weight allocation [7].

Complementary to algorithmic approaches, communication efficiency can be fundamentally improved through the judicious allocation of physical resources such as bandwidth and power. While extensive resource allocation strategies have been developed for conventional communication systems [8][10], these are ill-suited for FL due to a fundamental difference in objective. Traditional networks optimize for stable, high data rates to facilitate rapid individual message delivery. In contrast, FL systems do not require each device to communicate as quickly as possible; rather, their collective goal is to converge to an optimal global model rapidly and efficiently. This distinction motivates the need for FL-specific resource management strategies tailored to industrial settings.

Accordingly, a growing body of research has investigated communication resource allocation over wireless channels to enhance FL efficiency. In [11], a resource allocation framework was proposed to minimize training latency in IIoT environments. To mitigate delays caused by stragglers, the authors introduced a greedy client sampling algorithm that considers both channel conditions and data quality. Similarly, the work in [12] jointly optimized bandwidth allocation and device energy budgets to minimize overall consumption. Recognizing that IIoT devices often operate under stringent energy constraints, the authors proposed a strategy that selects clients and allocates bandwidth based on real-time battery levels and channel quality.

In [13], the joint client selection and bandwidth allocation problem was formulated where clients providing high-quality updates are prioritized. Furthermore, [14] proposed strategies to maximize the number of devices capable of satisfying a predefined time constraint without requiring prior channel state information. This was achieved through a combinatorial multi-armed bandit algorithm and a bandwidth allocation strategy analogous to the water-filling algorithm. Finally, [15] investigated the joint optimization of client selection and bandwidth allocation under strict round deadlines, exploring the fundamental tradeoff between the number of participating clients and the allocated bandwidth per device.

Several studies [16][18] pursued bandwidth allocation strategies that ensure all devices complete communication simultaneously. In particular, [16] assumed synchronous initiation; thus, bandwidth is allocated to have the same transmission rate across selected devices. In addition, [17] relaxed this to asynchronous initiation and established optimality when transmission power scales linearly with bandwidth. Extending to fixed transmission power, [18] demonstrated optimality of bandwidth allocation which ensures the simultaneous completion of transmission.

Several studies [16][18] have pursued bandwidth allocation strategies aimed at ensuring all devices complete their transmission simultaneously. Specifically, the authors of [16] assumed synchronous initiation; consequently, bandwidth is allocated to maintain uniform transmission rates across all selected devices. This approach was further extended in [17], which relaxed the assumption to asynchronous initiation and established optimality under the condition that transmission power scales linearly with bandwidth. Addressing scenarios with fixed power budgets where power does not scale with bandwidth, [18] demonstrated that optimality is still achieved through bandwidth allocation strategies that enforce simultaneous transmission completion.

However, all of the above schemes [12][18] share two critical limitations that restrict their applicability to realistic IIoT deployments. First, they assume that all selected devices share the entire bandwidth simultaneously. In practice, IIoT environments involve heterogeneous devices — such as robotic arms, programmable logic controllers, and smart sensors — with widely varying local training times. Under simultaneous allocation, devices that finish training early must wait idly, incurring unnecessary time overhead. Alternatively, sequentially allocating the full bandwidth to subsets of devices can substantially reduce total communication time by exploiting this heterogeneity. Second, these works assume a fixed SNR per device, independent of bandwidth allocation. Under realistic fixed transmission power constraints, which arise from hardware limitations or safety regulations in industrial settings [19], noise power varies with allocated bandwidth, rendering SNR a function of bandwidth and significantly complicating the optimization problem.

To address these gaps, we propose a novel bandwidth allocation policy for IIoT-based FL that introduces device partitioning into the communication schedule. Our approach explicitly accounts for fixed transmission power constraints, under which SNR is treated as a function of bandwidth, faithfully reflecting practical industrial conditions. The proposed policy is agnostic to the underlying FL algorithm, making it compatible with methods addressing data heterogeneity [20], [21], model compression [4], [22], [23], and federated dropout [24], [25]. Furthermore, we formally prove that our partitioning-based policy outperforms the optimal simultaneous bandwidth allocation strategy, regardless of the client scheduling algorithm employed.

The main contributions of this paper are summarized as follows:

  • We investigate a practical bandwidth allocation problem for industrial FL under fixed device transmission power, faithfully capturing hardware constraints prevalent in IIoT deployments.

  • We propose a novel device partitioning policy in which heterogeneous devices are divided into sequential groups, each granted exclusive access to the full bandwidth. We formally prove that this policy outperforms any simultaneous allocation scheme, irrespective of the client scheduling algorithm.

  • We develop a low-complexity partitioning algorithm that achieves near-optimal performance while avoiding the prohibitive computational cost of exhaustive search.

  • We show that the proposed policy reduces energy consumption of IIoT devices as a byproduct of shorter transmission durations, which is particularly beneficial for energy-constrained devices in industrial deployments.

  • We conduct extensive experiments on real-world datasets, demonstrating significant reductions in FL round completion time and energy consumption across diverse bandwidth, computation load, and participation rate configurations.

The remainder of this paper is organized as follows. Section 2 describes the system model. Section 3 presents the bandwidth allocation problem formulation. Section 4 details the proposed policy. Section 5 reports experimental results on real-world datasets. Finally, Section 6 concludes the paper.

2 System Model↩︎

2.1 Federated Learning↩︎

Figure 1: Federated learning system over IIoT networks

We consider an FL system deployed over an IIoT network consisting of a central server and \(N\) heterogeneous IIoT devices, denoted by the set \(\mathcal{N} = \{1, \dots, N\}\) as illustrated in Fig. 1. Each device \(n \in \mathcal{N}\) holds a private local dataset \(\mathcal{D}_n\) of size \(D_n = |\mathcal{D}_n|\), and the global dataset is given by \(\mathcal{D} = \bigcup_{n=1}^N \mathcal{D}_n\).

The objective of FL is to collaboratively train a global model that minimizes the average loss over \(\mathcal{D}\) without exposing the local data of any device. Given a loss function \(f(\mathbf{w}; x)\) for model parameters \(\mathbf{w}\) and data sample \(x\), the global loss function is defined as \[\begin{align} F(\mathbf{w}) = \frac{1}{|\mathcal{D}|} \sum_{x \in \mathcal{D}} f(\mathbf{w}; x). \label{eq:g95loss} \end{align}\tag{1}\]

a

b

c

d

Figure 2: Procedures of federated learning.

To minimize \(F(\mathbf{w})\) while preserving data privacy, the FL system proceeds iteratively over multiple communication rounds through four stages: global model broadcasting, local updating, local model transmission, and global aggregation, as illustrated in Fig. 2.

At the start of round \(k\), the server selects a subset of devices \(\mathcal{N}_k \subseteq \mathcal{N}\) with \(|\mathcal{N}_k| = q N\), where \(q \in (0, 1]\) is the participation fraction, and broadcasts the current global model \(\overline{\mathbf{w}}_k\) to all selected devices. Each device \(n \in \mathcal{N}_k\) initializes its local model as \(\mathbf{w}_{k,n} = \overline{\mathbf{w}}_k\) and performs local training via mini-batch stochastic gradient descent (SGD). For a mini-batch \(\mathcal{Q}_{n} \subset \mathcal{D}_n\) of size \(Q\), the stochastic gradient computed by device \(n\) is \[\begin{align} \mathbf{g}_{k,n} = \frac{1}{Q} \sum_{x \in \mathcal{Q}_{n}} \nabla f\left(\mathbf{w}_{k,n}; x\right). \end{align}\] Each device then updates its local model over \(E\) epochs, where a single update step is given by \[\begin{align} \mathbf{w}^{\text{next}}_{k,n} = \mathbf{w}^{\text{current}}_{k,n} - \eta \mathbf{g}_{k,n}, \end{align}\] with \(\eta\) denoting the learning rate.

Upon completing local training, each device \(n\) transmits its updated local model \(\mathbf{w}'_{k,n}\) to the server via the uplink wireless channel. The server then aggregates the received models as \[\begin{align} \overline{\mathbf{w}}_{k+1} = \sum_{n \in \mathcal{N}_k} \frac{D_n}{\sum_{n' \in \mathcal{N}_k} D_{n'}} \mathbf{w}'_{k,n}. \end{align}\] This process repeats until the global model achieves a target optimality gap of \(\epsilon\). Based on convergence analyses in [26], [27], we assume that \(K\) communication rounds are required to reach this target.

2.2 Computing Time and Communication Time↩︎

In each FL round, round time comprises the local computing time at the IIoT devices and communication time. Since the edge server has substantial computational and communication resources, the time for global aggregation and transmission of global model is considered negligible.

2.2.1 Computing Time↩︎

IIoT devices exhibit heterogeneous computing capabilities, characterized by CPU cycle frequencies drawn from finite set \(\mathcal{C}\). Let \(c_n \in \mathcal{C}\) denote the CPU cycle frequency of device \(n\). Assuming that processing one data sample requires \(\rho\) CPU cycles, the computing time for device \(n\) to complete \(E\) local epochs over its dataset of size \(D_n\) is \[\begin{align} \zeta_n = \frac{\rho E D_n}{c_n}. \label{eq:comp95time} \end{align}\tag{2}\] Without loss of generality, devices are indexed in ascending order of computing time, i.e., \(\zeta_n \leq \zeta_{m}\) for all \(n < m\).

2.2.2 Communication Time↩︎

The transmission time for the server to broadcast the global model is considered negligible due to the server’s high transmission power [28], [29]. Accordingly, the communication bottleneck is dominated by the transmission of local models by IIoT devices. Let \(B\) denote the total available system bandwidth. When a bandwidth of \(b_n \leq B\) is allocated to device \(n\) in round \(k\), the transmission rate is \[\begin{align} R_{k,n}(b_n) = b_n \log_2\!\left(1 + \frac{|h_{k,n}|^2 P}{b_n N_0} \right), \label{eq:capacity} \end{align}\tag{3}\] where \(P\) is the transmit power, \(N_0\) is the noise power spectral density, and \(h_{k,n}\) is the channel fading coefficient. We assume that the channel remains static within each round but varies across rounds [30], [31]. Given that the size of local model is equal to \(W\) bits, the transmission time for device \(n\) in round \(k\) is \[\begin{align} \delta_{k,n}(b_n) = \frac{W}{b_n \log_2\!\left(1 + \dfrac{|h_{k,n}|^2 P}{b_n N_0}\right)}. \label{eq:comm95time} \end{align}\tag{4}\]

3 Problem Formulation↩︎

In each round, IIoT devices can begin transmission only after completing their local computations. Due to heterogeneous computing capabilities, devices finish their local updates at different times. To exploit this heterogeneity and improve communication efficiency, we allow bandwidth to be reassigned once a device completes its transmission.

Under this reallocation strategy, the selected devices \(\mathcal{N}_k\) can be partitioned into multiple ordered subsets, where each subset is granted exclusive access to the full bandwidth \(B\). Once all devices in one subset complete their transmissions, the next subset begins. We formalize this as follows.

Definition 1. A bandwidth allocation policy \(\pi\) is a mapping that assigns a partition of \(\mathcal{N}_k\) and a corresponding bandwidth vector to each subset: \[\begin{align} \pi_{\mathcal{N}_k} = (\mathcal{P}, \mathcal{B}), \end{align}\] where \(\mathcal{P} = \{\mathcal{S} \mid \mathcal{S} \subset \mathcal{N}_k\}\) is an ordered partition of \(\mathcal{N}_k\), and \(\mathcal{B} = \{\mathbf{b}_{\mathcal{S}} \in \mathbb{R}^{|\mathcal{S}|} \mid \forall \mathcal{S} \in \mathcal{P}\}\) is the set of bandwidth allocation vectors, each satisfying the total bandwidth constraint \(B\).

We denote by \(\phi\) the class of single-partition policies, i.e., \(\phi_{\mathcal{N}_k} = \bigl(\{\mathcal{N}_k\}, \{\mathbf{b}_{\mathcal{N}_k} \}\bigr)\), in which all selected devices share the bandwidth simultaneously.

The subsets in \(\mathcal{P}\) are ordered by transmission sequence. For a device \(n\) belonging to subset \(\mathcal{S}\), its transmission cannot begin until all devices in the preceding subset have completed their transmissions. The round time of device \(n \in \mathcal{S}\) in round \(k\) is therefore given by \[\begin{align} \tau_{k,n}(\pi_{\mathcal{N}_k}) = \max\bigl\{\zeta_n,\, \tau_{k,\mathcal{S}'}(\pi_{\mathcal{N}_k})\bigr\} + \delta_{k,n}(b_n(\pi_{\mathcal{N}_k})), \label{eq:round95time95client} \end{align}\tag{5}\] where \(\mathcal{S}'\) is the preceding subset of \(\mathcal{S}\), and \(\tau_{k,\mathcal{S}'}(\pi_{\mathcal{N}_k})\) denotes the completion time of devices’ transmission in subset \(\mathcal{S}'\), defined as \[\begin{align} \tau_{k,\mathcal{S}'}(\pi_{\mathcal{N}_k}) = \max_{n' \in \mathcal{S}'} \tau_{k,n'}(\pi_{\mathcal{N}_k}). \label{eq:round95time95part} \end{align}\tag{6}\] Since the round completes when the last subset finishes, the round time for round \(k\) is \[\begin{align} \tau_k(\pi_{\mathcal{N}_k}) = \max_{\mathcal{S} \in \mathcal{P}}\, \tau_{k,\mathcal{S}}(\pi_{\mathcal{N}_k}). \label{eq:round95time} \end{align}\tag{7}\] Summing over all \(K\) rounds, the total training time is \[\begin{align} \tau = \sum_{k=1}^{K} \max_{\mathcal{S} \in \mathcal{P}}\, \tau_{k,\mathcal{S}}(\pi_{\mathcal{N}_k}). \end{align}\]

The round time is lower-bounded by the time for the device with the longest computation time to complete its transmission when allocated the full bandwidth \(B\). Formally, \[\begin{align} \tau_k(\pi_{\mathcal{N}_k}) \geq \zeta_{\hat{n}} + \delta_{k,\hat{n}}(B), \end{align}\] where \(\hat{n} = \arg\max_{n \in \mathcal{N}_k} \zeta_n\). We denote this lower bound as \(\tau_k^{\text{LB}} = \zeta_{\hat{n}} + \delta_{k,\hat{n}}(B)\). To isolate the effect of bandwidth allocation, we define the round time gap \[\begin{align} \Delta_k(\pi_{\mathcal{N}_k}) = \max_{\mathcal{S} \in \mathcal{P}}\, \tau_{k,\mathcal{S}}(\pi_{\mathcal{N}_k}) - \tau_k^{\text{LB}}, \label{eq:gap} \end{align}\tag{8}\] and seek to minimize its average over all rounds: \[\begin{align} \bar{\Delta} = \frac{1}{K} \sum_{k=1}^{K} \Delta_k(\pi_{\mathcal{N}_k}). \label{eq:avg95gap} \end{align}\tag{9}\] The optimal bandwidth allocation policy is obtained by solving \[\begin{align} \boldsymbol{P:} \quad \min_{\pi \in \mathcal{F}}\;\bar{\Delta}, \label{prob:min95bw} \end{align}\tag{10}\] where \(\mathcal{F}\) denotes the set of all feasible policies.

4 Bandwidth Allocation Policy with Device Partitioning↩︎

Solving 10 requires enumerating all feasible partitions of \(\mathcal{N}_k\), resulting in combinatorial complexity that grows exponentially with \(|\mathcal{N}_k|\). To gain insight toward a tractable solution, we first analyze the two-device case \(|\mathcal{N}_k| = 2\).

Proposition 1. For a given communication round \(k\), let \(\mathcal{N}_k = \{i, j\}\) with \(\zeta_i \leq \zeta_j\). The optimal policy \(\pi^*\) minimizing the round time gap is \[\begin{align} \pi^* = \begin{cases} \bigl(\{i\}, \{j\},\, [B],\, [B]\bigr) & \text{if } \tau_{k,i}^* \leq \tau^{\text{th}}, \\[4pt] \bigl(\{i,j\},\, [B - \tilde{b}_j,\, \tilde{b}_j]\bigr) & \text{otherwise}, \end{cases} \label{eq:min95round95t} \end{align}\qquad{(1)}\] where \(\tau^*_{k,i} = \zeta_i + \delta_{k,i}(B)\) is the minimum round time for device \(i\) when allocated the full bandwidth, and the threshold is \[\begin{align} \tau^{\text{th}} = \zeta_j + \delta_{k,j}(\tilde{b}_j) - \delta_{k,j}(B), \end{align}\] with \(\tilde{b}_j\) denoting the bandwidth allocated to device \(j\) such that both devices complete transmission simultaneously.

All feasible policies can be classified into two cases: those in which devices \(i\) and \(j\) share the bandwidth simultaneously, and those in which each device is allocated the bandwidth exclusively in sequence. In the latter case, allocating the full bandwidth to each device individually minimizes per-device communication time, since \(\delta_{k,n}(b_n)\) is strictly decreasing in \(b_n\).

When devices \(i\) and \(j\) share the bandwidth simultaneously, the round time is minimized when both devices complete their transmissions at the same time. To see this, note that the round time equals the completion time of the later-finishing device. If one device finishes earlier than the other, transferring bandwidth from the faster device to the slower one reduces the round time. This reallocation can be applied repeatedly until both devices complete transmission simultaneously, at which point the round time is minimized.

It therefore suffices to compare the policy that ensures simultaneous completion of transmission against the policy that allocates the full bandwidth to each device sequentially.

The round time gap of sequential allocation is given by \[\begin{align} \Delta_k(\pi_{\{i,j\}}) &= \max\{\tau^*_{k,i},\, \zeta_j\} + \delta_{k,j}(B) - \tau^{\text{LB}}_k, \\ &= \left(\tau^*_{k,i} - \zeta_j\right)^+, \end{align}\] where \((x)^+ = \max\{x, 0\}\) denotes the rectified linear unit function.

The round time gap of the simultaneous allocation with equal completion times is given by \[\begin{align} \Delta_k(\pi_{\{i,j\}}) &= \zeta_j + \delta_{k,j}(\tilde{b}_j) - \tau^{\text{LB}}_k, \\ &= \delta_{k,j}(\tilde{b}_j) - \delta_{k,j}(B). \end{align}\]

Sequential allocation achieves a smaller or equal round time gap if and only if \[\begin{align} \left(\tau^*_{k,i} - \zeta_j\right)^+ \leq \delta_{k,j}(\tilde{b}_j) - \delta_{k,j}(B). \label{ineq:cond95prop1} \end{align}\tag{11}\] Rewriting 11 gives \[\begin{align} \tau^*_{k,i} - \delta_{k,j}(B) \leq \zeta_j + \delta_{k,j}(\tilde{b}_j), \end{align}\] which is equivalent to \(\tau_{k,i}^* \leq \tau^{\text{th}}\), completing the proof.

Proposition 1 reveals that when device \(i\) completes both computation and communication quickly, sequential allocation outperforms simultaneous sharing. This stands in contrast to the result of [17], which showed that simultaneous allocation with equal completion times is optimal under a single partition. Proposition 1 demonstrates that partitioning devices into multiple subsets can strictly outperform the best single-partition policy. This insight is generalized to arbitrary \(|\mathcal{N}_k|\) in the following theorem.

Theorem 1 (Superiority of Partitioning). If there exists a subset \(\mathcal{S} \subset \mathcal{N}_k\) such that \[\begin{align} \tau_{k,\mathcal{S}}\!\left(\phi^*_{\mathcal{S}}\right) \leq \min_{n \in \mathcal{S}^c} \zeta_n, \label{ineq:cond95prop2} \end{align}\qquad{(2)}\] where \(\mathcal{S}^c = \mathcal{N}_k \setminus \mathcal{S}\) and \(\phi^*_{\mathcal{S}} = \bigl(\{\mathcal{S}\}, \{(\mathbf{b}^k_{\mathcal{S}})^* \}\bigr)\), then \[\begin{align} \Delta_k\!\left(\theta^*_{\mathcal{N}_k}\right) \leq \Delta_k\!\left(\phi^*_{\mathcal{N}_k}\right), \label{ineq:prop2} \end{align}\qquad{(3)}\] where \(\theta^*_{\mathcal{N}_k} = \bigl(\{\mathcal{S}, \mathcal{S}^c\},\, \{(\mathbf{b}^k_{\mathcal{S}})^*, (\mathbf{b}^k_{\mathcal{S}^c})^*\}\bigr)\) and \(\phi^*_{\mathcal{N}_k} = \bigl(\{\mathcal{N}_k\},\, \{(\mathbf{b}^k_{\mathcal{N}_k})^*\}\bigr)\).

In other words, whenever condition ?? holds, partitioning \(\mathcal{N}_k\) into \(\mathcal{S}\) and \(\mathcal{S}^c\) achieves a strictly lower round time gap than the optimal single-partition policy.

From 7 , \[\begin{align} \tau_k(\theta^*_{\mathcal{N}_k}) = \max\bigl\{\tau_{k,\mathcal{S}} (\theta^*_{\mathcal{N}_k}),\, \tau_{k,\mathcal{S}^c} (\theta^*_{\mathcal{N}_k})\bigr\}. \end{align}\]

Since \(\theta^*_{\mathcal{N}_k}\) is a policy which partitions \(\mathcal{N}_k\) into \(\mathcal{S}\) and \(\mathcal{S}^c\) and allocates entire bandwidth to each part optimally, the round time for each subset is equal to that of single partition policy with optimal bandwidth allocation. In other words, \[\begin{align} \tau_{k, \mathcal{S}}( \theta^*_{\mathcal{N}_k}) = \tau_{k, \mathcal{S}} (\phi^*_{\mathcal{S}}) \label{eq:prop2952} \end{align}\tag{12}\]

Since \(\delta_{k,n}(b^*_n(\theta^*_{\mathcal{N}_k})) > 0\) for all \(n\), \[\begin{align} \tau_{k, \mathcal{S}^{c}} ( \theta^*_{\mathcal{N}_k}) > \max_{n \in \mathcal{S}^c} \zeta_n \end{align}\] Since \(\delta_{k,n}(b^*_n(\theta^*_{\mathcal{N}_k})) > 0\) for all \(n\), \[\begin{align} \tau_{k,\mathcal{S}^c}(\theta^*_{\mathcal{N}_k}) > \max_{n \in \mathcal{S}^c} \zeta_n \geq \min_{n \in \mathcal{S}^c} \zeta_n \geq \tau_{k,\mathcal{S}}(\theta^*_{\mathcal{N}_k}), \end{align}\] where the last inequality follows from ?? and 12 . Hence, we have \(\tau_{k, \mathcal{S}^{c}} ( \theta^*_{\mathcal{N}_k}) >\tau_{k, \mathcal{S}}( \theta^*_{\mathcal{N}_k})\). Therefore, the round time using \(\theta^*_{\mathcal{N}_k}\) becomes \[\begin{align} \tau_k( \theta^*_{\mathcal{N}_k}) = \tau_{k, \mathcal{S}^{c}} ( \theta^*_{\mathcal{N}_k}). \label{eq:prop2953} \end{align}\tag{13}\]

Furthermore, using 5 and \(\underset{n \in \mathcal{S}^c}{\min} \zeta_n \geq \tau_{k, \mathcal{S}}( \theta^*_{\mathcal{N}_k})\) , we can rewrite 13 as \[\begin{align} \tau_k( \theta^*_{\mathcal{N}_k}) & = \max_{n \in \mathcal{S}^{c} } \left\lbrace \zeta_n + \delta_{k,n} ( b^*_n(\theta^*_{\mathcal{N}_k}) ) \right\rbrace. \label{eq:prop2951} \end{align}\tag{14}\]

For the single-partition policy, the round time using \(\phi^*_{\mathcal{N}_k}\) can be written as \[\begin{align} \tau_k(\phi^*_{\mathcal{N}_k}) = \max_{n \in \mathcal{N}_k} \left\lbrace \zeta_n + \delta_{k,n}(b^*_n(\phi^*_{\mathcal{N}_k})) \right\rbrace . \end{align}\] Let \(\tilde{n} = \underset{n \in \mathcal{N}_k}{\arg\max} \left\lbrace \zeta_n + \delta_{k,n}(b^*_n(\phi^*_{\mathcal{N}_k})) \right\rbrace\). Then, consider the following bandwidth allocation scheme for \(\mathcal{S}^c\). \[\begin{align} \tilde{b}_n = \begin{cases} b^*_n(\phi^*_{\mathcal{N}_k}) & \text{ for } n \neq \tilde{n} , \\ B - \sum_{n \neq \tilde{n}} b^*_n(\phi^*_{\mathcal{N}_k}) & \text{ for } n = \tilde{n} \end{cases} \end{align}\] Since the same amount of bandwidth is allocated for \(n \neq \tilde{n}\), \(\delta_{k, n}(\tilde{b}_{n}) = \delta_{k, n}(b^*_{n}(\phi^*_{\mathcal{N}_k}))\). Moreover, as \(|\mathcal{S}^c| < |\mathcal{N}_k|\), \(\tilde{b}_{\tilde{n}} > b^*_{\tilde{n}}(\phi^*_{\mathcal{N}_k})\), which leads to \(\delta_{k, \tilde{n}}(\tilde{b}_{\tilde{n}}) < \delta_{k, \tilde{n}}(b^*_{\tilde{n}}(\phi^*_{\mathcal{N}_k}))\).

Consequently, we have \[\begin{align} \tau_k(\phi^*_{\mathcal{N}_k}) \geq \max_{n \in \mathcal{S}^c}\left[ \zeta_n + \delta_{k,n}(\tilde{b}_n) \right]. \label{ineq:prop2951} \end{align}\tag{15}\] Since \(\theta^*_{\mathcal{N}_k}\) allocates bandwidth optimally to \(\mathcal{S}^c\), we have \[\begin{align} \tau_k(\theta^*_{\mathcal{N}_k}) \leq \max_{n \in \mathcal{S}^c}\left[ \zeta_n + \delta_{k,n}(\tilde{b}_n) \right]. \label{ineq:prop2952} \end{align}\tag{16}\] Combining 15 and 16 and subtracting both sides by \(\tau^{\text{LB}}_k\) yields \[\begin{align} \tau_k(\theta^*_{\mathcal{N}_k}) - \tau_k^{\text{LB}} \leq \tau_k(\phi^*_{\mathcal{N}_k})- \tau_k^{\text{LB}}, \end{align}\] which is equivalent to ?? .

Figure 3: Device Partitioning Bandwidth Policy

The key insight of Theorem 1 is that whenever a subset \(\mathcal{S}\) can complete all transmissions before the remaining devices finish their local computations, partitioning is guaranteed to reduce the round time gap relative to any single-partition policy with optimal bandwidth allocation.

Since the feasible set of 10 grows exponentially with \(|\mathcal{N}_k|\), exhaustive search is computationally intractable in practice. Nevertheless, condition ?? in Theorem 1 provides a verifiable criterion for when partitioning is beneficial: it is always advantageous when devices exhibit large gaps in their computing times.

Building on this insight, we propose the Device Partitioning Bandwidth Policy (DPBP), whose pseudocode is given in Algorithm 3. DPBP applies partitioning only when condition ?? is satisfied. Starting from an empty subset \(\mathcal{S}\), devices are added sequentially in ascending order of computing time. After each addition, the condition is checked: if satisfied, \(\mathcal{S}\) is committed as one partition and DPBP is applied recursively to the remaining devices \(\mathcal{S}^c\); if not, the next device is added and the check is repeated. If no valid subset is found after all devices are considered, the remaining devices form the final subset. In the degenerate case where this occurs at the initial call, all devices are placed in a single partition, recovering the standard simultaneous allocation.

Figure 4: Timeline illustration of DPBP

For each subset \(\mathcal{S} \in \mathcal{P}\), the intra-subset bandwidth allocation \(\overline{\mathbf{b}^k}_{\mathcal{S}}\) is chosen so that all devices in \(\mathcal{S}\) complete transmission simultaneously using entire bandwidth, as illustrated in Fig. 4. This allocation strategy was originally shown to be optimal under a single partition in [17]; however, since [17] assumes transmission power proportional to allocated bandwidth, we modify the algorithm to accommodate the fixed transmission power model considered here.

The complexity of determining the bandwidth allocation for each subset scales linearly with \(|\mathcal{N}_k| = qN\). The number of subsets that DPBP considers is proportional to \(N\), its overall complexity is \(\mathcal{O}(N^2)\). As the partitioning and bandwidth allocation are performed at the server, this computational burden remains manageable given the server’s processing capabilities [32].

5 Experiments↩︎

5.1 Setting↩︎

We use two widely used real-world dataset to evaluate the performance of our proposed bandwidth allocation policy for FL. The first dataset is GC10-Det [33], which is a collection of 2,300 metallic surface images. Each image show different 10 defects on metallic surfaces such as silk spot, welding line, and other scratches. By training through this images, models can learn to detect various types of defects on metallic surfaces, which is crucial for quality control in manufacturing processes. The second dataset is CIFAR10 [34], which is a widely used dataset for image classification tasks. It consists of 60,000 color images of size \(32 \times 32\) pixels, categorized into 10 classes, including airplanes, cars, birds, and cats. The GC10-Det dataset is used to evaluate the performance of our proposed policy in a real-world industrial application scenario, while the CIFAR-10 dataset serves as a benchmark for general federated learning scenario.

5.1.1 GC10-Det dataset↩︎

We consider \(N = 20\) IIoT devices with i.i.d. local data and \(K = 300\) communication rounds. In each round, all devices participate. (i.e., \(q=1\)) The CPU cycle frequency of each device \(c_n\) is drawn uniformly at random from \(\mathcal{C} = \{1, 5, 10, 20\} \times 10^6\) cycles/s. Each device trains a CNN with eight convolutional layers and two fully-connected layers, with model size \(W = 20\) MB. We set \(\rho = 10{,}000\) CPU cycles per sample and \(E = 5\) local epochs. The total bandwidth is \(B = 30\) MHz. The transmit power and noise power spectral density are set \(\frac{P}{N_0} = 20\) dB/Mhz. The wireless channels follow Rayleigh fading.

5.1.2 CIFAR-10 dataset↩︎

We consider \(N = 100\) IIoT devices with heterogeneous local data distributed according to a Dirichlet distribution with parameter \(0.1\), and \(K = 100\) rounds. In each round, \(30\%\) of devices are selected to participate. In other words, we set \(q=0.3\). Each device trains a CNN with two convolutional layers and three fully-connected layers, with model size \(W = 5\) MB. We set \(\rho = 1{,}000\), \(E = 3\), and \(B = 20\) MHz. All other parameters are identical to those in the GC10-Det experiment.

We compare the proposed policy against three baselines: single partition (SP) [17], which is known to be optimal among all single-partition strategies; channel-aware (CA) [16], which allocates bandwidth such that all devices achieve equal transmission rates; and uniform allocation, which assigns equal bandwidth to every device.

5.2 Results↩︎

5.2.1 Performance Comparision↩︎

Figure 5: Test accuracy over time for the GC10-Det dataset.
Table 1: Total training time for the GC10-Det dataset.
Method Proposed SP CA Uniform
Training time (s) 1064.89 1281.25 1781.31 1839.46
Figure 6: Average round time gap for GC10-Det dataset.
Figure 7: Total uplink energy consumption for GC10-Det dataset.

Fig. 5 and Table 1 show the test accuracy and total training time for the GC10-Det dataset. The proposed policy achieves the shortest training time among all methods. Fig. 6 further shows the average round time gap, where the proposed policy most closely approaches the theoretical lower bound. Notably, it yields a substantial reduction over SP, which is itself optimal among single-partition strategies. This demonstrates that device partitioning effectively reduces the communication overhead of FL over IIoT networks.

Fig. 7 reports the total uplink energy consumed by IIoT devices across all rounds. The proposed policy achieves the lowest energy consumption, as allowing devices to transmit with higher bandwidth reduces transmission duration. In contrast, SP exhibits the highest energy consumption despite its second-best round time performance. Under SP, devices that finish computation early must continue transmitting at reduced bandwidth while waiting for slower devices to complete both computation and transmission. This prolonged transmission duration leads to substantially higher energy use.

Figure 8: Test accuracy over time for the CIFAR-10 dataset.
Table 2: Total training time for the CIFAR-10 dataset.
Method Proposed SP CA Uniform
Training time (s) 133.14 140.87 201.93 203.59
Figure 9: Average round time gap for CIFAR-10 dataset.
Figure 10: Total uplink energy consumption for CIFAR-10 dataset.

Results for CIFAR-10 are shown in Figs. 8, 9, and Table 2, exhibiting trends consistent with those observed for GC10-Det. The proposed policy again nearly achieves the lower bound on round time gap. However, the margin over SP is reduced compared to the GC10-Det setting. This is attributable to the smaller model size, which reduces per-round communication time and consequently narrows the absolute deviation of the round time gap from the lower bound across all methods.

5.2.2 Impact of Bandwidth and Computation Load↩︎

Figure 11: Average round time gap under varying total bandwidth.
Figure 12: Average round time gap under varying computation load.

Fig. 11 shows the average round time gap as a function of total bandwidth \(B\), with all other settings identical to the CIFAR-10 experiment. Reducing \(B\) increases the round time gap for all policies, as lower bandwidth directly increases per-device transmission time, making it harder to satisfy the partitioning condition ?? . This is confirmed at \(B = 5\) MHz, where the proposed policy degrades to the performance of SP, indicating that partitioning no longer occurs.

Conversely, as \(B\) increases, the performance gap among policies diminishes. For example, the average round time gaps of the proposed policy and uniform allocation are \(0.03\) s and \(1.03\) s at \(B = 30\) MHz, compared to \(1.21\) s and \(2.93\) s at \(B = 10\) MHz. Under practical noise models, noise power grows proportionally with bandwidth, causing the transmission rate to saturate at large \(B\). Consequently, further increasing \(B\) yields diminishing reductions in communication time, compressing the differences among policies.

Fig. 12 illustrates the effect of computation load on bandwidth allocation performance. As computation load increases, the round time gaps of the proposed policy and SP decrease, while those of CA and uniform allocation remain largely unchanged. Under heterogeneous computing capabilities, higher computation load amplifies the spread of per-device computation times, which makes the partitioning condition easier to satisfy and allows the proposed policy to approach the lower bound. SP similarly benefits by concentrating bandwidth on the straggler, reducing its transmission time. In contrast, CA and uniform allocation lack mechanisms that adapt to computation heterogeneity, so their performance is insensitive to computation load.

Figure 13: Heatmap of average round time gap across bandwidth and computation load settings.
Figure 14: Heatmap of average number of partition groups across bandwidth and computation load settings.
Figure 15: Heatmap of average size of the last group in partition across bandwidth and computation load settings.

Figs. 13 and 14 present heatmaps of the average round time gap and average number of partition groups across a range of bandwidth and computation load values. Both increasing \(B\) and increasing computation load help the proposed policy approach the lower bound, as larger bandwidth shortens communication time and higher computation load increases the spread of computing times — both of which facilitate partitioning.

Comparing the two heatmaps, round time gaps within \(0.1\) s are achieved only when multiple partitions are formed. However, a large number of groups is not necessary to achieve near-lower-bound performance. For instance, with \(\rho = 3{,}000\) fixed, increasing \(B\) from \(20\) to \(50\) MHz raises the average number of groups from \(5.99\) to \(9.78\), yet the proposed policy already achieves the lower bound at \(B = 20\) MHz. This indicates that once the straggler’s computation time becomes the bottleneck, the critical factor is ensuring that the straggler receives the full bandwidth \(B\) — not the total number of groups. In other words, any policy that grants the final partition exclusive access to the full bandwidth achieves near-lower-bound performance regardless of how many groups precede it. This can be verified in Fig. 15, which clearly indicates that the average round time gap is proportional to the size of last group in partition.

5.2.3 Impact of Scheduling Algorithm and Number of Devices↩︎

Figure 16: Average round time gap under different scheduling algorithms.

Fig. 16 compares the performance of bandwidth allocation policies under three scheduling schemes: random, communication-priority, and computation-priority for the same setting of CIFAR10 experiment. Random scheduling selects participating devices with equal probability. Communication-priority scheduling selects devices with the highest channel gains, minimizing communication cost. Computation-priority scheduling selects devices with the shortest local computation times.

Across all scheduling schemes, the proposed policy achieves the smallest round time gap. However, the advantage over baseline methods narrows under computation-priority scheduling. When devices with similar computation times are selected, the spread of per-device computation times is small, which makes the partitioning condition harder to satisfy and reduces the frequency of partitioning. Additionally, under computation-priority scheduling, CA performs comparably to the proposed and SP policies: since selected devices begin transmission at nearly the same time, ensuring equal transmission rates across devices is sufficient to achieve near-simultaneous completion.

Among the three scheduling strategies, communication-priority scheduling yields the smallest round time gap for the proposed policy. High channel gains reduce per-device communication time, which allows finer-grained partitioning and reduces the size of the final group. With fewer devices competing for bandwidth in the last partition, each device receives a larger share of \(B\), further reducing the straggler’s transmission time.

Figure 17: Average round time gap for different number of IIoT devices

Fig. 17 shows the average round time gap as a function of the number of IIoT devices \(N\). As \(N\) increases, the performance gap between the proposed policy and the baselines widens. In particular, the round time gaps of CA and uniform allocation grow approximately linearly with \(N\), whereas those of the proposed policy and SP increase at a significantly slower rate.

This divergence is driven by the fact that, given a fixed total bandwidth, the bandwidth available per device diminishes as the network scales. Since the communication round time is dictated by the straggler, the performance of CA and Uniform—which fail to adaptively assign bandwidth to the straggler—degrades as the per-device allocation shrinks. Conversely, our proposed policy and SP dynamically determine bandwidth based on the computation times of the devices, resulting in a sublinear growth of the average round time gap.

The performance margin between the proposed policy and SP also widens at large \(N\). While SP adaptively allocates bandwidth, it is constrained to assign non-zero bandwidth to every selected device simultaneously. At large \(N\), this constraint forces the straggler to share bandwidth with many other devices, causing its communication time to notably exceed the theoretical lower bound. The proposed policy alleviates this by grouping devices into sequential subsets, thereby granting the final subset, which contains the straggler exclusive access to the full bandwidth \(B\), regardless of \(N\).

Interestingly, the average round time gap of our proposed policy does not follow a strictly monotonic trend. For instance, the gap at \(N=40\) is larger than at \(N=50\), with similar behavior observed between \(N=70\) and \(N=80\). This occurs because our policy achieves a smaller round time gap when the size of the final group in the partition is minimized. Since the final group size is not a monotone function of \(N\), adding a device with a sufficiently long computation time can cause it to be assigned to its own group, effectively reducing the final group size and thereby decreasing the round time gap.

6 Conclusion↩︎

In this paper, we investigated bandwidth allocation for FL over IIoT networks and proposed DBBP, a bandwidth allocation policy. In each communication round, DPBP partitions the selected IIoT devices into multiple ordered subsets and grants each subset exclusive access to the full bandwidth sequentially. We formally proved that this partitioning approach achieves a lower round time gap than the optimal single-partition strategy, irrespective of the underlying scheduling algorithm. Extensive experiments on the GC10-Det and CIFAR-10 datasets confirmed that DPBP consistently reduces total training time and uplink energy consumption compared to existing bandwidth allocation methods, validating the practical benefit of device partitioning in FL systems deployed over IIoT networks.

Several directions remain open for future work. First, as energy efficiency is a critical constraint in IIoT deployments, extending DPBP to incorporate explicit energy consumption objectives is a natural next step. Second, joint optimization of device scheduling and bandwidth partitioning presents a promising avenue for further reducing learning time, as the current work treats scheduling as an independent component. Scheduling algorithm to improve the performance of DBBP further can be optimized. Finally, while DPBP is designed for IIoT settings, its applicability extends to any domain where edge devices compete for limited bandwidth resources, including vehicular networks, mobile federated learning, and smart manufacturing.

References↩︎

[1]
A. H. Sodhro, S. Pirbhulal, and V. H. C. de Albuquerque, “Artificial intelligence-driven mechanism for edge computing-based industrial applications,” IEEE Transactions on Industrial Informatics, vol. 15, no. 7, pp. 4235–4243, 2019, doi: 10.1109/TII.2019.2902878.
[2]
D. C. Nguyen et al., “Federated learning for industrial internet of things in future industries,” IEEE Wireless Communications, vol. 28, no. 6, pp. 192–199, 2021, doi: 10.1109/MWC.001.2100102.
[3]
Y. Lu, X. Huang, K. Zhang, S. Maharjan, and Y. Zhang, “Communication-efficient federated learning for digital twin edge networks in industrial IoT,” IEEE Transactions on Industrial Informatics, vol. 17, no. 8, pp. 5709–5718, 2021, doi: 10.1109/TII.2020.3010798.
[4]
S. M. Shah and V. K. Lau, “Model compression for communication efficient federated learning,” IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 9, 2021.
[5]
W. Luping, W. Wei, and L. Bo, CMFL: Mitigating communication overhead for federated learning,” in IEEE 39th international conference on distributed computing systems (ICDCS), 2019, pp. 954–964.
[6]
W. Liu, L. Chen, Y. Chen, and W. Zhang, “Accelerating federated learning via momentum gradient descent,” IEEE Transactions on Parallel and Distributed Systems, vol. 31, no. 8, 2020.
[7]
H. Wu and P. Wang, “Fast-convergent federated learning with adaptive weighting,” IEEE Transactions on Cognitive Communications and Networking, vol. 7, no. 4, 2021.
[8]
D. Kivanc, G. Li, and H. Liu, “Computationally efficient bandwidth allocation and power control for OFDMA,” IEEE transactions on wireless communications, vol. 2, no. 6, pp. 1150–1158, 2003.
[9]
N. Jindal, J. G. Andrews, and S. Weber, “Bandwidth partitioning in decentralized wireless networks,” IEEE Transactions on Wireless Communications, vol. 7, no. 12, pp. 5408–5419, 2008, doi: 10.1109/T-WC.2008.071220.
[10]
X. Liu, Z. Qin, Y. Gao, and J. A. McCann, “Resource allocation in wireless powered IoT networks,” IEEE Internet of Things Journal, vol. 6, no. 3, pp. 4935–4945, 2019.
[11]
W. Gao, Z. Zhao, G. Min, Q. Ni, and Y. Jiang, “Resource allocation for latency-aware federated learning in industrial internet of things,” IEEE Transactions on Industrial Informatics, vol. 17, no. 12, pp. 8505–8513, 2021, doi: 10.1109/TII.2021.3073642.
[12]
J. Xu and H. Wang, “Client selection and bandwidth allocation in wireless federated learning networks: A long-term perspective,” IEEE Transactions on Wireless Communications, vol. 20, no. 2, pp. 1188–1200, 2020.
[13]
H. Ko, J. Lee, S. Seo, S. Pack, and V. C. Leung, “Joint client selection and bandwidth allocation algorithm for federated learning,” IEEE Transactions on Mobile Computing, vol. 22, no. 6, 2021.
[14]
J. Kuang, M. Yang, H. Zhu, and H. Qian, “Client selection with bandwidth allocation in federated learning,” in IEEE global communications conference (GLOBECOM), 2021, pp. 01–06.
[15]
Z. Zhao, J. Xia, L. Fan, X. Lei, G. K. Karagiannidis, and A. Nallanathan, “System optimization of federated learning networks with a constrained latency,” IEEE Transactions on Vehicular Technology, vol. 71, no. 1, 2021.
[16]
J. Ren, Y. He, D. Wen, G. Yu, K. Huang, and D. Guo, “Scheduling for cellular federated edge learning with importance and channel awareness,” IEEE Transactions on Wireless Communications, vol. 19, no. 11, 2020.
[17]
W. Shi, S. Zhou, and Z. Niu, “Device scheduling with fast convergence for wireless federated learning,” in IEEE international conference on communications (ICC), 2020, pp. 1–6.
[18]
W. Shi, S. Zhou, Z. Niu, M. Jiang, and L. Geng, “Joint device scheduling and resource allocation for latency constrained wireless federated learning,” IEEE Transactions on Wireless Communications, vol. 20, no. 1, 2020.
[19]
T. Zhao, F. Li, and L. He, “DRL-based joint resource allocation and device orchestration for hierarchical federated learning in NOMA-enabled industrial IoT,” IEEE Transactions on Industrial Informatics, vol. 19, no. 6, pp. 7468–7479, 2023, doi: 10.1109/TII.2022.3170900.
[20]
Y. Zhao, M. Li, L. Lai, N. Suda, D. Civin, and V. Chandra, “Federated learning with non-iid data,” arXiv preprint arXiv:1806.00582, 2018.
[21]
T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith, “Federated optimization in heterogeneous networks,” Proceedings of Machine learning and systems, vol. 2, pp. 429–450, 2020.
[22]
F. Sattler, S. Wiedemann, K.-R. Müller, and W. Samek, “Robust and communication-efficient federated learning from non-iid data,” IEEE Transactions on Neural Networks and Learning Systems, vol. 31, no. 9, 2019.
[23]
J. Xu, W. Du, Y. Jin, W. He, and R. Cheng, “Ternary compression for communication-efficient federated learning,” IEEE Transactions on Neural Networks and Learning Systems, vol. 33, no. 3, 2020.
[24]
S. Caldas, J. Konečny, H. B. McMahan, and A. Talwalkar, “Expanding the reach of federated learning by reducing client resource requirements,” arXiv preprint arXiv:1812.07210, 2018.
[25]
D. Wen, K.-J. Jeon, and K. Huang, “Federated dropout—a simple approach for enabling federated learning on resource constrained devices,” IEEE wireless communications letters, vol. 11, no. 5, pp. 923–927, 2022.
[26]
X. Li, K. Huang, W. Yang, S. Wang, and Z. Zhang, “On the convergence of fedavg on non-iid data,” arXiv preprint arXiv:1907.02189, 2019.
[27]
Y. J. Cho, J. Wang, and G. Joshi, “Towards understanding biased client selection in federated learning,” in International conference on artificial intelligence and statistics, 2022, pp. 10351–10375.
[28]
Z. Yang, M. Chen, W. Saad, C. S. Hong, and M. Shikh-Bahaei, “Energy efficient federated learning over wireless communication networks,” IEEE Transactions on Wireless Communications, vol. 20, no. 3, 2020.
[29]
J. Feng, L. Liu, Q. Pei, and K. Li, “Min-max cost optimization for efficient hierarchical federated learning in wireless edge networks,” IEEE Transactions on Parallel and Distributed Systems, vol. 33, no. 11, 2021.
[30]
W. Guo, R. Li, C. Huang, X. Qin, K. Shen, and W. Zhang, “Joint device selection and power control for wireless federated learning,” IEEE Journal on Selected Areas in Communications, vol. 40, no. 8, 2022.
[31]
J. Yao, W. Xu, Z. Yang, X. You, M. Bennis, and H. V. Poor, “Wireless federated learning over resource-constrained networks: Digital versus analog transmissions,” IEEE Transactions on Wireless Communications, vol. 23, no. 10, pp. 14020–14036, 2024.
[32]
J. Gong, J. S. Thompson, S. Zhou, and Z. Niu, “Base station sleeping and resource allocation in renewable energy powered cellular networks,” IEEE Transactions on Communications, vol. 62, no. 11, 2014.
[33]
X. Lv, F. Duan, J. Jiang, X. Fu, and L. Gan, “Deep metallic surface defect detection: The new benchmark and detection network,” Sensors, vol. 20, no. 6, p. 1562, 2020.
[34]
A. Krizhevsky, G. Hinton, et al., “Learning multiple layers of features from tiny images,” 2009.

  1. K. Kim was with the School of Electrical and Electronics Engineering, Pusan National University, Busan, 46241, Korea (e-mail : km.kim@pusan.ac.kr).↩︎

  2. J. Song is with the School of Electrical and Electronics Engineering, Pusan National University, Busan, 46241, Korea (e-mail : jsong@pusan.ac.kr).↩︎