Do Newer Lightweight CNNs Perform Better Under Resource Constraints?
A Controlled Multigenerational Study of Architecture, Initialization, Training Budget, and Efficiency
January 01, 1970
Newer lightweight convolutional neural network designs are often presented as offering improved predictive performance and deployment efficiency, but these advantages require evaluation under controlled downstream conditions. This study compares nine lightweight CNN model packages across CIFAR-10, CIFAR-100, and Tiny ImageNet using a shared training and evaluation protocol. Predictive performance is assessed using top-1 accuracy, macro F1, and top-5 accuracy, while resource demand is measured through parameter count, FP32 parameter storage, multiply-accumulate operations, standardized batch size 1 latency on an NVIDIA L4 and an AMD Ryzen 5 5500U CPU, and peak PyTorch CUDA allocated tensor memory. Accuracy and resource tradeoffs are further examined using independently calculated point-estimate Pareto frontiers. EfficientNetV2-S records the highest observed top-1 accuracy on CIFAR-10 and CIFAR-100 at 97.57% and 86.98%, whereas RepViT-M1.0 leads Tiny ImageNet at 79.87%, exceeding EfficientNetV2-S by 1.14 percentage points. EfficientNet-B0 records 97.35%, 86.13%, and 78.08% top-1 accuracy, remaining within 0.22, 0.85, and 1.79 percentage points of the best result on CIFAR-10, CIFAR-100, and Tiny ImageNet, respectively. Compared with EfficientNetV2-S, it uses approximately 79% fewer parameters and 86% fewer GMACs, while compared with RepViT-M1.0, it uses approximately 35% fewer parameters and 64% fewer GMACs. EfficientNet-B0 also belongs to each independently evaluated accuracy-resource Pareto frontier across all three datasets, supporting its role as the most consistently competitive intermediate-budget option. MobileNetV3-Small achieves 96.12%, 82.43%, and 70.95% accuracy, records the lowest GMAC count, and is the fastest model under both evaluated CPU thread settings. MobileNetV4-Conv-S uses approximately 60% more parameters and three times the GMACs of MobileNetV3-Small, yet records pretrained accuracy lower by 0.16, 0.53, and 1.04 percentage points across the three datasets. Under random initialization, MobileNetV3-Small achieves higher final test accuracy than MobileNetV4-Conv-S by 2.55, 1.76, and 0.99 percentage points, with paired test-set intervals excluding zero for the fixed trained models on all three datasets. The initialization study further shows that EfficientNet-B0 remains 3.29, 10.10, and 17.54 percentage points below its pretrained counterpart even after 100 epochs of scratch training, despite requiring approximately five times the recorded training time. At the smallest end of the model range, SqueezeNet1.1 has the lowest parameter count and peak CUDA allocation, but its substantially weaker predictive performance prevents it from matching the overall low-resource profile of MobileNetV3-Small. Latency rankings differ considerably between the evaluated L4 and CPU execution environments, showing that theoretical computation alone does not reliably predict measured inference performance. Overall, newer designs provide selective rather than universal gains. EfficientNetV2-S leads the CIFAR tasks, RepViT-M1.0 leads Tiny ImageNet, EfficientNet-B0 provides the most consistent general tradeoff, and MobileNetV3-Small emerges as the strongest evaluated option under severe resource constraints.
This manuscript is currently under consideration at a peer reviewed venue. The preprint may be updated following feedback received during the review process.
Keywords: lightweight convolutional neural networks, image classification, resource constraints, Pareto analysis, transfer learning, CPU latency, peak tensor memory, controlled benchmarking
Lightweight convolutional neural networks remain central to image classification on mobile, embedded, and resource-limited systems. Their design objective is rarely a single scalar quantity. A deployment may limit parameter storage, arithmetic cost, peak memory, latency, or training time, while still requiring an acceptable level of predictive quality. These quantities are related but not interchangeable. A model with few GMACs may execute slowly when its operators are fragmented or poorly supported, while a larger dense network may run efficiently on a highly parallel accelerator. Model selection under resource constraints is therefore a multiobjective problem rather than a simple accuracy leaderboard.
A second challenge is architectural turnover. New lightweight networks are typically introduced with architecture-specific pretraining, augmentation, distillation, input resolution, and hardware targets. Comparing their public checkpoints under one downstream recipe is practically useful, but it does not isolate architecture alone. This study therefore distinguishes two questions. The nine-model benchmark evaluates public pretrained architecture and checkpoint combinations under controlled downstream conditions. Separate scratch experiments provide narrower evidence about EfficientNet-B0 training budget and the MobileNetV3-Small versus MobileNetV4-Conv-S ordering.
The term controlled refers to preprocessing, downstream optimization, input size, evaluation, and resource measurement. It does not imply that upstream pretraining is controlled. The term multigenerational refers to a representative sample of design strategies from 2016 to 2024 rather than a complete chronology of every efficient CNN. The pool retains the seven established architectures from the earlier controlled benchmark and adds RepViT-M1.0 and MobileNetV4-Conv-S as two 2024 designs representing recent reparameterized and universal mobile-block directions. All nine backbones have public ImageNet-1K checkpoints, reproducible implementations, standard 224 pixel input support, fewer than approximately 25 million parameters, and distinct design mechanisms.
The study addresses five research questions:
RQ1: How do nine public pretrained model packages compare in top-1 accuracy, macro F1, and top-5 accuracy under one downstream protocol?
RQ2: Which models remain point estimate Pareto efficient under parameters, GMACs, L4 latency, CPU latency, and peak CUDA tensor memory?
RQ3: How strongly do theoretical resource indicators correspond to measured latency across the evaluated GPU and CPU execution environments?
RQ4: How do separate 20 epoch and 100 epoch EfficientNet-B0 scratch schedules compare in accuracy and recorded wall clock time?
RQ5: Does MobileNetV4-Conv-S improve upon MobileNetV3-Small under the evaluated pretrained and scratch settings?
The main contributions are:
A controlled downstream comparison of nine lightweight CNN model packages across three datasets, with explicit checkpoint provenance.
A multidimensional resource evaluation using parameters, FP32 storage, GMACs, L4 latency distributions, consumer CPU latency, and peak CUDA allocated tensor memory.
Point estimate Pareto frontiers, transparent example budgets, execution-environment rank analysis, and descriptive rank correlations that replace informal claims about a universal best balance.
Focused training studies using cumulative wall clock time, extended EfficientNet-B0 scratch schedules, MobileNet cross-generation learning trajectories, and paired test-set uncertainty.
Lightweight CNN research has progressed through several complementary strategies. SqueezeNet reduced parameter storage through Fire modules and aggressive use of \(1\times1\) convolutions [1]. ResNet established residual learning as a robust conventional baseline [2]. MobileNetV2 introduced inverted residual blocks and linear bottlenecks [3], while ShuffleNetV2 shifted attention from nominal FLOPs toward memory access cost, channel fragmentation, and practical operator design [4]. MobileNetV3 combined hardware-aware neural architecture search, squeeze-and-excitation, and specialized nonlinearities [5]. EfficientNet formalized compound scaling across depth, width, and resolution [6], and EfficientNetV2 introduced fused mobile blocks and training-aware scaling to improve both optimization and inference efficiency [7]. Together, these works show that compactness can be pursued through parameter compression, block redesign, scaling policy, or hardware-aware search rather than one universal architectural recipe.
Recent efficient networks increasingly distinguish theoretical arithmetic from realized latency. MobileOne uses structural reparameterization so that a multi-branch training graph can be converted into a simpler inference graph [8]. FasterNet explicitly argues that low FLOPs do not guarantee fast execution and introduces partial convolution to reduce redundant memory access [9]. FastViT similarly combines structural reparameterization with hybrid convolutional blocks to improve latency and accuracy tradeoffs [10]. RepViT transfers efficient macro and micro design choices associated with mobile vision transformers back into a pure CNN [11]. MobileNetV4 introduces the Universal Inverted Bottleneck, mobile multi-query attention, and a hardware-oriented search framework designed to remain competitive across CPUs, DSPs, GPUs, and accelerators [12].
These studies motivate measuring parameters, arithmetic, memory, and latency separately. They also show why a model optimized for one device or runtime may not preserve the same ordering elsewhere. More recent compact CNNs continue to explore alternative mechanisms: StarNet studies multiplicative feature expansion [13], LSNet combines large-field perception with small-field aggregation [14], and UniConvNet expands effective receptive fields while preserving the proposed spatial distribution of influence [15]. These models are relevant to the evolving design landscape, but they are not added to the experimental pool because the present study retains a fixed established benchmark base and adds only the two 2024 designs central to its cross-generation question.
Large backbone studies demonstrate that rankings depend on the target task, upstream data, pretraining method, and adaptation procedure. Goldblum et al. compare a broad set of supervised, self-supervised, vision-language, and randomly initialized backbones across classification, detection, retrieval, and distribution-shift tasks [16]. Pegeot et al. examine compact-model transfer under upstream and downstream constraints, including linear probing and full fine tuning [17]. The RCV2023 challenges evaluate training and inference under explicit time, memory, and latency budgets [18]. Jeevan and Sethi compare pretrained backbones across multiple visual domains and reduced-data settings [19]. Guerin et al. formalize target-specific backbone selection in low-data regimes and show that no universal backbone dominates a large model pool [20].
This literature supports two distinctions used here. First, a shared downstream recipe provides protocol equality but does not isolate architecture from checkpoint provenance or guarantee architecture-specific optimality. Second, resource-aware selection requires explicit budgets or Pareto reasoning rather than an informal claim that one network has the best overall balance. The present analysis therefore treats each public pretrained architecture and checkpoint combination as a model package, and it calculates bivariate accuracy-resource frontiers independently for each resource dimension.
ImageNet pretraining often improves downstream accuracy and optimization efficiency, but the magnitude of the benefit depends on target data, task scale, and training schedule [21], [22]. Scratch and pretrained comparisons must therefore distinguish initialization from optimization exposure. Distillation further complicates public-checkpoint comparisons because it can improve a model package without changing the downstream architecture definition. These considerations motivate the separate EfficientNet-B0 schedule study and the matched MobileNet scratch comparison.
The pretrained results for the seven established architectures and the original fixed-budget EfficientNet-B0 initialization records were previously reported in [23]. RepViT-M1.0, MobileNetV4-Conv-S, the fresh 100 epoch schedules, unified latency and memory measurements, Pareto analyses, cumulative-time analyses, and paired prediction analyses are introduced in the present manuscript.
The benchmark uses CIFAR-10, CIFAR-100, and Tiny ImageNet. Ten percent of each official training set is reserved for validation through a fixed stratified split. The official CIFAR test sets are retained for final evaluation. Tiny ImageNet does not distribute labels for its official test set, so its labeled validation set is used as the benchmark test set, while the internal validation partition is drawn only from the original training images. Tiny ImageNet is derived from ImageNet and shares its class taxonomy, so it is not fully independent of ImageNet pretraining.
| Dataset | Classes | Train | Validation | Test |
|---|---|---|---|---|
| CIFAR-10 | 10 | 45,000 | 5,000 | 10,000 |
| CIFAR-100 | 100 | 45,000 | 5,000 | 10,000 |
| Tiny ImageNet | 200 | 90,000 | 10,000 | 10,000 |
The seven established architectures are instantiated from torchvision using public ImageNet-1K weights distributed with the library. RepViT-M1.0 uses repvit_m1_0.dist_450e_in1k from timm. This checkpoint was trained with distillation by the
RepViT authors. The timm implementation retains two classifier heads for distilled variants and averages their outputs during evaluation. The standardized 200-class RepViT therefore contains 6,584,396 parameters, and both heads are included consistently in
parameter count, storage, GMACs, latency, and memory measurement. MobileNetV4-Conv-S uses mobilenetv4_conv_small.e2400_r224_in1k from timm, a public checkpoint trained with a MobileNetV4-inspired recipe rather than an official Google release.
For every pretrained run, the source classifier is replaced for the target dataset and the full network is fine tuned.
| Model | Year | Design family | Implementation | Public checkpoint | Provenance note |
|---|---|---|---|---|---|
| SqueezeNet1.1 | 2016 | Fire modules | torchvision | ImageNet-1K | torchvision-maintained ImageNet-1K weights |
| ResNet18 | 2016 | Residual blocks | torchvision | ImageNet-1K | torchvision-maintained ImageNet-1K weights |
| MobileNetV2 | 2018 | Inverted residuals | torchvision | ImageNet-1K | torchvision-maintained ImageNet-1K weights |
| ShuffleNetV2 x1.0 | 2018 | Channel split and shuffle | torchvision | ImageNet-1K | torchvision-maintained ImageNet-1K weights |
| EfficientNet-B0 | 2019 | Compound scaling | torchvision | ImageNet-1K | torchvision-maintained ImageNet-1K weights |
| MobileNetV3-Small | 2019 | Searched mobile blocks | torchvision | ImageNet-1K | torchvision-maintained ImageNet-1K weights |
| EfficientNetV2-S | 2021 | Fused mobile blocks | torchvision | ImageNet-1K | torchvision-maintained ImageNet-1K weights |
| RepViT-M1.0 | 2024 | ViT-inspired reparameterized CNN | timm 1.0.27 | repvit_m1_0.dist_450e_in1k | Distilled checkpoint, dual heads averaged at evaluation |
| MobileNetV4-Conv-S | 2024 | Universal Inverted Bottleneck | timm 1.0.27 | mobilenetv4_conv_small.e2400_r224_in1k | Public third-party checkpoint, not an official Google release |
All images are processed at 224 by 224 pixels. Training uses random resized crop with scale 0.80 to 1.00 and aspect ratio 0.90 to 1.10, bicubic interpolation, horizontal flipping, color jitter, and ImageNet normalization. Validation and test images are resized deterministically and normalized with the same channel statistics. The 224 pixel setting preserves compatibility with source checkpoints, but it is a transfer adaptation choice rather than a native-resolution study. Upscaling does not create new image detail and increases computation relative to the original dataset resolution.
Models are optimized with AdamW using an initial learning rate of \(3\times10^{-4}\), weight decay \(10^{-4}\), PyTorch default coefficients \(\beta=(0.9,0.999)\) and \(\epsilon=10^{-8}\), label smoothing 0.1, mixed precision, and gradient clipping at 1.0. CosineAnnealingLR uses \(T_{\max}\) equal to
the configured schedule length and a minimum learning rate of \(3\times10^{-6}\). Training batch size is 32 and evaluation batch size is 64. The common schedule has a maximum of 20 epochs with early stopping patience 5.
Validation accuracy is the checkpoint criterion, with macro F1 used as a tie breaker. Training loaders shuffle with a seeded generator and drop the final incomplete batch; validation and test loaders use deterministic ordering. The training loader uses
pin_memory=True, seeded worker initialization, and persistent workers when worker processes are available. cuDNN deterministic mode is enabled and cuDNN benchmark mode is disabled. All reported training runs use seed 42 unless otherwise noted,
and the primary benchmark reports one completed training run per setting.
| Setting | Value | Setting | Value |
|---|---|---|---|
| Input resolution | 224 by 224 | Optimizer | AdamW |
| Maximum epochs | 20 | Initial learning rate | \(3\times10^{-4}\) |
| Training batch size | 32 | Weight decay | \(10^{-4}\) |
| Evaluation batch size | 64 | Label smoothing | 0.1 |
| Scheduler | Cosine annealing | Gradient clipping | 1.0 |
| Precision | Mixed precision | Checkpoint criterion | Validation accuracy |
| Early stopping | Patience 5 | Fine tuning | Full network |
EfficientNet-B0 is evaluated under ImageNet pretrained initialization with the common schedule, random initialization with the same maximum 20 epoch schedule, and a fresh random initialization run using a cosine schedule configured for 100 epochs. The 100 epoch runs start from a new random initialization rather than continuing the completed 20 epoch schedule. They use early stopping patience 15. Comparisons therefore describe two distinct schedules, not the effect of simply appending 80 epochs.
MobileNetV3-Small and MobileNetV4-Conv-S are evaluated under the public pretrained protocol and under matched 20 epoch scratch training. These experiments measure learning behavior under one fixed downstream budget. They do not estimate architecture-specific maximum scratch performance.
All recorded training-time analyses use the NVIDIA L4 environment reported in the saved experiment metadata. The main software environment used Python 3.12.13, PyTorch 2.11.0 with CUDA 12.8, torchvision 0.26.0, and timm 1.0.27 for the recent models.
Static efficiency is described by total parameters, FP32 parameter storage, and GMACs for input shape \(1\times3\times224\times224\). Cross-model resource measurements use standardized 200-class variants so that one classifier size is applied consistently. These figures are standardized proxies for comparing architecture instances and are not the exact CIFAR-10 classifier footprints.
NVIDIA L4 latency uses PyTorch eager execution, FP32 model parameters, FP16 CUDA autocast, batch size 1, 50 warmups, and 500 timed forward passes measured with CUDA Events. Host transfer and preprocessing are excluded. The same run records baseline allocated tensor memory, peak allocated tensor memory, incremental inference peak above the loaded-model baseline, and allocator-reserved memory. The main memory metric is peak PyTorch CUDA allocated tensor memory, not total device memory.
Consumer CPU latency is measured on an AMD Ryzen 5 5500U using Windows 11, PyTorch 2.12.1 CPU, FP32, MKLDNN, batch size 1, and contiguous NCHW input. One-thread and four-thread modes are evaluated in fresh subprocesses. Each configuration uses 30 warmups followed by three rounds of 100 measured passes. Median and p95 latency are emphasized because occasional operating-system activity can affect tail measurements.
For each dataset and resource metric, a model is point estimate Pareto optimal when no other evaluated model is both at least as accurate and no more resource demanding, with one strict improvement, following the standard bivariate Pareto concept in multiobjective optimization [24]. Each frontier is calculated independently using top-1 accuracy and one resource dimension at a time; the study does not construct one simultaneous multidimensional frontier. These frontiers do not include training-seed uncertainty, so membership near small accuracy gaps may change after retraining.
For selected model pairs, the stored predictions are compared over the same ordered test examples. Accuracy differences are summarized with 50,000 paired bootstrap replicates [25] and exact two-sided McNemar tests [26]. The prediction files do not contain explicit sample identifiers, but true-label sequences match row by row and the evaluation loaders use deterministic ordering. The comparisons are exploratory and are not adjusted for multiplicity; they are used to qualify selected observed differences rather than support confirmatory architecture-level hypothesis testing. This analysis quantifies uncertainty over fixed test examples for the trained models. It does not capture variability across independent training runs.
Table ¿tbl:tab:predictive? reports top-1 accuracy and macro F1. EfficientNetV2-S ranks first on CIFAR-10 and CIFAR-100. RepViT-M1.0 ranks third on both CIFAR datasets but leads Tiny ImageNet at 79.87%, exceeding EfficientNetV2-S by 1.14 points and EfficientNet-B0 by 1.79 points. MobileNetV3-Small remains competitive despite its compact resource profile. MobileNetV4-Conv-S records lower observed top-1 accuracy than MobileNetV3-Small by 0.16, 0.53, and 1.04 points across the three datasets.
Macro F1 closely follows top-1 accuracy because the evaluation datasets are class balanced. Top-5 accuracy provides additional information on the higher-class-count tasks. EfficientNet-B0 leads CIFAR-100 top-5 accuracy at 97.29%, while RepViT-M1.0 leads Tiny ImageNet at 93.09%.
| CIFAR-10 | CIFAR-100 | Tiny ImageNet | ||||
|---|---|---|---|---|---|---|
| 2-3(lr)4-5(lr)6-7 Model | Acc. (%) | Macro F1 | Acc. (%) | Macro F1 | Acc. (%) | Macro F1 |
| EfficientNetV2-S | 97.57 | 0.9757 | 86.98 | 0.8697 | 78.73 | 0.7866 |
| EfficientNet-B0 | 97.35 | 0.9735 | 86.13 | 0.8615 | 78.08 | 0.7801 |
| RepViT-M1.0 | 97.25 | 0.9725 | 85.40 | 0.8534 | 79.87 | 0.7981 |
| ResNet18 | 96.52 | 0.9652 | 82.53 | 0.8250 | 68.22 | 0.6815 |
| MobileNetV3-Small | 96.12 | 0.9611 | 82.43 | 0.8239 | 70.95 | 0.7093 |
| MobileNetV2 | 95.99 | 0.9599 | 81.45 | 0.8144 | 71.41 | 0.7126 |
| MobileNetV4-Conv-S | 95.96 | 0.9595 | 81.90 | 0.8183 | 69.91 | 0.6979 |
| ShuffleNetV2 x1.0 | 94.78 | 0.9477 | 80.38 | 0.8036 | 69.35 | 0.6925 |
| SqueezeNet1.1 | 93.73 | 0.9371 | 73.34 | 0.7330 | 60.41 | 0.6030 |
The pretrained results for the seven established architectures and the original fixed-budget EfficientNet-B0 initialization records were previously reported in [23]. RepViT-M1.0, MobileNetV4-Conv-S, the fresh 100 epoch schedules, unified latency and memory measurements, Pareto analyses, cumulative-time analyses, and paired prediction analyses are new to this manuscript.
Full top-5 results are provided in Appendix 8.1.
Table 5 reports standardized 200-class resource measurements. EfficientNetV2-S is the largest and most computationally demanding model. EfficientNet-B0 uses approximately 79% fewer parameters and 86% fewer GMACs than EfficientNetV2-S while trailing it by only 0.22 and 0.85 points on the CIFAR tasks. Relative to the Tiny ImageNet leader, RepViT-M1.0, EfficientNet-B0 uses approximately 35% fewer parameters and 64% fewer GMACs while trailing by 1.79 points.
MobileNetV3-Small defines the lowest arithmetic budget at 0.0617 GMACs. It uses 37% fewer parameters and 67% fewer GMACs than MobileNetV4-Conv-S, while recording higher observed accuracy on all three datasets. SqueezeNet1.1 has the lowest parameter count and the smallest peak CUDA allocation, but its substantially lower CIFAR-100 and Tiny ImageNet accuracy limits its practical profile.
| Model | Params (M) | FP32 (MiB) | GMACs | Peak CUDA allocated (MiB) | Incremental inference peak (MiB) |
|---|---|---|---|---|---|
| EfficientNetV2-S | 20.434 | 77.948 | 2.9009 | 93.11 | 3.76 |
| EfficientNet-B0 | 4.264 | 16.265 | 0.4141 | 31.18 | 4.98 |
| RepViT-M1.0 | 6.584 | 25.117 | 1.1420 | 37.74 | 2.57 |
| ResNet18 | 11.279 | 43.026 | 1.8236 | 62.24 | 9.46 |
| MobileNetV3-Small | 1.723 | 6.572 | 0.0617 | 17.51 | 1.13 |
| MobileNetV2 | 2.480 | 9.461 | 0.3265 | 24.34 | 4.98 |
| MobileNetV4-Conv-S | 2.749 | 10.487 | 0.1848 | 22.84 | 2.51 |
| ShuffleNetV2 x1.0 | 1.459 | 5.564 | 0.1519 | 17.79 | 2.39 |
| SqueezeNet1.1 | 0.825 | 3.147 | 0.2800 | 7.08 | 3.35 |
Across CIFAR-10 and CIFAR-100, the parameter frontier contains SqueezeNet1.1, ShuffleNetV2 x1.0, MobileNetV3-Small, EfficientNet-B0, and EfficientNetV2-S. The GMAC frontier contains MobileNetV3-Small, EfficientNet-B0, and EfficientNetV2-S. Tiny ImageNet changes the ordering: MobileNetV2 and RepViT-M1.0 become important frontier members, reflecting the stronger performance of these models on the more difficult 200-class task.
Peak allocated memory is strongly associated with parameter count in this model pool, with Spearman \(\rho=0.97\). It nevertheless captures activation, buffer, output, and temporary tensor demand that FP32 parameter storage alone does not. Appendix Figure 8 shows that the peak-memory frontiers broadly mirror the parameter frontiers, while ResNet18 has the largest incremental inference peak at 9.46 MiB.
The corresponding accuracy versus peak-memory plots are provided in Appendix 8.2.
Table 6 reports median and p95 latency. ResNet18 is the fastest L4 model at 3.038 ms median latency, narrowly ahead of SqueezeNet1.1. MobileNetV3-Small is the fastest CPU model in both thread modes at 14.20 and 14.25 ms. MobileNetV4-Conv-S is approximately 59% slower than MobileNetV3-Small in one-thread median latency and approximately 32% slower in four-thread median latency.
RepViT-M1.0 scales strongly from 199.80 ms with one thread to 61.78 ms with four threads, but remains substantially slower than MobileNetV3-Small and EfficientNet-B0 on the evaluated CPU. SqueezeNet1.1 and ResNet18 also benefit substantially from four-thread execution. These results show that parallel scaling and operator support differ materially across architectures.
| Model | NVIDIA L4 (ms) | CPU, 1 thread (ms) | CPU, 4 threads (ms) | |||
|---|---|---|---|---|---|---|
| 2-3(lr)4-5(lr)6-7 | Median | p95 | Median | p95 | Median | p95 |
| EfficientNetV2-S | 22.086 | 22.860 | 213.12 | 215.95 | 112.94 | 131.57 |
| EfficientNet-B0 | 9.832 | 10.000 | 69.50 | 77.05 | 43.24 | 49.65 |
| RepViT-M1.0 | 16.113 | 16.766 | 199.80 | 202.26 | 61.78 | 68.54 |
| ResNet18 | 3.038 | 3.154 | 127.33 | 129.05 | 46.95 | 58.63 |
| MobileNetV3-Small | 6.444 | 6.594 | 14.20 | 14.64 | 14.25 | 14.90 |
| MobileNetV2 | 6.437 | 6.564 | 41.83 | 42.87 | 31.13 | 45.32 |
| MobileNetV4-Conv-S | 6.389 | 6.707 | 22.58 | 23.12 | 18.76 | 20.06 |
| ShuffleNetV2 x1.0 | 7.554 | 7.694 | 34.21 | 35.15 | 25.43 | 29.15 |
| SqueezeNet1.1 | 3.054 | 3.196 | 67.19 | 81.75 | 23.97 | 27.08 |
The execution-environment-specific ranking shift is substantial. ResNet18 ranks first on the L4 but seventh in one-thread CPU latency. MobileNetV3-Small ranks fifth on the L4 but first under both CPU thread modes. SqueezeNet1.1 ranks second on the L4, fifth with one CPU thread, and third with four CPU threads. Figure 4 visualizes these changes directly.
Descriptive Spearman rank correlations [27] reinforce the same conclusion. GMACs correlate strongly with CPU median latency, \(\rho=0.95\) for one thread and \(\rho=0.93\) for four threads, but weakly with L4 median latency, \(\rho=0.27\). L4 and one-thread CPU latency rankings correlate only moderately, \(\rho=0.40\). These values describe the nine-model sample and should not be generalized as population estimates.
The complete descriptive correlation table is provided in Appendix 8.3.
Point estimate Pareto frontiers describe all nondominated options but do not select one model without a deployment budget. Table 7 gives transparent illustrative profiles. The thresholds are examples derived from natural breakpoints in the evaluated model pool, not universal deployment standards.
| Profile | Budget | CIFAR-10 and CIFAR-100 | Tiny ImageNet |
|---|---|---|---|
| Very compact storage | At most 2M parameters | MobileNetV3-Small | MobileNetV3-Small |
| Severe arithmetic limit | At most 0.10 GMAC | MobileNetV3-Small | MobileNetV3-Small |
| Compact general deployment | At most 5M parameters and 0.50 GMAC | EfficientNet-B0 | EfficientNet-B0 |
| Low CPU latency | At most 25 ms, one thread | MobileNetV3-Small | MobileNetV3-Small |
| Accuracy-oriented deployment | No strict compact budget | EfficientNetV2-S | RepViT-M1.0 |
EfficientNet-B0 is notable because it belongs to each independently calculated bivariate accuracy-resource frontier for all six resource dimensions: parameters, GMACs, L4 latency, one-thread CPU latency, four-thread CPU latency, and peak CUDA allocated memory. This does not make it a universal winner, but it supports its role as the most consistently competitive intermediate-budget option in the evaluated pool.
Under the common schedule, ImageNet pretraining improves EfficientNet-B0 test accuracy by 5.38 points on CIFAR-10, 14.90 points on CIFAR-100, and 20.65 points on Tiny ImageNet. A fresh 100 epoch scratch schedule increases scratch accuracy by 2.09, 4.80, and 3.11 points relative to the separate 20 epoch schedule, but the pretrained model still leads by 3.29, 10.10, and 17.54 points.
The longer schedules require approximately five times the recorded training time: 159.65 versus 31.59 minutes on CIFAR-10, 154.31 versus 31.55 minutes on CIFAR-100, and 314.52 versus 63.55 minutes on Tiny ImageNet. The gains are therefore meaningful but expensive, especially on the higher-class-count datasets.
| Dataset | Pretrained test | Scratch 20 test | Scratch 100 test | Remaining gap | 20 epoch time (min) | 100 epoch time (min) | Time multiplier |
|---|---|---|---|---|---|---|---|
| CIFAR-10 | 97.35 | 91.97 | 94.06 | 3.29 | 31.59 | 159.65 | 5.06 |
| CIFAR-100 | 86.13 | 71.23 | 76.03 | 10.10 | 31.55 | 154.31 | 4.89 |
| Tiny ImageNet | 78.08 | 57.43 | 60.54 | 17.54 | 63.55 | 314.52 | 4.95 |
Under pretrained initialization, MobileNetV3-Small records higher observed test accuracy than MobileNetV4-Conv-S by 0.16 points on CIFAR-10, 0.53 points on CIFAR-100, and 1.04 points on Tiny ImageNet. The paired intervals cross zero for the two CIFAR differences, while the Tiny ImageNet interval is 0.23 to 1.84 points.
Under scratch training, MobileNetV3-Small leads by 2.55, 1.76, and 0.99 points. The paired intervals are 1.90 to 3.21, 0.86 to 2.67, and 0.14 to 1.84 points, with exact McNemar \(p<0.05\) for all three datasets. The learning curves show different behavior across tasks. MobileNetV3-Small leads throughout CIFAR-10, overtakes MobileNetV4-Conv-S on CIFAR-100, and remains close but ends higher on Tiny ImageNet. The cumulative-time curves do not reverse the final ordering.
The result supports a conditional conclusion. The evaluated later-generation MobileNetV4-Conv-S model package does not improve upon MobileNetV3-Small under the tested public checkpoints, downstream recipe, scratch budget, CPU environment, parameter budget, or GMAC budget. It does not establish that MobileNetV4 is universally inferior under architecture-specific training or on its intended deployment targets.
Figure 7 summarizes selected paired accuracy differences. Several small pretrained differences are not clearly separated from zero, including EfficientNetV2-S versus EfficientNet-B0 on CIFAR-10, EfficientNet-B0 versus RepViT-M1.0 on CIFAR-10, and MobileNetV3-Small versus MobileNetV4-Conv-S on both CIFAR datasets. In contrast, RepViT’s Tiny ImageNet advantages and the three MobileNet scratch differences are supported by intervals that exclude zero.
The numerical paired-comparison table is provided in Appendix 8.4.
The results do not support a monotonic relationship between release year and downstream performance. RepViT-M1.0 provides a clear selective gain on Tiny ImageNet, where it leads top-1 accuracy, macro F1, and top-5 accuracy. EfficientNetV2-S remains the highest observed CIFAR model. MobileNetV4-Conv-S provides the opposite result: it does not exceed MobileNetV3-Small under the evaluated pretrained or scratch conditions and is less favorable in parameters, GMACs, peak CUDA memory, and CPU latency.
Accuracy rankings are highly consistent between CIFAR-10 and CIFAR-100, with Spearman \(\rho=0.98\), but less consistent between either CIFAR dataset and Tiny ImageNet, with \(\rho=0.77\) and \(\rho=0.73\). This supports the view that later designs may improve a particular target task without becoming universal winners.
The benchmark supports several model profiles rather than one universal ranking. EfficientNetV2-S is the accuracy-oriented choice for the CIFAR tasks when its higher resource demand is acceptable. RepViT-M1.0 is the accuracy-oriented Tiny ImageNet choice. EfficientNet-B0 is the most consistently Pareto-efficient intermediate option. It remains close to the dataset leaders while belonging to each independently calculated bivariate accuracy-resource frontier.
MobileNetV3-Small provides the highest observed accuracy under the evaluated severe arithmetic and one-thread CPU budgets. It has the lowest GMAC count, leads both CPU latency modes, and records substantially stronger accuracy than the smallest-storage SqueezeNet1.1. SqueezeNet remains useful when parameter storage or peak CUDA allocation is the dominant constraint, but its predictive losses become large on CIFAR-100 and Tiny ImageNet.
Parameters, GMACs, latency, and peak tensor memory answer different deployment questions. GMACs correspond strongly to latency on the tested CPU, but poorly to L4 latency. Dense conventional operations in ResNet18 execute efficiently on the L4 despite high arithmetic demand. MobileNetV3-Small shows the reverse execution-environment shift: it is not the fastest L4 model, but it decisively leads the CPU benchmark. Operator fusion, memory access, kernel selection, parallelism, and thread overhead therefore influence realized latency in addition to nominal computation.
The execution-environment comparison should remain narrow. The findings describe one NVIDIA L4 FP16 autocast environment and one AMD Ryzen 5 5500U FP32 environment. They do not establish rankings on mobile phones, ARM boards, neural processing units, DSPs, or production deployment runtimes.
The EfficientNet-B0 results separate initialization from training duration. Pretraining yields large advantages under the common schedule. The fresh 100 epoch scratch schedules produce higher observed accuracy than the separate shorter schedules, but the remaining gaps are larger on CIFAR-100 and Tiny ImageNet. The approximately fivefold increase in recorded training time also shows why equal epoch counts and equal computational exposure are not equivalent concepts.
The MobileNet experiment addresses a different question: comparative learning under one constrained schedule. MobileNetV3-Small reaches stronger final scratch performance and generally stronger learning trajectories. This reduces, but does not eliminate, checkpoint provenance as an explanation for the pretrained ordering.
Paired test-set analysis prevents small point differences from being overstated. The CIFAR-10 EfficientNetV2-S versus EfficientNet-B0 difference and the CIFAR pretrained MobileNet differences are not clearly separated by finite test-set uncertainty. The Tiny ImageNet RepViT advantages and all three MobileNet scratch advantages are better supported. However, these results condition on one completed training run and cannot estimate the variance produced by independent initialization, data order, or nondeterministic operations.
The principal limitation is the absence of repeated training seeds. All predictive results are point estimates from one completed training seed, and paired test-set intervals do not substitute for training-run uncertainty. Close differences should therefore be interpreted cautiously.
The shared recipe provides protocol equality rather than architecture-specific opportunity equality. The separate EfficientNet-B0 scratch schedules use the same nominal seed but are not a matched-initial-state intervention; they differ in schedule length, cosine trajectory, and early-stopping patience, so the comparison is descriptive rather than a causal estimate of adding epochs. Individual networks may benefit from different learning rates, warmup, augmentation, regularization, distillation, input resolution, or longer schedules. The MobileNet scratch comparison measures behavior under one fixed budget and does not claim fully optimized scratch performance.
The pretrained benchmark remains confounded by upstream checkpoint provenance. RepViT uses a distilled checkpoint with dual classifier heads, and MobileNetV4-Conv-S uses a public timm checkpoint rather than an official Google checkpoint. The main comparison is therefore best understood as public model-package selection under common downstream conditions rather than pure causal isolation of architecture.
The datasets are all natural-image classification benchmarks, and Tiny ImageNet is related to the ImageNet pretraining source. The fixed 224 pixel adaptation does not test native-resolution behavior. The model pool is representative but not exhaustive and omits several recent lightweight CNNs. The resource analysis does not include quantization, energy, mobile-device execution, peak training memory, or optimized runtimes such as TensorRT, ONNX Runtime, and Core ML. The CPU benchmark includes three timing rounds, whereas the L4 distribution is based on one warmed sequence of 500 timed passes; the two environments therefore provide different forms of repeatability evidence.
Within the evaluated model pool, datasets, and execution environments, the results do not support a consistent advantage from architectural recency. EfficientNetV2-S records the highest observed CIFAR accuracy, RepViT-M1.0 leads Tiny ImageNet, EfficientNet-B0 belongs to every independently calculated bivariate accuracy-resource frontier, and MobileNetV3-Small provides the highest observed accuracy under the evaluated severe arithmetic and CPU-latency budgets.
The results also show why a single resource proxy is insufficient. GMACs correspond strongly to latency on the tested CPU but weakly to L4 latency. Peak CUDA tensor memory largely follows parameter count, while incremental allocation still reveals differences in activation and temporary-tensor demand. The fresh 100 epoch EfficientNet-B0 schedules improve over the separate shorter scratch schedules but leave substantial pretrained gaps, and the evaluated MobileNetV4-Conv-S package does not improve upon MobileNetV3-Small under the tested checkpoints or scratch protocol.
The appropriate conclusion is conditional rather than universal. Newer designs can improve particular datasets and deployment profiles, but model selection continues to depend on checkpoint provenance, target data, optimization budget, resource definition, and the complete execution environment.
| Model | CIFAR-10 (%) | CIFAR-100 (%) | Tiny ImageNet (%) |
|---|---|---|---|
| EfficientNetV2-S | 99.84 | 96.69 | 91.79 |
| EfficientNet-B0 | 99.87 | 97.29 | 92.42 |
| RepViT-M1.0 | 99.82 | 96.58 | 93.09 |
| ResNet18 | 99.72 | 95.42 | 86.33 |
| MobileNetV3-Small | 99.76 | 96.48 | 89.32 |
| MobileNetV2 | 99.71 | 95.92 | 89.57 |
| MobileNetV4-Conv-S | 99.68 | 95.54 | 86.98 |
| ShuffleNetV2 x1.0 | 99.70 | 96.19 | 88.69 |
| SqueezeNet1.1 | 99.16 | 93.65 | 83.13 |
| Comparison | Spearman \(\rho\) |
|---|---|
| GMACs versus L4 median latency | 0.27 |
| GMACs versus CPU one-thread median latency | 0.95 |
| GMACs versus CPU four-thread median latency | 0.93 |
| L4 versus CPU one-thread latency | 0.40 |
| L4 versus CPU four-thread latency | 0.50 |
| CPU one-thread versus CPU four-thread latency | 0.95 |
| Parameters versus peak CUDA allocated memory | 0.97 |
| CIFAR-10 versus CIFAR-100 accuracy rank | 0.98 |
| CIFAR-10 versus Tiny ImageNet accuracy rank | 0.77 |
| CIFAR-100 versus Tiny ImageNet accuracy rank | 0.73 |
| Setting | Dataset | Model A | Model B | Difference | 95% interval | McNemar \(p\) |
|---|---|---|---|---|---|---|
| Pretrained | CIFAR-10 | EfficientNetV2-S | EfficientNet-B0 | 0.22 | [-0.10, 0.54] | 0.196 |
| Pretrained | CIFAR-100 | EfficientNetV2-S | EfficientNet-B0 | 0.85 | [0.22, 1.48] | 0.0085 |
| Pretrained | Tiny ImageNet | RepViT-M1.0 | EfficientNetV2-S | 1.14 | [0.42, 1.85] | 0.0018 |
| Pretrained | Tiny ImageNet | RepViT-M1.0 | EfficientNet-B0 | 1.79 | [1.06, 2.51] | \(<0.00001\) |
| Pretrained | CIFAR-10 | MobileNetV3-Small | MobileNetV4-Conv-S | 0.16 | [-0.24, 0.56] | 0.458 |
| Pretrained | CIFAR-100 | MobileNetV3-Small | MobileNetV4-Conv-S | 0.53 | [-0.16, 1.22] | 0.139 |
| Pretrained | Tiny ImageNet | MobileNetV3-Small | MobileNetV4-Conv-S | 1.04 | [0.23, 1.84] | 0.0124 |
| Scratch | CIFAR-10 | MobileNetV3-Small | MobileNetV4-Conv-S | 2.55 | [1.90, 3.21] | \(<0.00001\) |
| Scratch | CIFAR-100 | MobileNetV3-Small | MobileNetV4-Conv-S | 1.76 | [0.86, 2.67] | 0.00015 |
| Scratch | Tiny ImageNet | MobileNetV3-Small | MobileNetV4-Conv-S | 0.99 | [0.14, 1.84] | 0.0255 |
Table 12 summarizes the point estimate frontier members used in the main analysis.
| Dataset | Resource | Frontier members in increasing resource order |
|---|---|---|
| CIFAR-10 | Parameters | SqueezeNet1.1, ShuffleNetV2 x1.0, MobileNetV3-Small, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-10 | GMACs | MobileNetV3-Small, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-10 | L4 median latency | ResNet18, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-10 | CPU median latency | MobileNetV3-Small, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-10 | Peak CUDA memory | SqueezeNet1.1, MobileNetV3-Small, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-100 | Parameters | SqueezeNet1.1, ShuffleNetV2 x1.0, MobileNetV3-Small, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-100 | GMACs | MobileNetV3-Small, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-100 | L4 median latency | ResNet18, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-100 | CPU median latency | MobileNetV3-Small, EfficientNet-B0, EfficientNetV2-S |
| CIFAR-100 | Peak CUDA memory | SqueezeNet1.1, MobileNetV3-Small, EfficientNet-B0, EfficientNetV2-S |
| Tiny ImageNet | Parameters | SqueezeNet1.1, ShuffleNetV2 x1.0, MobileNetV3-Small, MobileNetV2, EfficientNet-B0, RepViT-M1.0 |
| Tiny ImageNet | GMACs | MobileNetV3-Small, MobileNetV2, EfficientNet-B0, RepViT-M1.0 |
| Tiny ImageNet | L4 median latency | ResNet18, MobileNetV4-Conv-S, MobileNetV2, EfficientNet-B0, RepViT-M1.0 |
| Tiny ImageNet | CPU median latency | MobileNetV3-Small, MobileNetV2, EfficientNet-B0, RepViT-M1.0 |
| Tiny ImageNet | Peak CUDA memory | SqueezeNet1.1, MobileNetV3-Small, MobileNetV2, EfficientNet-B0, RepViT-M1.0 |
Training augmentation uses bicubic random resized cropping with a scale of 0.80 to 1.00 and an aspect ratio of 0.90 to 1.10, horizontal flipping with probability 0.5, and color jitter values of 0.15 for brightness, contrast, and saturation and 0.05 for
hue. Evaluation uses direct bicubic resizing to \(224\times224\) without random cropping. ImageNet channel means and standard deviations are used for normalization. Training uses seeded data loading,
drop_last=True, deterministic cuDNN mode, and disabled cuDNN benchmarking. One multiply-accumulate operation is counted as one MAC and reported in GMACs accordingly.