TMP: Tree-structured Mixed-policy Pruning for Large-scale Image Generation and Editing

Peizhen Zhang, Yang Li, Xunsong Li, Songtao Liu3, Zewen Liu, Qiangqiang Hu, Guotong Guo, Jupeng Ding, Yifu Sun, coopersli, Jian Zhang, Zhao Zhong4, Liefeng Bo
Multimodal Model Department, Tencent


Abstract

Modern image generation model rapidly grows their sizes to meet high-fidelity image synthesis. However, they gradually become unaffordable for their enormous parameter consumption and computation budget that lead to massive resources requirement and gpu memory footprint. In this paper, we propose TMP, the first Tree-structured Mixed-policy Pruning framework that generalizes prevalent image tasks (T2I and TI2I) and architectures (Mixture-of-Experts (MoE) and Diffusion transformer (DiT)). It could be applied to the step-distilled models and contribute as the last stage. We perform experiments upon current open-sourced SOTA HunyuanImage-3.0 instruct and a popular efficient model Z-Image turbo. The proposed pruning framework manages to compress HunyuanImage 3.0 from 80B to 20B parameters at 75% reduction ratio, sacrificing limited generation quality. We also optimize to enable the inference of the pruned 20B version of HunyuanImage 3.0 on a single 24GB 4090 GPU by engineering skills. The inference script and model weight have been integrated into the existing HunyuanImage3.0 open-source github1  and huggingface2  repository. Besides, we prove the efficacy of TMP by compressing Z-Image turbo from 6B to 4B (33% reduction) with negligible degradation.

Figure 1: Single-reference image editing examples comparing the originalHunyuanImage-3.0 80B model and our 20B pruned model.Despite a 4\times reduction in model size, the pruned model preservesediting fidelity and visual quality across diverse editing instructions.
Figure 2: Dual-reference image editing examples comparing the originalHunyuanImage-3.0 80B model and our 20B pruned model.The pruned model preserves multi-image reasoning andreference-following capabilities despite substantial model compression.

1 Introduction↩︎

In the past one year, closed-sourced commercial models like Nano-Banana [1] and GPT-Image [2] series dominate the image synthesis scenario. Recently, excellent open-sourced models like HunyuanImage-3.0 [3] and FLUX.2 [4] have been developed to catch them. However, they are normally too costly for practical usage. Typically, HunyuanImage-3.0, albeit achieve SOTA performance among open-sourced models, consume 80B parameters and is recommended to infer at the lowest requirement of 3\(\times\)​80GB GPUs, hindering its popularity in community. To this end, structured pruning is one of the most straightforward methods to alleviate the burden by compressing the original model into a smaller one and is normally hardware-friendly. However, existing structured pruning methods are customized deeply for DiT architecture whereas HunyuanImage-3.0 is built upon an MoE-based decoder-only LLM, i.e., Hunyuan-A13B. The challenge lies in the fact that only an extremely high degree of parameter reduction can make the resulting model lightweight enough to run on low-cost devices. Existing methods could hardly achieve this (normally meet severe performance drop when reducing ratio \(\geq 15\%\)). In this paper, we propose TMP, a structured pruning framework designed to overcome these challenges. TMP involves a joint pruning strategy, followed by a two-stage post-prune recovery training. The algorithm could be directly applied to the pruning of DiT-based architecture that we generalize to Z-Image turbo without much effort.

The proposed pruning framework includes a structured pruning followed by a recovery training. The recovery training has two stages: (1) A tree-structured local feature supervision running in divide-and-conquer manner. It is instantiated as a combination of both off-policy and on-policy distillation (OPD). (2) A brief output distillation to match the flow matching velocity prediction. Normally, if network parameters have undergone pruning, each layer-wise inference would inevitably suffer from representation misalignment. In the end-to-end perspective, such errors will accumulate across layers and ultimately lead to model collapse due to cascading error propagation. Preliminary works like [5] introduce local supervision to mitigate this issue. Basically, tokens processed by a teacher layer are fed into its corresponding succeeding student layer. The supervision is imposed by matching the student outputs against the outputs of the subsequent teacher layer. Recently, some structured pruning works like FastFLUX [6] and PPCL [7] adopt the local supervision design in their algorithm. Our method, TMP-Pruner differs from theirs in two key aspects, which inspire its name: (a) Tree-structured Merging: We maintain the intervals with a tree-structured merging. Hierarchically, two adjacent intervals (layer sequences) may be merged into a single one, in which case their separate distillation objectives are replaced by a unified loss after merging. This induces a curriculum-learning-like behavior, the model gradually shifts toward newly formed unified objectives as the local supervision regions get progressively merged. Eventually, the objective will reduce to a single end-to-end token representation matching. (b) On-policy Distillation: We extra introduce an on-policy feature distillation and extend it to a mixed policy distillation that further align the pruned model to the original model.

In short, our contributions are fourfold:

  • We propose a unified structured pruning framework that involves a heuristic pruning followed by a mixed policy distillation that refines the teacher-student alignment in intermediate token space.

  • We devise a novel tree-structured manner of curriculum learning scheme to ease the optimization.

  • Experiments demonstrate that our proposed framework could be applied to SOTA open-sourced text-to-image/image-editing model at a huge parameter reduction ratio, making it become affordable, with limited losses of generation quality.

  • Our framework is versatile towards different backbone architectures (MoE and DiT), model sizes (80B HunyuanImage-3.0 and 6B Z-Image turbo) and image tasks (image editing and text-to-image generation).

2 Related Work↩︎

Generative models for high-fidelity image generation have become popular in recent years. They tend to obtain better performance at the cost of increasing model capacity, typically parameter consumption. This imposes enormous memory consumption that either prohibits the model from running on a single consumer-level GPU or leads to low throughput in large-scale deployment. Structured pruning is intuitive to relieve this by slimming the network architecture to obtain a lightweight version. It involves a model pruning stage followed by a recovery training stage.

2.1 Structured Pruning↩︎

The ways to determine pruned model architecture could be divided into five types: (a) Importance metric (b) Probabilistic modeling (c) NAS-based sampling (d) Cheap replacement and (e) Time-step sensitive allocation. (a) Molchanov et al. [8] came up with a layer importance metric to determine which parameters to prune. Diff-Pruning [9] adopted a time-step-aware version. LD-Pruner [10] further proposed an operator-level metric. Xie et al. [11] used cosine similarity given input and output of each layer to measure the importance. OBS-DIFF [12] invented a second-order sensitivity based on Hessian matrix computation. PPCL [7] combined linear probing with CKA-analysis to conduct contiguous layer set removal. (b) EcoDiff [13] learned a differentiable mask for sparsification. TinyFusion [14] proposed a differentiable masked sampling that learned shallow DiT sub-networks to reduce denoising computation through adaptive depth reduction. (c) ALTER [15] trained a supernet and used dynamic routing for block and layer pruning. E-DiT [16] trained differentiable router modules for block skipping and dimension reduction in linear layers. (d) Unlike hard pruning, some works conducted cheap substitution of existing modules. FastFLUX [6] progressively replaced the DiT blocks as linear modules that will join a sandwich training scheme. Amber-Image [17] proposed a method that could merge two-stream architecture in MMDiT [18] into a unified stream. However, the sandwich training in FastFLUX demands constructing new local supervision datasets every time when a transformer block get replaced. The two-stream merging strategy in Amber-Image is only specialized for two-stream DiT architecture. (e) Some works introduced a time step-sensitive design related to the reverse diffusion nature during runtime. MosaicDiff [19] handcrafted a sparsity curve that set model variants with different sizes under different time steps. Diff-ES [20] automated it by setting the sparsity schedule with evolutionary search. The former requires deploying a series of pruning variants that is unfriendly towards low-memory devices. The latter attains the entire supernet during inference that demands the same memory budget as the original model in the worst-case scenarios.

2.2 Recovery Training↩︎

After architecture pruning, recovery training will transfer knowledge (KD) [21] from the original model to the pruned model. Besides naive KD, some works emphasized on optimizing the loss design. HierarchicalPrune [22] used reverse loss weighting to suppress the learning of layers with high importance. IGSM [23] proposed a second-order Jacobian matching loss inspired by Finite-Time Lyapunov Exponents that makes the model more robust towards input perturbation in the denoising trajectory. However, the loss weighting strategy requires human efforts for delicate adjustment. The second-order metric brings about heavy computation budget. Regardless of the choice of loss function, the optimization process is fundamentally driven by the low-level intermediate feature representations, higher-level latent or noise estimation. These methods, in principle, still exhibit off-policy behavior due to their dependence on teacher-generated trajectories. Recently, in language modeling tasks, on-policy distillation (OPD) has been shown to achieve faster and more effective distillation than off-policy distillation [24], [25]. The key is that student rollout affects future supervision distribution. In this paper, we devise an on-policy paradigm that constructs supervisory target given student intermediate representations. It runs in mixed mode that alternates off-and-on policy distillation.

Moreover, existing structured pruning works are restricted to text-to-image task with DiT architecture, not involving architecture like MoE adopted by modern SOTA models and image editing task. Our method is able to prune HunyuanImage-3.0 at high parameter reduction ratio.

3 Methodology↩︎

3.1 Structured Pruning↩︎

HunyuanImage-3.0 [3] is an 80B MoE-based image editing model built upon Hunyuan-A13B [26] where the MoE layers encompass 98% of the parameters. We perform a joint expert and width pruning that add up to 75% of parameter count reduction, resulting in a 20B pruned model. (1) width pruning: We apply magnitude pruning to reduce the intermediate size of the MLPs in MoE-FFN from 3072 to 2048. (2) expert pruning: We aggregate the gating scores over a calibration set and keep only 24 out of 64 experts per layer. Z-Image turbo [27] is a 6B single-stream DiT-based image generation model. We apply width pruning to it that is highly similar to what we have done above for HunyuanImage-3.0. We reduce the expansion rate of the MLP by \(37.5\%\) and achieve an overall 33% of parameter count reduction. This results in a 4B pruned model.

3.2 Recovery Training↩︎

The recovery training runs in two-staged fashion. The first stage is about intermediate token distillation. The second stage further refines by velocity prediction distillation. We parameterize the full model as \(F_\theta\). It consists of three components: a collection of input encoders and embedders \(E_\theta\) (VAE, timestep embedders and any other involved encoders), a transformer backbone \(\pi_\theta\) and two parallel prediction heads — a velocity head \(v_\theta\) for image latents and an autoregressive head \(p_\theta\) for text tokens. Across all training stages, \(E_\theta\) is kept frozen. In the mixed-policy feature-distillation stage we unfreeze only \(\pi_\theta\); in the subsequent prediction-alignment stage we additionally unfreeze the heads \({v_\theta, p_\theta}\). We generalize \(\pi_\theta\) (with a slight abuse of notation) from the next-token distribution to the sequence of intermediate features produced by the student backbone, thereby moving the on-/off-policy distinction from the output layer down to intermediate representations.

3.2.1 Tree-Structured Mixed-policy Feature Distillation↩︎

Supposed there are total \(L\) layers in each model, we set \(L\) intervals correspondingly where each interval contains one layer at the beginning. We conduct the distillation fashion in a tree-structured, divide-and-conquer manner.

Mixed policy distillation. Supposed there are M intervals in current training iteration, we illustrate the local distillation of interval \(i\) without loss of generality. It satisfies \(L=\sum_{i=1}^{M} B_i\) where \(B_i\) denotes the number of consecutive layers covered by interval \(i\) . Supposed interval \(i\) begins at layer \(l\), we introduce how local feature distillation is conducted between the original model \(\Pi_*\) and the pruned model \(\Pi_\theta\). In each iteration, we feed the token sequence \(x \sim p_{\mathrm{data}}\) into the teacher model to perform end-to-end forward propagation. In this procedure, we collect input tokens into layer \(l\), said \(X_*^{l-1}\) and the output tokens by layer \(l+B_i-1\), said \(X_*^{l+B_{i}-1} \equiv \Pi_{*}^{1 \rightarrow l+B_i-1}(x)\). For student, instead of regular forwarding, we feed the input tokens into interval \(i\) of teacher into that of the student to get the output tokens. We make it learn towards corresponding teacher output tokens, forming off-policy distillation:

\[\label{eq:offpolicydistill} \mathcal{L}_{\mathrm{off}}^{(i)} = \min_{\theta} \mathbb{E}_{x \sim p_{\mathrm{data}}, X_{*}^{l-1} \sim \Pi_{*}^{1 \rightarrow l-1}(x)} \left[ \frac{1}{N} \mathcal{D} \left( X_*^{l+B_{i}-1}, \Pi_{\theta}^{\,l \rightarrow l+B_i-1}(X_{*}^{\,l-1}) \right) \right]\tag{1}\]

where \(N\) is the number of tokens. \(\mathcal{D}(\cdot, \cdot)\) represents the feature discrepancy. We instantiate it as L2 distance. Besides the off-policy distillation, we enhance the intermediate token representation learning towards the original model by introducing on-policy distillation to the feature distillation. We instantiate the on-policy feature distillation by swapping the roles described above. The key lies in utilizing the student roll-outs to join the learning target construction. We condition on the imperfect student representation and modulate it by local original model consecutive layers to construct learning target and define the distillation as:

\[\label{eq:onpolicydistill} \mathcal{L}_{\mathrm{on}}^{(i)} = \min_{\theta} \mathbb{E}_{x \sim p_{\mathrm{data}}, X_{\theta}^{l-1} \sim \Pi_{\theta}^{1 \rightarrow l-1}(x)} \left[ \frac{1}{N} \mathcal{D} \left( \mathrm{sg} \left[ \Pi_{*}^{\,l \rightarrow l+B_i-1}(X_{\theta}^{l-1})\right], X_{\theta}^{l+B_{i}-1} \right) \right]\tag{2}\] where \(\mathrm{sg}\left[\cdot\right]\) denotes the stop-gradient operator. We combine both mimicking loss, leading to a mixed policy distillation manner. The ultimate learning objective considering all current \(M\) intervals is \(\mathcal{L}_{\mathrm{mixed}}=\frac{1}{M}\Sigma_{i=1}^{M}\left[\mathcal{L}_{off}^{(i)}+\mathcal{L}_{on}^{(i)}\right]\)

Tree-structured Interval Merging. To gradually mitigate the propagated error towards end-to-end optimization, we adopt a bottom-up binary tree fashion of interval merging. As every \(K\) training iterations go by, two adjacent intervals are merged into one, halving the interval number.

3.2.2 Prediction Alignment↩︎

After hierarchical mixed policy feature distillation, we use output distillation5 to further align the original model and the pruned model. It is accomplished by distilling the velocity prediction between both models as shown in Eq. 3 .

\[\label{eq:velocity95distill} \mathcal{L}_{\mathrm{velocity}} = \mathbb{E}_{x_t,t,c} \left[ \left\| v_{\theta}(x_t,t,c) - v_{*}(x_t,t,c) \right\|_2^2 \right]\tag{3}\]

4 Experiments↩︎

Figure 3: Human preference comparison between pruned 20B v.s. original HunyuanImage-3.0 80B on private image editing benchmark. The evaluation shows a limited performance degradation (-2.5\%) on human preference.

4.1 About HunyuanImage-3.0↩︎

(a) Original model v.s. pruned model Fig. 3 shows the comparison between HunyuanImage-3.0 original 80B model and its pruned 20B model by our proposed structured pruning method. Notably, there is only -2.5% of human preference dropping. The comparison is made upon private image editing dataset. Fig. 1 and Fig. 2 have shown some image editing results comparison between these two models.

(b) Memory-Efficient Inference To enable deployment on consumer-grade GPUs, we develop a memory-efficient inference framework based on FP8 quantization and dynamic module offloading. The main model is quantized to FP8 precision, significantly reducing the memory footprint. To avoid memory spikes during initialization, model weights are first loaded into CPU memory and selectively transferred to the GPU according to the execution stage. In addition, several memory-intensive components, including the visual encoder and VAE, are activated on demand and immediately offloaded after use. Intermediate activations and CUDA caches are also released between major generation stages to minimize memory fragmentation and peak allocation. With these optimizations, our model can perform 1024\(\times\)​1024 image generation with 8 sampling steps on a single NVIDIA RTX 4090 (24 GB), requiring less than 24 GB of GPU memory at peak usage.

Table 1: Quantitative evaluation results on OneIG-EN.
Model Alignment Text Reasoning Style Diversity Overall
Qwen-Image 20B  [28] 0.882 0.891 0.306 0.418 0.197 0.539
Z-Image 6B [27] 0.881 0.987 0.280 0.387 0.194 0.546
Z-Image-Turbo 6B [27] 0.840 0.994 0.298 0.368 0.139 0.528
PPCL-OPPO-10B [7] 0.839 0.860 0.249 0.359 0.121 0.485
Amber-Image-10B [17] 0.867 0.938 0.278 0.298 0.137 0.504
Amber-Image-6B [17] 0.829 0.917 0.284 0.287 0.135 0.490
Z-Image-Turbo pruned-4B (Ours) 0.840 0.980 0.305 0.364 0.161 0.530
Table 2: Quantitative evaluation results on OneIG-ZH.
Model Alignment Text Reasoning Style Diversity Overall
Qwen-Image 20B [28] 0.825 0.963 0.267 0.405 0.279 0.548
Z-Image 6B [27] 0.793 0.988 0.266 0.386 0.243 0.535
Z-Image-Turbo 6B [27] 0.782 0.982 0.276 0.361 0.134 0.507
PPCL-OPPO-10B [7] 0.854 0.878 0.268 0.365 0.130 0.499
Amber-Image-10B [17] 0.798 0.975 0.221 0.362 0.153 0.502
Amber-Image-6B [17] 0.779 0.953 0.208 0.345 0.143 0.486
Z-Image-Turbo pruned-4B (Ours) 0.784 0.946 0.261 0.352 0.158 0.500
Table 3: Quantitative evaluation results on LongText-Bench.
Model LongText-Bench-EN LongText-Bench-ZH
Qwen-Image 20B [28] 0.943 0.946
Z-Image [27] 0.935 0.936
Z-Image-Turbo [27] 0.917 0.926
PPCL-OPPO-10B [7] 0.871 0.885
Amber-Image-10B [17] 0.911 0.915
Amber-Image-6B [17] 0.870 0.876
Z-Image-Turbo pruned-4B (Ours) 0.889 0.882

4.2 About Z-Image turbo↩︎

Both PPCL-OPPO and Amber-Image are among the best T2I models obtained via structured pruning in recent literature. They are built upon Qwen-Image (better than our experimenting teacher Z-Image-Turbo in the metrics shown in below involved benchmarks). Table 1, Table 2 and Table 3 show the comparison of the pruned Z-Image Turbo 4B model by our method to others on OneIG-ZH/EN [29] and LongText benchmarks [30] respectively. Notably, the 4B-pruned model outperforms the pruned models by these methods (OneIG-EN) or is comparable (OneIG-ZH, LongText-Bench) to them despite containing 1.5 to 2.5 \(\times\) fewer parameters.

5 Conclusion↩︎

We presented TMP, a structured pruning framework engineered to compress modern, large-scale image synthesis models. By advancing beyond standard distillation with a tree-structured merging strategy and mixed-policy feature alignment, the proposed framework effectively scales to intricate MoE-based and DiT architectures. TMP achieves an aggressive compression ratio—most notably reducing HunyuanImage-3.0 from 80B to 20B parameters—while preserving high-fidelity generation quality. This work bridges the gap between state-of-the-art generative capabilities and resource-constrained hardware deployment. Moving forward, we aim to adapt this compression paradigm to other generative modalities and investigate deeper hardware-level optimization.

References↩︎

[1]
Google, Commercial image generation and editing system“Nano-banana.” https://deepmind.google/, 2025.
[2]
OpenAI, Commercial image generation model“GPT-image.” https://openai.com, 2025.
[3]
S. Cao et al., “Hunyuanimage 3.0 technical report,” arXiv preprint arXiv:2509.23951, 2025.
[4]
Black Forest Labs, Accessed: 2026-05-12“FLUX.2.” https://bfl.ai/blog/flux-2, 2025.
[5]
Y. Wang, Z. Ni, S. Song, L. Yang, and G. Huang, “Revisiting locally supervised learning: An alternative to end-to-end training,” arXiv preprint arXiv:2101.10832, 2021.
[6]
F. Cai, Y. Guo, J. Li, W. Li, J. Chen, and X. Fang, “FastFLUX: Pruning FLUX with block-wise replacement and sandwich training,” in Proceedings of the AAAI conference on artificial intelligence, 2026, vol. 40, pp. 2507–2515.
[7]
J. Ma, Q. Peng, X. Zhu, P. Xie, C. Chen, and H. Lu, “Pluggable pruning with contiguous layer distillation for diffusion transformers,” arXiv preprint arXiv:2511.16156, 2025.
[8]
P. Molchanov, A. Mallya, S. Tyree, I. Frosio, and J. Kautz, “Importance estimation for neural network pruning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 11264–11272.
[9]
G. Fang, X. Ma, and X. Wang, “Structural pruning for diffusion models, 2023,” URL https://arxiv. org/abs/2305.10924, 2023.
[10]
T. Castells, H.-K. Song, B.-K. Kim, and S. Choi, “Ld-pruner: Efficient pruning of latent diffusion models using task-agnostic insights,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 821–830.
[11]
E. Xie et al., “Sana 1.5: Efficient scaling of training-time and inference-time compute in linear diffusion transformer,” arXiv preprint arXiv:2501.18427, 2025.
[12]
J. Zhu, H. Wang, M. Su, Z. Wang, and H. Wang, “OBS-diff: Accurate pruning for diffusion models in one-shot,” arXiv preprint arXiv:2510.06751, 2025.
[13]
Y. Zhang et al., “Effortless efficiency: Low-cost pruning of diffusion models,” arXiv e-prints, pp. arXiv–2412, 2024.
[14]
G. Fang, K. Li, X. Ma, and X. Wang, “Tinyfusion: Diffusion transformers learned shallow,” in Proceedings of the computer vision and pattern recognition conference, 2025, pp. 18144–18154.
[15]
X. Yang et al., “Alter: All-in-one layer pruning and temporal expert routing for efficient diffusion generation,” Advances in Neural Information Processing Systems, vol. 38, pp. 128571–128599, 2026.
[16]
J. Wang et al., “Elastic diffusion transformer,” arXiv preprint arXiv:2602.13993, 2026.
[17]
C. Yang, T. Li, Y. Zhang, and J. Gao, “Amber-image: Efficient compression of large-scale diffusion transformers,” arXiv preprint arXiv:2602.17047, 2026.
[18]
P. Esser et al., “Scaling rectified flow transformers for high-resolution image synthesis,” in Forty-first international conference on machine learning, 2024.
[19]
B. Guo, S. Tang, C. Zeng, and Z. Shen, “Mosaicdiff: Training-free structural pruning for diffusion model acceleration reflecting pretraining dynamics,” in Proceedings of the IEEE/CVF international conference on computer vision, 2025, pp. 1655–1664.
[20]
Z. Liu, S. Tang, Z. Wu, X. Yuan, and Z. Shen, “Diff-ES: Stage-wise structural diffusion pruning via evolutionary search,” arXiv preprint arXiv:2603.05105, 2026.
[21]
G. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531, 2015.
[22]
Y. D. Kwon, R. Li, S. Li, D. Li, S. Bhattacharya, and S. I. Venieris, “Hierarchicalprune: Position-aware compression for large-scale diffusion models,” in Proceedings of the AAAI conference on artificial intelligence, 2026, vol. 40, pp. 22716–22724.
[23]
C. Zheng and E. Shlizerman, “IGSM: Improved geometric and sensitivity matching for finetuning pruned diffusion models,” arXiv preprint arXiv:2506.05398, 2025.
[24]
K. Lu and T. M. Lab, https://thinkingmachines.ai/blog/on-policy-distillation“On-policy distillation,” Thinking Machines Lab: Connectionism, 2025, doi: 10.64434/tml.20251026.
[25]
R. Agarwal et al., “On-policy distillation of language models: Learning from self-generated mistakes,” in ICLR, 2024.
[26]
Tencent Hunyuan Team, “Hunyuan-A13B technical report,” Tencent, 2025. [Online]. Available: https://github.com/Tencent-Hunyuan/Hunyuan-A13B.
[27]
H. Cai et al., “Z-image: An efficient image generation foundation model with single-stream diffusion transformer,” arXiv preprint arXiv:2511.22699, 2025.
[28]
C. Wu et al., “Qwen-image technical report,” arXiv preprint arXiv:2508.02324, 2025.
[29]
J. Chang et al., “Oneig-bench: Omni-dimensional nuanced evaluation for image generation,” arXiv preprint arXiv:2506.07977, 2025.
[30]
Z. Geng et al., “X-omni: Reinforcement learning makes discrete autoregressive image generative models great again,” arXiv preprint arXiv:2507.22058, 2025.

  1. https://github.com/Tencent-Hunyuan/HunyuanImage-3.0.↩︎

  2. https://huggingface.co/tencent/HunyuanImage-3.0.↩︎

  3. Corresponding author↩︎

  4. Project leader↩︎

  5. For HunyuanImage-3.0 which demands self-recaption capability, we extra apply a Kullback-Leibler (KL) divergence between the next-token prediction by the language modeling heads of both models↩︎