Matching Tasks to Objectives: Fine-Tuning and Prompt-Tuning Strategies for Encoder-Decoder Pre-trained Language Models


Table 1: Comperison of proposed method with related works on CommonSenseQA and OpenBookQA datasets.
Methods CommonSenseQA OpenBookQA
2-4 (lr)5-7
30
425
1700
30
298
991
RoBERTa-large 29.66 58.47 37.00 41.47
MHGRN 29.01 50.23 38.00 39.73
QA-GNN 32.95 50.15 33.53 42.40
GreaseLM 22.80\(^\star\) 63.09\(^\star\) 39.00\(^\star\) 42.20\(^\star\)
GSC [@base-GSC] 31.02\(^\star\) 65.83\(^\star\) 29.60\(^\star\) 42.40\(^\star\)
SAFE 36.45 65.16 38.80 44.93
MVP-Tuning 48.99 67.12 39.60 56.00
Our Methods (Fine-Tuning)
T5-large (baseline) 21.04 23.12 65.15 23.34 43.08 55.31
T5-MP 25.11 25.53 59.02 21.26 54.42 \(\mathbf{57.52}^d\)
T5-MAP \(\textbf{48.21}_c^d\) \(60.05^{d}\) 68.12 \(35.84_c\) 40.24 \(44.3^{l}\)
T5-AP \(47.01^l\) \(\mathbf{61.22}^d\) \(66.25^d\) 39.71 46.72 48.63
Our Methods (Prompt-Tuning)
T5-PreMAP \(35.28^d\) \(\underline{52.35}^d\) \(\underline{59.21}^l\) \(22.63^l\) \(34.33^l\) \(33.15^l\)
T5-PreAP \(19.14^d\) \(33.41^l\) \(59.17\) \(3.15^d\) \(32.71^l\) \(33.27^l\)