TRUST: Item-Calibrated Interval Evidence for Temporal Session-Based Recommendation


Abstract

Temporal signals have been widely used in session-based recommendation to infer user interest. Existing temporal session-based recommenders primarily rely on absolute interval values, implicitly assuming that the same interval carries similar interest signals across items. However, we empirically find that this assumption does not hold: each item has its own interval distribution, so an interval should be interpreted relative to the item it belongs to. Based on this observation, we propose TRUST, a framework that evaluates each observed interval relative to the empirical interval distribution of the corresponding item. Specifically, we propose a score function to guide global neighbor sampling, session graph encoding, and final interest aggregation. Experiments on public datasets show that TRUST consistently improves over representative temporal and non-temporal baselines, and plug-in experiments further show that the proposed scoring function can improve existing temporal session recommenders as a model-agnostic method. Component-wise ablations further show that calibrating the temporal signals within each module, rather than removing the module itself, consistently improves neighbor sampling, session graph encoding, and interest aggregation.

1 Introduction↩︎

Temporal session-based recommendations (TSBRs) aim to predict a user’s next interaction by modeling both the order and timing of interactions within an ongoing session. Unlike conventional session-based recommendation (SBR), which mainly relies on item co-occurrence and transition order[1][3], TSBRs further use temporal signals as user-interest evidence for estimating the user’s current interest [4][7]. For example, if a user’s interaction with an item is consistent with common interval patterns, this may indicate stronger interest. In contrast, interactions that fall substantially below typical interval patterns, such as very short interval time, may suggest weaker or even negative interest [8][11].

Existing TSBR methods usually interpret such intervals according to a global interval distribution over all training interactions. Prior analyses show that intervals often cluster around a typical range and become sparse at very short or very long durations (as Overall label shown in 2) [7], [12]. This observation motivates TSBRs to weight session interactions by their temporal behavior: interactions with typical intervals are treated as stronger evidence of user interest, while unusually short or long ones indicate less interest [13]. However, users do not naturally spend the same amount of time on every item, even when they are interested [10], [14], [15]. As illustrated in 1 (a), a user spends 10 seconds on a laptop, a phone case, and a camera. Existing TSBRs treat them as the same evidence of user interest because the intervals are the same. When item-specific typical intervals (reference interval) are considered, the interpretation changes: the reference interval is 200 seconds for the laptop, 10 seconds for the phone case, and 150 seconds for the camera. Relative to these references, the phone-case interaction is temporally typical, whereas the laptop and camera interactions are abnormal. Thus, intervals should be interpreted relative to the item being viewed, rather than only by their absolute value. We refer to this item-relative consistency between an observed interval and the item’s usual interval distribution as temporal reliability.

a

Figure 1: No caption. a — Motivation for item-calibrated temporal interpretation. The same observed interval can carry different user-interest evidence when evaluated against item-specific temporal references.

This motivates our main objective: to make temporal intervals reliable evidence for temporal session-based recommendation by interpreting each interval relative to the item being viewed, rather than treating the same interval as equally meaningful across all items. Realizing this objective is not straightforward: (1) how to verify whether item-level interval heterogeneity makes globally interpreted intervals unreliable for TSBR; (2) how to estimate item-calibrated temporal signal without becoming overconfident for sparsely observed items; and (3) how to use this reliability signal in global relation modeling, local session transition modeling, and final interest generation.

We therefore propose a framework named Temporal Reliability Under item-Specific inTerval evidence for Session-based Recommendation (TRUST), which addresses these challenges with three corresponding designs. First, we provide empirical evidence through distributional and perturbation analyses in 3, showing that item-specific interval patterns differ from the overall interval distribution and that this mismatch affects performance of TSBRs. Second, we propose a score function (ITSF) to calibrate the heterogeneity of item-specific interval distribution: ITSF estimates an interval’s reliability score from the interval’s empirical quantile within the viewed item-specific interval distribution, and relaxes the penalty when that distribution is estimated from few observations. Third, we use the ITSF to guide the neighbor sampling, graph learning, and generation of session representation in TRUST. Specifically, in the global graph, TRUST uses transitions with reliable intervals to build item neighborhoods, reducing the influence of noisy co-occurrences on collaborative signals. Within each session, TRUST helps the graph encoder distinguish transitions based on how well their intervals align with the corresponding items’ interval distributions. Finally, TRUST uses interval reliability to decide whether to focus on recent interactions or look back to earlier ones. The overall framework is depicted in 3. We summarize main contributions as follows:

  • To the best of our knowledge, we are the first to identify and empirically characterize the impact of uncalibrated intervals on the performance of existing TSBR methods. We show that interval observations are not comparable across items, and that uncalibrated intervals can lead to a drop in performance.

  • We propose an item-specific interval calibration score function that measures how well an observed interval matches the historical interval distribution of the corresponding item. The score function is model-agnostic and can be integrated into existing TSBRs.

  • We develop TRUST, a temporal reliability-guided session recommendation framework that uses the item-calibrated temporal signal to improve both collaborative evidence modeling and session interest learning.

  • Experiments on public datasets show that TRUST improves over strong temporal and non-temporal baselines. Plug-in experiments further demonstrate the model-agnostic applicability of the proposed score function, and ablation studies verify the contribution of each temporal reliability-guided mechanism within the component.

2 Related Work↩︎

Early recommender systems mainly infer user interests from historical user–item interactions and collaborative patterns, where item co-occurrence and latent user–item relations are used as the primary evidence for recommendation [16], [17]. In session-based recommendation, the task becomes more challenging because the model usually observes only a short anonymous interaction sequence. Existing session-based methods have therefore developed recurrent [18], attention-based [19], [20], and graph-based architectures [21], [22] to model interests. These methods have substantially improved the representation of item dependencies within sessions, but most of them treat an interaction mainly as an item identity or an ordered transition. As a result, the temporal evidence associated with each interaction is ignored. This limitation has motivated increasing attention to temporal recommendation models, which incorporate time-related signals to better characterize user interest.

Building on this, temporal recommendations further enhance recommendation by incorporating temporal signals to better capture user’s interest [1], [23][25]. In these scenarios, timestamps [26][28], interaction intervals [29][31], and temporal ordering [32], have been used to model preference drift [33], [34], recency effects [35], and sequential dependency [36], [37]. For TSBRs, some studies embed temporal signals into item representations to capture time-aware session preferences [38], [39], while others incorporate temporal information into session graph structures to refine transition modeling [32], [40][42]. Interval-aware methods further model temporal gaps between consecutive interactions to improve item-transition representation [43], [44]. Moreover, multi-interest temporal methods use time information to distinguish heterogeneous short-term interests within the same session [45], [46]. These designs improve the modeling of users’ interests by exploiting temporal signals from different perspectives.

3 Statistical Analysis↩︎

In this section, we examine whether intervals should be treated as globally comparable across items. We first analyze the distributional mismatch between item-specific interval patterns and the overall interval distribution. We then further examine, under item-specific interval calibration, whether intervals located at different positions in their item-specific distributions contribute equally to recommendation performance.

3.1 Intervals Are Not Globally Comparable↩︎

We first examine the extent to which temporal observations vary across items. For each item \(i\), let \(\mathcal{D}_i\) denote the set of its observed intervals and define the log-transformed observation set as \(\mathcal{X}_i=\{\log(1+d):d\in\mathcal{D}_i\}\). We use \(\mathcal{X}_i\) to construct the empirical item-specific interval distribution \(\hat{F}_i(x)\) [47], [48]. For the overall empirical interval distribution \(\hat{F}_D(x)\), we aggregate all log-transformed intervals. We retain items with at least 20 temporal observations in the descriptive analysis. Our intuition is that if temporal signals were globally comparable across items, item-specific distributions \(\hat{F}_i(x)\) should not depart substantially from \(\hat{F}_D(x)\). We evaluate this assumption on three datasets.

Figure 2 provides evidence of item-specific interval distribution disparity. The red violin denotes the overall interval distribution over eligible items, while the blue violins denote ten representative empirical item-specific distributions. The representative items are selected to span the low-to-high range of item-specific median interval values. Across all three datasets, item-specific interval distributions visibly differ from the aggregate distribution and from one another. The differences involve central tendency, dispersion, and tail behavior. This descriptive pattern indicates that even a temporal value located around the median or high-density region of the overall interval distribution may fall into a lower or upper percentile range under an item-specific distribution. Hence, interval observations are not directly comparable across items without considering item-specific scales.

Figure 2: Item-specific interval distribution disparity across datasets. The red violin denotes the overall interval distribution, and the blue violins denote ten representative item-specific distributions sampled across item-specific medians. The visible differences in location, dispersion, and tail behavior show that interval observations are not globally comparable across items.

To assess whether the item-specific intervals distributional disparity is statistically supported, we conduct an aggregate-quantile chi-square homogeneity test [49]. Our hypothesis is: if items do not have a different interval distribution, then the intervals corresponding to an item should be distributed similarly to the overall interval distribution. For the formal test, we sample 200 items in each dataset and partition the aggregate distribution into \(K=10\) quantile bins \(\{B_1,\ldots,B_K\}\). The null hypothesis is therefore \(H_0:\hat{F}_i(x)=\hat{F}_D(x),\forall i\), while the alternative is \(H_1:\exists i,\hat{F}_i(x)\neq\hat{F}_D(x)\). For item \(i\), the observed count in bin \(k\) is \(O_{ik}=\sum_{x\in\mathcal{X}_i}\mathbf{1}(x\in B_k)\), and the expected count under \(H_0\) is \(E_{ik}=n_ip_k\), where \(n_i=|\mathcal{X}_i|\) and \(p_k\approx 1/K\). The Pearson statistic is computed as: \[\chi^2=\sum_i\sum_k \frac{(O_{ik}-E_{ik})^2}{E_{ik}}.\]

Table 1: Aggregate-quantile goodness-of-fit test between item-specific interval distributions and the overall interval distribution .
Dataset Values \(\chi^2\) Cramer’s \(V\) \(p\)-value
Diginetica 257,454 36,096.63 0.125 \(<0.01\)
RetailRocket 199,550 39,972.79 0.149 \(<0.01\)
Nowplaying 376,894 424,951.01 0.354 \(<0.01\)

As shown in Table 1, the aggregate-quantile tests reject the null hypothesis. Because chi-square tests can become significant in large samples, we interpret the Pearson statistics together with Cramer’s \(V\) [50]. The effect sizes indicate modest but non-trivial item-specific heterogeneity on Diginetica and RetailRocket, and a stronger discrepancy on Nowplaying. These results show that item-specific intervals cannot be treated as a single distribution.

3.2 Impact of Item-specific Interval Disparity on Temporal Session Recommendation↩︎

The distributional disparity analysis does not show the impact of interval distribution on recommendation performance directly. We therefore use a matched perturbation experiment to test whether item-specific interval typicality or atypicality corresponds to performance.

We measure item-specific interval typicality using a two-standard-deviation range around the mean of item-specific intervals. We construct two matched groups of interactions. The outside group contains interactions whose intervals fall outside the two-standard-deviation range of their corresponding item-specific interval distributions. The inside group contains interactions that fall within the range, but are matched to the outside group by their percentile under the overall interval distribution. Thus, we could disturb the two groups with the same interval intervention to partition the impact of intervention magnitude on performance. Specifically, we perturb original intervals by one standard deviation, computed on the overall interval distribution. The perturbation is applied in a randomly sampled positive or negative direction; when the negative direction would produce an invalid negative interval, we resample the direction. For comparison, we use several GNN-based and Transformer-based TSBRs baselines: IGT, TMI-GNN, DT-GAT, and TE-GNN. Appendix 7 gives brief descriptions, and Table 2 reports the results.

Table 2: Impact of perturbing temporally typical and atypical interactions on Diginetica.
Model Variant Hit@5 Hit@10 MRR@5 MRR@10
IGT Original 27.43 38.99 15.55 17.08
Inside 27.06 38.31 15.10 16.59
Inside change \(-\)0.37 \(-\)0.68 \(-\)0.45 \(-\)0.49
Outside 27.61 38.92 15.70 17.21
Outside change 0.18 \(-\)0.07 0.15 0.13
TMI-GNN Original 25.95 37.23 14.56 16.07
Inside 25.55 36.79 14.10 15.58
Inside change \(-\)0.40 \(-\)0.44 \(-\)0.46 \(-\)0.49
Outside 26.11 37.18 14.58 16.04
Outside change 0.16 \(-\)0.05 0.02 \(-\)0.03
DT-GAT Original 28.19 39.19 16.36 17.83
Inside 27.38 38.66 15.81 17.31
Inside change \(-\)0.81 \(-\)0.53 \(-\)0.55 \(-\)0.52
Outside 27.89 39.24 16.25 17.75
Outside change \(-\)0.30 0.05 \(-\)0.11 \(-\)0.08
TE-GNN Original 28.08 39.18 15.58 17.10
Inside 27.49 38.53 15.28 16.74
Inside change \(-\)0.59 \(-\)0.65 \(-\)0.30 \(-\)0.36
Outside 28.03 39.11 15.75 17.26
Outside change \(-\)0.05 \(-\)0.07 0.17 0.16

Across all four models, perturbing the inside group degrades performance on both HR and MRR metrics, with particularly larger drops on MRR@5. This indicates that the interval within the item-specific typical range provides more stable evidence for modeling short-term intent. Perturbing the outside group does not produce a stable performance decline; in some cases, it even leads to marginal improvements. This suggests that atypical intervals in the two tails are not consistently informative for TSBRs and may sometimes act as noisy interval signals. The experiment serves as controlled diagnostic evidence that item-specific interval typicality affects the stability of temporal signals during training.

Above analyzes give two takeaways. First, temporal values are not globally comparable across items: a value that appears typical under the dataset-level distribution may be atypical under the corresponding item-specific distribution. Second, item-specific interval typicality in its distribution is associated with higher recommendation performance, since perturbing temporally typical interactions causes performance drops compared with perturbing atypical interactions.

4 Methodology↩︎

The analysis in §3.1 suggests that item-specific interval heterogeneity can limit interval signals when a model treats them through absolute value. TRUST addresses this issue through three stages, as illustrated in 3: (i) calibrating each available item-interval observation into a temporal reliability score using the item’s empirical interval reference; (ii) using this shared reliability signal to guide item-specific representation learning from both reliability-aware global neighbor sampling and reliability-calibrated session graph encoding; and (iii) generating the final session representation through reliability-constrained plastic interest aggregation before next-item prediction.

Figure 3: Overview of the proposed TRUST framework. Item-calibrated temporal reliability is estimated from empirical item-specific interval distributions and used to sample reliable global neighbors, weight session-graph transitions, and capture reliable interests during session representation generation.

Problem Formulation

We begin by defining the notations used. Let \(\mathcal{I}= \{1, \ldots, |\mathcal{I}|\}\) be the items set. A prediction instance consists of an observed item sequence \(\mathcal{S}=[v_1,\ldots,v_n]\), where \(v_i \in \mathcal{I}\) is the item at position \(i\). For each observed transition from \(v_i\) to \(v_{i+1}\), let \(d_i\) denote the interval observation on item \(v_i\). Our goal is to predict the next item \(v_{n+1}\) given \(v_1, \ldots, v_n\) and corresponding intervals.

4.1 Shared Item-Calibrated Temporal Scoring Function↩︎

We first define the shared Item-Calibrated Temporal Scoring Function (ITSF) used throughout TRUST. Building on 3.1, our goal is to measure how much an observed interval \(d_i\) deviates from the interval distribution of other users on the same item \(v_i\). In our framework, such deviation is treated as evidence that the observed behavior is less consistent with the item-specific temporal pattern, and thus should receive a lower temporal reliability score. Moreover, this evidence should be weighted by the number of observations available for the item [51], [52]: when an item has only a few intervals observed, the item-level temporal distribution is uncertain, and the model should be conservative in penalizing deviations. Therefore, we define the Item-Calibrated Temporal Scoring Function, which computes the temporal reliability score \(q_i(d_i)\) as: \[\begin{align} q_i(d_i) &= \exp\left[ -\frac{1}{2} \left( \frac{|z_i(d_i)|}{1+\frac{1}{\sqrt{n_i+\epsilon_q}}} \right)^\gamma \right], \end{align} \label{eq:uncertainty-temporal-reliability}\tag{1}\] where \(\gamma>0\) controls how sharply reliability decays as the effective temporal deviation increases, \(\epsilon_q\) is a small constant for numerical stability, and \(n_i\) is the number of observed intervals for the corresponding item \(i\). A smaller \(n_i\) reflects greater uncertainty and therefore relaxes the penalty associated with the temporal deviation \(|z_i(d_i)|\). It measures the empirical quantile position of interval \(d_i\) in \(v_i\)’s item-specific intervals distribution, which \(z_i(d_i)\) is defined as: \[z_i(d_i) = \Phi^{-1} \left( \operatorname{clip} \left( \hat{F}_i(\log(1+d_i)), \rho, 1-\rho \right) \right),\] where \(\Phi^{-1}(\cdot)\) is the inverse standard normal cumulative distribution function, and \(\rho\) is a small clipping constant used to avoid infinite values at the distribution tails.

4.2 Reliability-Guided Structural-Temporal Encoding↩︎

Using the ITSF from §4.1, we next describe how TRUST incorporates this shared temporal evidence into item-level representation learning. As shown in Fig. 3, encoders consist of two reliability-guided message sources: a global source obtained through reliability-aware neighbor sampling (RaNS), and a local source obtained from a reliability-calibrated session graph (RcSG).

4.2.1 Reliability-aware Global Graph Encoding (RaNS)↩︎

The global co-occurrence graph provides collaborative evidence beyond an individual session, but its neighbors are formed from all observed consecutive interactions and may therefore include relations supported mainly by temporally unreliable behavior. Directly sampling from this graph can propagate such noisy collaborative signals into item representations. The global branch addresses this challenge by using RaNS to prioritize neighbors supported by reliable intervals and then applying graph convolution to encode the resulting collaborative evidence.

To implement RaNS, we first construct an undirected global item co-occurrence graph \(\mathcal{G}_g=(\mathcal{V},\mathcal{E}_g)\) only from training sessions, where an edge \(\{i,j\}\) indicates that items \(i\) and \(j\) appear in a pair of consecutive interactions in at least one session. All interactions from item \(i\) to item \(j\) comprise a set \(\mathcal{O}_{i\rightarrow j}\)

After constructing the global graph, we further sample item neighbors from it. Because the sampled quality directly shapes the information propagated during graph representation learning, , instead of randomly sampling neighboring items [53], [54], we choose the neighboring items that have the strongest reliable collaborative signal. We quantify this signal using the edge weight. To filter out noise connections, we only retain edges whose aggregated reliability weight exceeds the threshold \(p\), and sample \(w\) neighboring items to form the neighboring sampling set \(\tilde{\mathcal{N}}_g(i)\):

\[\begin{gather} \tilde{\mathcal{N}}_g(i) = \operatorname{Sample}_w \left( \{j \in \mathcal{N}_g(i) \mid \omega_{\{i,j\}}>p\} \right),\\ \omega_{\{i,j\}} = \sum_{o\in\mathcal{O}_{i\rightarrow j}} q_i(d_o) + \sum_{o\in\mathcal{O}_{j\rightarrow i}} q_j(d_o), \end{gather} \label{eq:retained-global-neighbor}\tag{2}\] where edge weights \(\omega_{\{i,j\}}\) aggregate both the frequency of item co-occurrences and the temporal reliability of each observed interval, and \(d_o\) denotes the interval in observation \(o\). \(\mathcal{N}_g(i)\) is the whole set of neighboring items in the graph. If \(\tilde{\mathcal{N}}_g(i)\) is empty, we retain the self-loop instead of sampling. Otherwise, \(\operatorname{Sample}(\cdot)\) is a probability-based sampling function with normalized weight \(P(j\mid i)=\omega_{\{i,j\}}/\sum_{j'\in\mathcal{C}_g(i)}\omega_{\{i,j'\}}\), which is defined as:

  • If \(|\mathcal{C}_g(i)| \geq w\), we sample \(w\) neighbors without replacement according to \(P(j\mid i)\).

  • If \(|\mathcal{C}_g(i)|< w\), all retained neighbors are first included, and the remaining \(w-|\mathcal{C}_g(i)|\) positions are filled by sampling with replacement according to \(P(j\mid i)\).

After obtaining \(\tilde{\mathcal{N}}_g(i)\), we apply a gated graph convolution to aggregate global collaborative evidence. Let \(\mathbf{h}_i\) denote the representation of item \(i\), initialized by its learnable item embedding. We finally compute global graph convolution as \(\mathbf{m}_{i}^{(g)}\): \[\begin{align} \mathbf{m}_{i}^{(g)} = \sum_{j\in \tilde{\mathcal{N}}_g(i)} a_{ij}^{(g)} \mathbf{W}_g\mathbf{h}_{j}, \\ a_{ij}^{(g)} = \sigma\left( \mathbf{w}_g^\top [\mathbf{h}_i;\mathbf{h}_j] + b_g \right), \end{align} \label{eq:global-message}\tag{3}\] where \(\mathbf{W}_g\) is a learnable transformation matrix, \(\mathbf{w}_g\) and \(b_g\) are learnable parameters.

4.2.2 Reliability-calibrated Session Graph (RcSG)↩︎

As discussed in 3, although TSBR graph encoders utilize absolute intervals, they cannot account for distributional differences in intervals across items. RcSG addresses this by incorporating ITSF scores into directed session-graph message passing, so that local transitions are encoded according to their item-calibrated temporal reliability.

To do so, it constructs a directed session graph and uses the ITSF score as convolutional edge weights. We fuse temporal reliability with session graph learning through a temporal reliability-aware attention mechanism. Each item \(i\)’s representation \(\mathbf{m}_{i}^{(s)}\) is then computed as: \[\label{eq:session-local-attn} \begin{gather} \mathbf{m}_{i}^{(s)} = \sum_{j\in\mathcal{N}_s(i)} \beta_{ij}^{(s)} \mathbf{h}_j, \\ \beta_{ij}^{(s)} = \operatorname{softmax}_{j\in\mathcal{N}_s(i)} \left(e_{ij}^{(s)}\right), \\ e_{ij}^{(s)} = \operatorname{LeakyReLU} \Bigl( \mathbf{W}_1^{s} \bigl( \mathbf{h}_i \odot \mathbf{h}_j \bigr) + \mathbf{W}_2^{s} \mathcal{T}(q_i(d_i)) \Bigr). \end{gather}\tag{4}\] where, \(\mathcal{N}_s(i)\) denotes the outgoing neighboring items of \(i\) in the directed session graph, and \(\odot\) denotes element-wise multiplication. \(\mathcal{T}(\cdot)\) is a linear projection that maps the scalar reliability score into a \(d\)-dimensional temporal feature, while \(\mathbf{W}_1^{s}\) and \(\mathbf{W}_2^{s}\) project their inputs to scalar edge scores before the LeakyReLU activation.

4.3 Plastic Interest Aggregation (PIA)↩︎

A remaining challenge lies in how to effectively fuse the item representations learned by the two encoders into a unified session representation. This is non-trivial because existing methods commonly treat the last item as the short-term interest signal for matching other items to generate a session representation, although the last item alone may be noisy and insufficient. To address it, we propose Plastic Interest Aggregation, which models reliable short-term interests at multiple depths.

We first fuse two encoders’ outputs with temporal position information to get the fused item \(i\)’s representation \(\bar{\mathbf{h}}_{i}\): \[\begin{align} \bar{\mathbf{h}}_{i} &= \mathbf{m}_{i}^{(g)} + \mathbf{m}_{i}^{(s)} + \mathbf{p}_{\ell_i}^{(\mathrm{rev})}, \quad \ell_i = n-i+1, \end{align} \label{eq:reverse-position-fusion}\tag{5}\] where \(\ell_i\) is the reverse position index, and \(\mathbf{p}_{\ell_i}^{(\mathrm{rev})}\in\mathbb{R}^{d}\) is a learnable reverse positional embedding.

PIA then builds reliable interest queries at multiple depths. We assign a position-based reliability score to each position. To prevent data leakage, we assign the last position score as 1, the position-based reliability score \(\tilde{r}_i\) is: \[\tilde{r}_i = \begin{cases} q_{i}(d_i), & 1\le i<n,\\ 1, & i=n, \end{cases} \label{eq:reliability-coordinate-mass}\tag{6}\] where \(d_i\) is the observed interval on the item at position \(i\). For each depth-\(b\) view, PIA allocates a depth of \(b \in \mathcal{B}\) backward from the last item and assigns the position weight \(\xi_i^{(b)}\) as follows: \[\xi_i^{(b)} = \max\!\left( 0, \min\!\left( \tilde{r}_i, b-\sum_{k=i+1}^{n}\xi_k^{(b)} \right) \right), \label{eq:budget-allocation}\tag{7}\] where the weights are computed recursively for \(i=n,n-1,\ldots,1\). Small depths focus on the most recent reliable items; larger depths extend toward earlier items. After normalizing \(\xi_i^{(b)}\) as \(\bar{\xi}_i^{(b)}=\xi_i^{(b)}/(\sum_{j=1}^{n}\xi_j^{(b)}+\epsilon_{\mathrm{norm}})\), where \(\epsilon_{\mathrm{norm}}\) is a small constant for numerical stability, the depth-specific plastic query \(\mathbf{q}_p^{(b)}\) is: \[\mathbf{q}_p^{(b)} = \sum_{i=1}^{n} \bar{\xi}_i^{(b)} \bar{\mathbf{h}}_{i}. \label{eq:plastic-query}\tag{8}\]

For each query \(\mathbf{q}_p^{(b)}\) with depth \(b\), we apply additive attention over all items in the session to obtain the various view of session representation \(\mathbf{s}^{(b)}\): \[\begin{gather} \mathbf{s}^{(b)} = \sum_{i=1}^{n} \beta_i^{(b)} \bar{\mathbf{h}}_{i}, \\ \beta_i^{(b)} = \frac{ \exp(a_i^{(b)}) }{ \sum_{j=1}^{n} \exp(a_j^{(b)}) }, \\ a_i^{(b)} = \mathbf{w}_a^\top \sigma \!\left( \mathbf{W}_a \bar{\mathbf{h}}_{i} + \mathbf{W}_q \mathbf{q}_p^{(b)} + \mathbf{b}_a \right), \end{gather} \label{eq:reliability-constrained-attn}\tag{9}\] where \(\sigma\) is the Sigmoid activation, \(\mathbf{W}_a,\mathbf{W}_q\in\mathbb{R}^{d\times d}\) and \(\mathbf{w}_a\in\mathbb{R}^{d}\) are learnable parameters. The final session representation \(\mathbf{s}\) is obtained by averaging the available views: \[\mathbf{s} = \frac{1}{|\mathcal{B}|} \sum_{b\in\mathcal{B}} \mathbf{s}^{(b)}. \label{eq:session-rep}\tag{10}\] To this end, we complete the modelling of our proposed framework. The complete procedure is summarized in 4.

Training Objective:

The training objective is to minimize the cross-entropy loss \(\mathcal{L}\) over the full item softmax: \[\begin{align} \mathcal{L} &= -\log \frac{\exp(\hat{y}_{v^*})}{\sum_{j \in \mathcal{I}} \exp(\hat{y}_j)}, \\ \hat{y}_j &= \mathbf{s}^\top \mathbf{h}_j, \quad \forall\, j \in \mathcal{I}. \end{align} \label{eq:loss}\tag{11}\] where \(v^*=v_{n+1}\) is the ground-truth next item, \(\mathbf{h}_j \in \mathbb{R}^d\) is the shared embedding of item \(j\), and \(\mathcal{I}\) is the set of all items.

Figure 4: TRUST: Temporal Reliability Calibration Framework

5 Experiments↩︎

5.1 Experimental Setup↩︎

5.1.1 Datasets↩︎

We conduct experiments on three widely used public datasets: Diginetica, RetailRocket, and Nowplaying. Diginetica is derived from the CIKM Cup 2016 competition [55] and contains user click sequences on product pages, each annotated with the time spent. RetailRocket covers product-view, add-to-cart, and purchase events; following standard protocol we retain only view events [56]. Nowplaying is a music-domain dataset constructed from listening list, where each session records successive song interactions [57]. For all three datasets, we follow the standard preprocessing procedure [25], [58], [59]: sessions of length one are discarded, items appearing fewer than five times are removed, and the last item in each session is held out as the prediction target. Table 3 summarizes the statistics of datasets. The column Mean/Med. gives the ratio of the mean to the median of intervals.

Table 3: Dataset statistics of the three datasets
Dataset #Items Train sess. Test sess. Avg.len Mean/Med.
Diginetica 43,097 719,470 60,858 5.12 1.71
RetailRocket 50,018 721,266 61,969 4.35 2.56
Nowplaying 60,416 825,304 89,824 5.55 1.45

5.1.2 Evaluation Metrics↩︎

We adopt Hit Rate (HR@\(K\)) and Mean Reciprocal Rank (MRR@\(K\)) at \(K\in\{5,10\}\). HR@\(K\) measures whether the ground-truth next item appears in the top-\(K\) predicted list; MRR@\(K\) suggests whether a high rank for the ground-truth item. All metrics are reported as percentages.

5.1.3 Baselines↩︎

We compare TRUST against two groups of baselines. The non-temporal SBR group includes GRU4Rec [18], NARM [19], SR-GNN [21], FGNN [60], GCE-GNN [61], \(S^2\)-DHCN [62], MSGAT [63], and DMI-GNN [64]. The TSBRs group includes STAN [65], TASRec [38], TE-GNN [40], IGT [43], TMI-GNN [46], and DT-GAT [7]. The baselines description lies in Appendix 7.

Table 4: Overall recommendation performance (%). Best result per column is bold; second-best is underlined. ‘* indicates the difference between TRUST and the second-best model is statistically significant (\(p\)-value \(<\) 0.01 by a paired \(t\)-test based on 5 runs of the best performing model and the second-best performing model, respectively)
Diginetica RetailRocket Nowplaying
2-5(lr)6-9(lr)10-13 Method HR@5 HR@10 MRR@5 MRR@10 HR@5 HR@10 MRR@5 MRR@10 HR@5 HR@10 MRR@5 MRR@10
GRU4Rec 11.93 18.53 6.46 7.49 43.16 47.71 33.94 35.20 7.42 11.18 4.30 4.88
NARM 25.95 35.55 14.07 15.38 47.84 55.47 34.20 35.08 8.15 12.34 4.86 5.41
SR-GNN 27.15 38.65 15.29 16.82 48.66 56.04 36.25 37.24 9.31 13.67 5.52 6.10
FGNN 24.39 35.90 13.22 14.64 47.35 55.29 35.24 35.36 8.63 12.89 5.07 5.52
\(S^2\)-DHCN 28.35 39.87 15.68 17.53 48.94 56.94 36.15 36.92 9.74 14.05 5.76 6.35
MSGAT 24.74 35.58 13.99 15.47 47.92 56.00 35.72 36.71 8.92 13.21 5.31 5.87
DMI-GNN 28.08 40.16 16.27 17.91 51.26 58.05 38.20 39.17 12.34 17.09 7.67 8.01
STAN 25.71 38.70 15.64 17.12 49.40 56.28 37.75 38.68 9.03 13.04 5.34 5.88
TASRec 27.39 39.10 15.31 17.35 51.03 57.86 38.16 39.13 10.45 15.22 6.18 6.79
TE-GNN 28.08 39.18 15.58 17.10 47.40 54.85 34.98 35.98 11.10 15.91 6.60 7.23
IGT 27.43 38.99 15.55 17.08 48.42 55.71 36.20 36.96 10.82 15.47 6.43 7.03
TMI-GNN 25.95 37.23 14.56 16.07 50.58 57.87 37.45 38.44 9.77 14.32 5.80 6.41
DT-GAT 28.19 39.19 16.36 17.83 49.70 56.24 36.48 37.36 9.70 14.77 6.03 6.77
TRUST* 30.07* 41.79* 17.27* 18.82* 52.30* 59.87* 39.10* 40.22* 12.68* 17.35 7.79* 8.40*
Improve \(+6.07\%\) \(+4.06\%\) \(+5.56\%\) \(+5.08\%\) \(+2.03\%\) \(+3.14\%\) \(+2.36\%\) \(+2.68\%\) \(+2.76\%\) \(+1.52\%\) \(+1.56\%\) \(+4.87\%\)

5.1.4 Implementation Details↩︎

All parameters are initialized with Xavier-uniform initialization and optimized with Adam at an initial learning rate of \(10^{-3}\). Both item-embedding dimension and batch size is set as 100 to align with baselines. We apply early stopping based on HR@10, with a patience of 3 epochs. The ITSF is first estimated from the training dataset before training. The clipping constant \(\rho\) is set at 0.01. The multi-granularity plastic interest in PIA is set as \(\mathcal{B}=\{1,2,3\}\).

5.2 Research Questions↩︎

The experiments evaluate four aspects: overall recommendation accuracy, plug-in (ITSF) applicability to existing TSBRs, the contribution of each reliability-guided use of the signal (ablation), and computational practicality. Five questions are as follows: RQ1 evaluates whether TRUST improves next-item recommendation over non-temporal and temporal baselines. RQ2 tests whether the calibrated reliability signal is useful beyond TRUST as a plug-in replacement for raw interval. RQ3 isolates whether the three uses of the reliability signal, global sampling, session-graph transition weighting, and interest aggregation, each contribute to performance. RQ4 analyzes the sensitivity of TRUST to the main hyperparameters. RQ5 examines algorithm efficiency.

5.3 RQ1: Overall Performance Comparison↩︎

Table 4 reports the recommendation performance on Diginetica, RetailRocket, and Nowplaying. TRUST achieves the best result on almost all metrics and shows statistically significant gains over the strongest competing baselines in most cases. The magnitude of improvement varies across datasets, while gains remain consistent.

TBSRs do not always outperform strong non-temporal SBRs. For example, several temporal methods perform below DMI-GNN on Diginetica and Nowplaying. On the one hand, DMI-GNN is a recent state-of-the-art model that utilizes a dynamic multi-interest mechanism. Its performance also suggests that raw temporal intervals alone do not guarantee better recommendation accuracy. TRUST addresses this issue by calibrating temporal information at the item level before using it as evidence of user interest.

5.4 RQ2: Plug-In Applicability with ITSF↩︎

We also test whether the item-calibrated temporal signal helps models beyond the full TRUST architecture. Specifically, we integrate ITSF into four temporal baselines: IGT, TE-GNN, TMI-GNN, and DT-GAT. We utilize ITSF to replace the original interval used in these methods, while the session representation generation, training objective, optimizer, and other implementation details remain unchanged. We evaluate the resulting variants under the same training protocol as §5.1.4.

Table 5 compares the original temporal baselines with their ITSF-based variants. Replacing the original interval with ITSF improves HR@10 and MRR@10 for all four baselines. The gains vary across backbones and datasets, but the results show that ITSF could serve as a model-agnostic module for TSBRs.

Table 5: Performance of temporal SBR backbones with and without ITSF. “Cal.” denotes the variant integrated with ITSF.
Diginetica Nowplaying
2-3(lr)4-5 Method HR@10 MRR@10 HR@10 MRR@10
IGT 38.99 17.08 15.47 7.03
Cal. 39.24 17.38 15.94 7.47
Gain +0.64% +1.76% +3.04% +6.26%
TE-GNN 39.18 17.10 15.91 7.23
Cal. 39.53 17.46 16.37 7.28
Gain +0.89% +2.11% +2.89% +0.69%
TMI-GNN 37.23 16.07 14.32 6.41
Cal. 37.84 16.45 14.97 6.62
Gain +1.64% +2.36% +4.54% +3.28%
DT-GAT 39.19 17.83 14.77 6.77
Cal. 40.13 18.34 15.16 7.18
Gain +2.40% +2.86% +2.64% +6.06%

4pt

5.5 RQ3: Ablation Study↩︎

In this experiment, we verify the contribution of each TRUST component, especially the item-calibrated temporal signal in each module, by evaluating the following variants:

  • w/o RaNS_cal. Replace reliability-aware neighbor sampling with uniform co-occurrence-frequency sampling, removing the reliability-aware neighbor sampling in Eq. 2 .

  • w/o RcSG_cal. Remove the temporal reliability from the session-graph learning in Eq. 4 , reducing it to standard graph attention.

  • w UniPIA. Keep the PIA aggregation structure but set all reliability weights in PIA to one, so the backward multi-granularity plastic interest aggregation remains while reliability-weighted allocation is removed.

  • w LastQ. Replace PIA with a standard short-term interest-based session generation, which uses the last item as the short-term query.

  • w/o ITSF. Replace the item-calibrated temporal scoring function with the original absolute intervals, so that scoring is no longer estimated relative to each item’s empirical interval distribution.

Table 6 reports the ablation results. Each variant performs worse than the full model, but the magnitude of the drop varies by module. Performance decreases when RaNS is removed, suggesting that reliability-aware neighbor sampling helps prevent noisy global co-occurrence relations in global graph learning. Removing RcSG also reduces performance, especially on Nowplaying, where local transition reliability appears more important.

The two PIA-based variants give a direct view of the session generation step. UniPIA remains competitive because it maintains the multi-granularity plastic interest aggregation structure, but it falls slightly below full TRUST when reliability-weighted allocation is removed. LastQ is also close to the full model on Diginetica, yet it is weaker on both datasets, especially on Nowplaying MRR@10. These results suggest that the PIA structure matters, and the calibrated reliability weights still provide useful information when deciding how much recent interactions should contribute.

Removing ITSF results in a larger decline than removing either RaNS or RcSG alone, because ITSF provides the calibrated reliability signal used by the downstream reliability-guided modules. Once ITSF is replaced by absolute intervals, the model can no longer distinguish whether an observed interval is typical for the specific item under consideration.

Table 6: Ablation results on Diginetica and Nowplaying (%). Best result per column is bold.
Diginetica Nowplaying
2-3(lr)4-5 Variant HR@10 MRR@10 HR@10 MRR@10
w/o RaNS_cal 41.36 18.33 17.27 8.16
w/o RcSG_cal 41.44 18.23 16.99 7.97
w UniPIA 41.29 18.47 16.97 8.14
w LastQ 41.40 18.50 17.18 8.09
w/o ITSF 39.87 17.54 16.42 7.81
Full TRUST 41.79 18.82 17.35 8.40

4pt

5.6 RQ4: Hyperparameter Sensitivity↩︎

We examine two hyperparameters: the neighbor sampling size \(w\) on global graph and the temporal reliability decay exponent \(\gamma\) on ITSF. Figure 5 and Figure 6 report the results.

5.6.1 Effect of neighbor sampling size \(w\).↩︎

The neighbor sampling size \(w\) controls the number of neighboring items used in RaNS for graph learning. As shown in Figure 5, the effect of \(w\) is dataset-dependent but generally moderate. On Diginetica, HR@10 increases slightly as \(w\) grows, reaching its maximum at \(w=8\), whereas MRR@10 remains stable and is marginally higher for smaller neighbor sampling sizes. On Nowplaying, HR@10 also benefits from a larger sampling size and peaks at \(w=8\), while MRR@10 gradually decreases as \(w\) increases, suggesting that overly broad sampling may weaken top-rank precision in the music-domain setting. RetailRocket shows a different pattern: both HR@10 and MRR@10 improve substantially as \(w\) increases, with the best ranking quality around \(w=8\) and the highest HR@10 at \(w=12\). In the main experiments, we choose \(w\) based on validation performance, resulting in \(w=1\) for Diginetica, \(w=12\) for RetailRocket, and \(w=2\) for Nowplaying.

5.6.2 Effect of decay exponent \(\gamma\).↩︎

The decay exponent \(\gamma\) in Eq. 1 controls how sharply the reliability score decreases when an observed interval deviates from the item-specific intervals reference. Figure 6 shows that TRUST performs consistently across the tested range \(\gamma\in\{1.0,1.5,2.0,3.0\}\). On all three datasets, MRR@10 improves from \(\gamma=1.0\) to \(\gamma=2.0\), indicating that a Gaussian-like decay provides a useful balance between robustness and discrimination. When \(\gamma\) is further increased to \(3.0\), performance becomes slightly lower or nearly unchanged, suggesting that an overly sharp decay may over-penalize moderately atypical but still informative intervals. HR@10 follows a similar but less monotonic pattern, with the best values occurring around \(\gamma=2.0\). Therefore, we choose \(\gamma=2.0\) for three datasets.

a

b

c

Figure 5: Hyperparameter sensitivity of TRUST with respect to neighbor sampling size \(w\). HR@10 and MRR@10 are reported on Diginetica, RetailRocket, and Nowplaying..

a

b

c

Figure 6: Hyperparameter sensitivity of TRUST with respect to the temporal reliability decay exponent \(\gamma\). HR@10 and MRR@10 are reported on Diginetica, RetailRocket, and Nowplaying..

5.7 RQ5: Algorithm Efficiency↩︎

Last but not least, we ask whether our proposed method improves performance at the cost of high computational overhead. We compare TRUST with representative baselines on actual per-epoch training time and the number of convergence epochs (Table 7). TRUST adds two costs relative to SBRs: (i) one-time offline pre-computation of ITSF, and (ii) Multi-granularity of PIA. The precomputation step takes 37 seconds and runs once before training. During training, the ITSF reduces to a vectorized binary search over pre-sorted arrays, so ITSF does not require iterative parameter learning.

Table 7 further shows that TRUST takes longer per epoch than most temporal baselines and converges after more epochs on Diginetica. The full architecture has a moderate overhead in exchange for improvement of performance. ITSF pre-computation is small compared with model training, but PIA makes the model heavier than the last item-based session generation TSBRs.

Table 7: Actual running time on Diginetica.
Method Time/epoch (s) Conv.epochs
DMI-GNN 221 12
IGT 323 9
TE-GNN 182 10
TMI-GNN 209 10
DT-GAT 488 8
TRUST 439 14
+ ITSF precomp. 37

5pt

6 Conclusion↩︎

This paper presents TRUST, a temporal reliability calibration framework for session-based recommendation. TRUST calibrates the temporal interval as item-specific evidence whose meaning depends on the empirical interval pattern of the corresponding item. Experiments across three datasets show that this calibration improves recommendation performance and can also benefit several existing temporal SBR backbones. Future work can extend this line of work with tail-aware and shift-aware calibration strategies. The sample-size-dependent margin in ITSF improves the robustness of item-level temporal calibration for sparsely observed items. Future work can build on this design by incorporating one-sided tail calibration and category-level temporal references to further enhance temporal reliability estimation in long-tail recommendation scenarios.

7 Baseline Descriptions↩︎

GRU4Rec [18] introduces recurrent neural networks into session-based recommendation using GRU units to model intra-session sequential patterns for next-item prediction.

NARM [19] extends recurrent session modeling with an attention mechanism to jointly represent the user’s main intent and more localized sequential preference within a session.

SR-GNN [21] formulates each session as a directed graph and applies gated graph neural networks to learn item representations, followed by an attention-based readout.

FGNN [60] enhances graph-based session modeling with richer interaction structures to capture item dependencies from both sequential transitions and graph relations.

GCE-GNN [61] augments session-level graph modeling with global item co-occurrence information, allowing local session representations to benefit from broader collaborative context.

\(S^2\)-DHCN [62] models high-order item and session relationships through hypergraph convolution and introduces self-supervised learning to improve representation quality.

MSGAT [63] applies sparse graph attention to reduce noise in item-level and session-level representations for more robust graph-based session modeling.

DMI-GNN [64]: introduces multi-interest learning into session modeling, uses multiple positional patterns to encode different positional contexts, and applies dynamic multi-interest regularization to reduce redundant interest representations based on session length.

STAN [65] adapts neighborhood-based session recommendation with temporal decay, assigning lower importance to distant sessions and higher importance to recent interactions.

TASRec [38] incorporates historical sessions across different days using a temporally decaying mechanism to assign adaptive importance to past user-interest evidence.

TE-GNN [40] incorporates time-enhanced item transition information into graph neural networks, improving session representation by modeling temporally dependent interaction patterns.

IGT [43] explicitly exploits inter-item intervals within a session to refine item relation modeling and improve session representation construction.

TMI-GNN [46] introduces temporal-aware multi-interest modeling to capture evolving user preferences from time-sensitive item transitions within a session.

DT-GAT [7] uses time-aware graph and hyper-graph channels to incorporate interaction intervals and inter-session temporal differences into item-level and session-level representation learning.

References↩︎

[1]
Z. Li et al., “Graph and sequential neural networks in session-based recommendation: A survey,” ACM Computing Surveys, vol. 57, no. 2, pp. 1–37, 2024.
[2]
M. Choi, S. Lee, S. Park, and J. Lee, “Linear item-item models with neural knowledge for session-based recommendation,” in Proceedings of the 48th international ACM SIGIR conference on research and development in information retrieval, 2025, pp. 1666–1675.
[3]
X. Zhang, B. Xu, F. Ma, Z. Wang, L. Yang, and H. Lin, “Rethinking contrastive learning in session-based recommendation,” Pattern Recognition, vol. 169, p. 111924, 2026.
[4]
X. Yi, L. Hong, E. Zhong, N. N. Liu, and S. Rajan, “Beyond clicks: Dwell time for personalization,” in Proceedings of the 8th ACM conference on recommender systems, 2014, pp. 113–120.
[5]
V. Bogina and T. Kuflik, “Incorporating dwell time in session-based recommendations with recurrent neural networks.” RecTemp@ RecSys, vol. 1922, pp. 57–59, 2017.
[6]
Z. Xu et al., “Time interval aware graph neural networks for session-based recommendation,” in Pacific-asia conference on knowledge discovery and data mining, 2025, pp. 53–65.
[7]
L. Guo, S. Wu, D. Lu, L. Gao, and G. Xu, “Dual-channel time-aware graph attention network for session-based recommendation,” Information Sciences, p. 123289, 2026.
[8]
P. Yin, P. Luo, W.-C. Lee, and M. Wang, “Silence is also evidence: Interpreting dwell time for recommendation from psychological perspective,” in Proceedings of the 19th ACM SIGKDD international conference on knowledge discovery and data mining, 2013, pp. 989–997.
[9]
S. Gong and K. Q. Zhu, “Positive, negative and neutral: Modeling implicit feedback in session-based news recommendation,” in Proceedings of the 45th international ACM SIGIR conference on research and development in information retrieval, 2022, pp. 1185–1195.
[10]
R. Xie, L. Ma, S. Zhang, F. Xia, and L. Lin, “Reweighting clicks with dwell time in recommendation,” in Companion proceedings of the ACM web conference 2023, 2023, pp. 341–345.
[11]
Z. He et al., “A survey on user behavior modeling in recommender systems,” in Proceedings of the thirty-second international joint conference on artificial intelligence, 2023, pp. 6656–6664.
[12]
C. Wu et al., “Feedrec: News feed recommendation with various user feedbacks,” in Proceedings of the ACM web conference 2022, 2022, pp. 2088–2097.
[13]
Y. Kim, A. Hassan, R. W. White, and I. Zitouni, “Modeling dwell time to predict click-level satisfaction,” in Proceedings of the 7th ACM international conference on web search and data mining, 2014, pp. 193–202.
[14]
C. Liu, R. W. White, and S. Dumais, “Understanding web browsing behaviors through weibull analysis of dwell time,” in Proceedings of the 33rd international ACM SIGIR conference on research and development in information retrieval, 2010, pp. 379–386.
[15]
Y. Seki and M. Yoshida, “Analysis of user dwell time by category in news application,” in 2018 IEEE/WIC/ACM international conference on web intelligence (WI), 2018, pp. 732–735.
[16]
Y. Koren, R. Bell, and C. Volinsky, “Matrix factorization techniques for recommender systems,” Computer, vol. 42, no. 8, pp. 30–37, 2009.
[17]
S. Rendle, C. Freudenthaler, Z. Gantner, and L. Schmidt-Thieme, “BPR: Bayesian personalized ranking from implicit feedback,” in Proceedings of the twenty-fifth conference on uncertainty in artificial intelligence, 2009, pp. 452–461.
[18]
B. Hidasi, “Session-based recommendations with recurrent neural networks,” in The international conference on learning representations (ICLR), 2016.
[19]
J. Li, P. Ren, Z. Chen, Z. Ren, T. Lian, and J. Ma, “Neural attentive session-based recommendation,” in Proceedings of the 2017 ACM on conference on information and knowledge management, 2017, pp. 1419–1428.
[20]
R. Wang, X. Rui, and Z. Wang, “Category-aware dual channel graph neural networks for session-based recommendation: R. Wang et al.” Knowledge and Information Systems, vol. 68, no. 1, p. 3, 2026.
[21]
S. Wu, Y. Tang, Y. Zhu, L. Wang, X. Xie, and T. Tan, “Session-based recommendation with graph neural networks,” in Proceedings of the AAAI conference on artificial intelligence, 2019, vol. 33, pp. 346–353.
[22]
F. Yu, Y. Zhu, Q. Liu, S. Wu, L. Wang, and T. Tan, “TAGNN: Target attentive graph neural networks for session-based recommendation,” in Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval, 2020, pp. 1921–1924.
[23]
Y. Koren, “Collaborative filtering with temporal dynamics,” in Proceedings of the 15th ACM SIGKDD international conference on knowledge discovery and data mining, 2009, pp. 447–456.
[24]
W. Ye, S. Wang, X. Chen, X. Wang, Z. Qin, and D. Yin, “Time matters: Sequential recommendation with complex temporal information,” in Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval, 2020, pp. 1459–1468.
[25]
X. Zhang et al., “A survey on side information-driven session-based recommendation: From a data-centric perspective,” IEEE Transactions on Knowledge and Data Engineering, 2025.
[26]
Z. Fan, Z. Liu, J. Zhang, Y. Xiong, L. Zheng, and P. S. Yu, “Continuous-time sequential recommendation with temporal graph collaborative transformer,” in Proceedings of the 30th ACM international conference on information & knowledge management, 2021, pp. 433–442.
[27]
V. A. Tran, G. Salha-Galvan, B. Sguerra, and R. Hennequin, “Attention mixtures for time-aware sequential recommendation,” in Proceedings of the 46th international ACM SIGIR conference on research and development in information retrieval, 2023, pp. 1821–1826.
[28]
S. Lee, S. Park, J. Kim, M. Yoon, and J. Lee, “Enhancing time awareness in generative recommendation,” in Findings of the association for computational linguistics: EMNLP 2025, 2025, pp. 23917–23933.
[29]
Y. Ma, B. Narayanaswamy, H. Lin, and H. Ding, “Temporal-contextual recommendation in real-time,” in Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining, 2020, pp. 2291–2299.
[30]
J. Li, Y. Wang, and J. McAuley, “Time interval aware self-attention for sequential recommendation,” in Proceedings of the 13th international conference on web search and data mining, 2020, pp. 322–330.
[31]
J. Luo, W. Zhang, X. Zhang, and Y. Fang, “Time-aware adaptive side information fusion for sequential recommendation,” in Proceedings of the nineteenth ACM international conference on web search and data mining, 2026, pp. 469–478.
[32]
Z. Wan et al., “Spatio-temporal contrastive learning-enhanced GNNs for session-based recommendation,” ACM Transactions on Information Systems, vol. 42, no. 2, pp. 1–26, 2023.
[33]
X. Li, Y. Liu, Z. Liu, and P. S. Yu, “Time-aware hyperbolic graph attention network for session-based recommendation,” in 2022 IEEE international conference on big data (big data), 2022, pp. 626–635.
[34]
L. Heryawan, R. Pulungan, et al., “Trust decay-based temporal learning for dynamic recommender systems with concept drift adaptation,” IEEE Access, 2025.
[35]
M. Filipovic, B. Mitrevski, D. M. Antognini, E. Lejal Glaude, B. Faltings, and C.-C. Musat, “Modeling online behavior in recommender systems: The importance of temporal context,” in Perspectives 2021-proceedings of the perspectives on the evaluation of recommender systems workshop 2021, co-located with the 15th ACM conference on recommender systems, RecSys2021, 2021.
[36]
R. Zhang, Y. Gu, X. Shen, and H. Su, “Knowledge-enhanced session-based recommendation with temporal transformer,” arXiv preprint arXiv:2112.08745, 2021.
[37]
L. Xia, C. Huang, Y. Xu, and J. Pei, “Multi-behavior sequential recommendation with temporal graph transformer,” IEEE Transactions on Knowledge and Data Engineering, vol. 35, no. 6, pp. 6099–6112, 2022.
[38]
H. Zhou, Q. Tan, X. Huang, K. Zhou, and X. Wang, “Temporal augmented graph neural networks for session-based recommendations,” in Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval, 2021, pp. 1798–1802.
[39]
C. Guo, R. Xue, and J. Zhang, “Time-period-aware embedding regeneration for session-based recommendation,” in Proceedings of the 34th ACM international conference on information and knowledge management, 2025, pp. 4738–4742.
[40]
G. Tang, X. Zhu, J. Guo, and S. Dietze, “Time enhanced graph neural networks for session-based recommendation,” Knowledge-Based Systems, vol. 251, p. 109204, 2022.
[41]
Y. Li et al., “Spatiotemporal-aware session-based recommendation with graph neural networks,” in Proceedings of the 31st acm international conference on information & knowledge management, 2022, pp. 1209–1218.
[42]
Q. Chen, F. Jiang, X. Guo, J. Chen, K. Sha, and Y. Wang, “Combine temporal information in session-based recommendation with graph neural networks,” Expert Systems with Applications, vol. 238, p. 121969, 2024.
[43]
H. Wang, Y. Zeng, J. Chen, N. Han, and H. Chen, “Interval-enhanced graph transformer solution for session-based recommendation,” Expert Systems with Applications, vol. 213, p. 118970, 2023.
[44]
Z. Zuo, J. Lu, L. Tan, Q. Hu, B. Wang, and F. Liu, “Time highlighted multi-interest network for session-based news recommendation,” Intelligent Data Analysis, p. 1088467X261426714, 2026.
[45]
Z. Wang and Y. Shen, “Time-aware multi-interest capsule network for sequential recommendation,” in Proceedings of the 2022 SIAM international conference on data mining (SDM), 2022, pp. 558–566.
[46]
Q. Shen, S. Zhu, Y. Pang, Y. Zhang, and Z. Wei, “Temporal aware multi-interest graph neural network for session-based recommendation,” in Asian conference on machine learning, 2023.
[47]
S. Xu, Y. Zhu, H. Jiang, and F. C. Lau, “A user-oriented webpage ranking algorithm based on user attention time.” in AAAI, 2008, vol. 8, pp. 1255–1260.
[48]
S. Xu, H. Jiang, and F. C.-M. Lau, “Mining user dwell time for personalized web search re-ranking,” in International joint conference on artificial intelligence (IJCAI 2011), 2011.
[49]
A. Agresti, Categorical data analysis. John Wiley & Sons, 2013.
[50]
R. L. Wasserstein and N. A. Lazar, “The ASA statement on p-values: Context, process, and purpose,” The American Statistician, vol. 70. Taylor & Francis, pp. 129–133, 2016.
[51]
V. Kuleshov, N. Fenner, and S. Ermon, “Accurate uncertainties for deep learning using calibrated regression,” in International conference on machine learning, 2018, pp. 2796–2804.
[52]
C. Han, P. Castells, P. Gupta, X. Xu, and V. Salaka, “Addressing cold start in product search via empirical bayes,” in Proceedings of the 31st ACM international conference on information & knowledge management, 2022, pp. 3141–3151.
[53]
Z.-H. Deng, C.-D. Wang, L. Huang, J.-H. Lai, and S. Y. Philip, “G 3 SR: Global graph guided session-based recommendation,” IEEE transactions on neural networks and learning systems, vol. 34, no. 12, pp. 9671–9684, 2022.
[54]
Z. Ou, X. Zhang, Y. Zhu, S. Lyu, J. Liu, and T. Ao, “LS-TGNN: Long and short-term temporal graph neural network for session-based recommendation,” in Proceedings of the AAAI conference on artificial intelligence, 2025, vol. 39, pp. 12426–12434.
[55]
Q. Zhao, Y. Zhang, D. Friedman, and F. Tan, “E-commerce recommendation with personalized promotion,” in Proceedings of the 9th ACM conference on recommender systems, 2015, pp. 219–226.
[56]
S. Wang, L. Cao, Y. Wang, Q. Z. Sheng, M. A. Orgun, and D. Lian, “A survey on session-based recommender systems,” ACM Computing Surveys (CSUR), vol. 54, no. 7, pp. 1–38, 2021.
[57]
M. Ludewig and D. Jannach, “Evaluation of session-based recommendation algorithms: M. Ludewig, d. jannach,” User Modeling and User-Adapted Interaction, vol. 28, no. 4, pp. 331–390, 2018.
[58]
Z. Pan, F. Cai, W. Chen, H. Chen, and M. De Rijke, “Star graph neural networks for session-based recommendation,” in Proceedings of the 29th ACM international conference on information & knowledge management, 2020, pp. 1195–1204.
[59]
D. Yu, Q. Li, H. Yin, and G. Xu, “Causality-guided graph learning for session-based recommendation,” in Proceedings of the 32nd ACM international conference on information and knowledge management, 2023, pp. 3083–3093.
[60]
R. Qiu, J. Li, Z. Huang, and H. Yin, “Rethinking the item order in session-based recommendation with graph neural networks,” in Proceedings of the 28th ACM international conference on information and knowledge management, 2019, pp. 579–588.
[61]
Z. Wang, W. Wei, G. Cong, X.-L. Li, X.-L. Mao, and M. Qiu, “Global context enhanced graph neural networks for session-based recommendation,” in Proceedings of the 43rd international ACM SIGIR conference on research and development in information retrieval, 2020, pp. 169–178.
[62]
X. Xia, H. Yin, J. Yu, Q. Wang, L. Cui, and X. Zhang, “Self-supervised hypergraph convolutional networks for session-based recommendation,” in Proceedings of the AAAI conference on artificial intelligence, 2021, vol. 35, pp. 4503–4511.
[63]
S. Qiao, W. Zhou, J. Wen, H. Zhang, and M. Gao, “Bi-channel multiple sparse graph attention networks for session-based recommendation,” in Proceedings of the 32nd ACM international conference on information and knowledge management, 2023, pp. 2075–2084.
[64]
M. Lv, X. Liu, and Y. Xu, “Dynamic multi-interest graph neural network for session-based recommendation,” in Proceedings of the AAAI conference on artificial intelligence, 2025, vol. 39, pp. 12328–12336.
[65]
D. Garg, P. Gupta, P. Malhotra, L. Vig, and G. Shroff, “Sequence and time aware neighborhood for session-based recommendations: STAN,” in Proceedings of the 42nd international ACM SIGIR conference on research and development in information retrieval, 2019, pp. 1069–1072.

  1. Linjiang Guo, Nitin Bisht, and Yifan Yin are with the School of Computer Science, University of Technology Sydney, Australia. E-mails: linjiang.guo@student.uts.edu.au, nitin.bisht@student.uts.edu.au, Yifan.Yin-2@student.uts.edu.au.↩︎

  2. Shiqing Wu is with the Faculty of Data Science, City University of Macau, Macau SAR. E-mail: sqwu@cityu.edu.mo.↩︎

  3. Guandong Xu is with The Education University of Hong Kong, Hong Kong SAR. E-mail: gdxu@eduhk.hk.↩︎

  4. Corresponding authors.↩︎