July 01, 2026
Proactive assistance with large language models (LLMs) has received growing attention in the human computer interaction (HCI) community. However, most past work on proactive LLMs’ assistance has focused on adult users and task-oriented settings, leaving open how such systems could support children, whose interests and needs are often expressed through gaze and other nonverbal behaviors rather than explicit requests. In this study, we focus on two key challenges of proactive assistance in children’s picture exploration: when to provide assistance and what assistance to provide based on children’s nonverbal behaviors. To address these challenges, we introduce Ollie, a gaze-informed proactive artificial intelligence (AI) assistant that offers short narrative descriptions based on where a child is looking. Ollie uses children’s gaze to estimate their attention, identify their current visual focus, and select a related picture region for the LLM to verbally describe. In a within-subject experiment, we compared gaze-informed assistance with random assistance. Results show that gaze-informed assistance kept children’s attention on their current focus for a longer period of time, and guided them more effectively to related picture regions. Children, parents, and a participating kindergarten teacher viewed Ollie positively and consider that it better matched children’s interests when compared with the random assistance. This work shows the feasibility of using gaze as an implicit input for proactive AI assistance for children and provides design implications for future child-centered AI systems.
<ccs2012> <concept> <concept_id>00000000.0000000.0000000</concept_id> <concept_desc>Do Not Use This Code, Generate the Correct Terms for Your Paper</concept_desc> <concept_significance>500</concept_significance> </concept> <concept> <concept_id>00000000.00000000.00000000</concept_id> <concept_desc>Do Not Use This Code, Generate the Correct Terms for Your Paper</concept_desc> <concept_significance>300</concept_significance> </concept> <concept> <concept_id>00000000.00000000.00000000</concept_id> <concept_desc>Do Not Use This Code, Generate the Correct Terms for Your Paper</concept_desc> <concept_significance>100</concept_significance> </concept> <concept> <concept_id>00000000.00000000.00000000</concept_id> <concept_desc>Do Not Use This Code, Generate the Correct Terms for Your Paper</concept_desc> <concept_significance>100</concept_significance> </concept> </ccs2012>
The advance of large language models (LLMs) has created new opportunities to use AI as a proactive partner that offers timely, context-aware assistance by anticipating users’ needs in a wide range of activities [1], [2]. While reactive assistance requires users to stop their ongoing task and formulate explicit commands or queries, typically in the form of written or spoken prompts, proactive assistance can provide support without requiring the user to interrupt their task to initiate it and can address needs that are difficult or cumbersome to articulate verbally [2], [3]. This is particularly relevant for a wider adoption of AI-based assistance as formulating effective prompts remains a known barrier for many users, who struggle to translate their goals into queries that elicit useful responses [4], [5], given the functional flexibility of LLM-based assistants: users must not only articulate their need, but also identify which of the assistant’s many possible capabilities would best support them at a given moment [5], [6].
Proactive assistance can be especially useful for children, such as preschool and early elementary school children, who may not yet have sufficient skills to formulate clear prompts, request specific help, or even recognize what kind of support they might benefit from during an activity. At the same time, prior work has shown that children often communicate their interests and intentions through non-verbal behaviors, including gaze patterns [7] and facial expressions [8], which can indicate internal states such as attention, engagement, confusion, curiosity, and emotional response beyond verbal exchange. However, little is known about how to leverage information encoded in children’s behavioral signals, such as eye gaze, for the design of proactive AI assistance, especially in open-ended activities where their interests are dynamic and often not explicitly stated.
In this study, we focus on providing proactive assistance for children’s picture exploration, a common activity used in books to support children’s visual literacy [9], narrative understanding [10], and emotional interpretation [11]. In practice, children often benefit from adult support during such activities, for example through storytelling, guiding questions, or open dialogue with peers [12]. Nevertheless, such support is not always available or easy to provide. Shared reading requires active communication and participation from both adults and children, while working parents may have limited time for these interactions [13]. Moreover, even when adult support is available, providing such support can be challenging for parents and teachers, who may struggle to come up with engaging stories or questions that align with children’s moment-to-moment attention and interests [14]. This challenge has motivated a range of efforts to use AI to support children’s engagement during reading, including picturebook storytelling [14], visual drawing [15], [16], and interactive narratives [17]. Still, most existing approaches rely on children or parents to provide explicit input, leaving a gap for real-time, autonomous systems that can offer meaningful narration and storytelling during picture exploration by observing children’s behavior alone.
However, developing such promising proactive assistance for children still poses significant practical challenges. First, a fundamental challenge is when to provide assistance: how to determine suitable moments for intervention based on the child’s ongoing (dis)engagement during picture exploration. Second, when the timing is identified, what assistance to provide: what story to tell about the picture so that it matches a child’s interest and engages them in the exploration process. To address these challenges, we use children’s gaze as a real-time behavioral signal for deciding when to provide assistance and what visual content the assistance should build on. As illustrated by the gaze-informed assistance in 1, we developed Ollie, a proactive AI assistant that starts by telling a story about an area a child is currently engaged with and then guides their attention toward another area they previously fixated on and that is visually related. Concretely, we use gaze information to decide when to provide assistance by using a Hidden Markov Model (HMM) to estimate whether the child is engaged with the currently viewed area, or Area of Interest (AOI). Gaze history within the picture also helps Ollie decide what to talk about by identifying another relevant AOI that can be connected to the child’s current focus.
To evaluate the effectiveness of Ollie, we compared gaze-based assistance with a random-assistance baseline, in which both the primary and secondary AOIs were randomly selected from the picture while using the same narration structure. We examined three aspects: how effectively the assistance guided children’s attention to the narrated AOIs, how children incorporated the assistance into their post-exploration verbal descriptions, and how children and accompanying adults perceived the two approaches. Our results show that gaze-based assistance provided more effective visual guidance than the random one, where AI assistance is provided randomly without gaze information. The gaze-based assistance was also preferred by most children, parents, and the participating teacher, who perceived it as better aligned with children’s interests compared with the random assistance.
Our work contributes the following:
We developed Ollie, a gaze-informed proactive AI assistant for children’s picture exploration that uses children’s gaze to inform both when support should be offered and what visual content the LLM should build on.
Through a user study comparing gaze-informed assistance with random assistance, we show that gaze-informed narration can better sustain children’s attention, guide them toward related picture regions, and make the assistance feel more aligned with children’s interests.
Beyond picture exploration, our findings demonstrate that eye gaze can be a promising signal for proactive AI assistance in open-ended activities, where users’ needs and interests are difficult to anticipate or articulate.
Picturebook reading plays an important role in children’s language development and long-term cognitive growth [18]–[22]. It supports language acquisition [23], narrative comprehension [22], memory [20], creativity [24], and overall school readiness [25]. To facilitate children’s picture reading, digital technologies have been explored as a way to extend traditional printed books, with reported benefits for maintaining attention, vocabulary learning, and reading comprehension [26], [27]. Early efforts mainly added multimedia features to storybooks, such as sound effects, animations, colorful pictures, and varied text styles [23], [28]. More recent systems provide support in multimodal, adaptive, and socially engaging ways [29]–[34]. These include multimedia narration and animation [30], interactive elements and gamification for engagement [35], remote shared-reading tools for social-emotional connection [32], and personalized learning environments supported by adaptive systems [36]–[38].
However, one persistent challenge remains for these support designs: children do not always clearly express what they are interested in, confused by, or ready to explore next [39]–[43]. This has motivated research that uses behavioral signals, such as eye gaze, to analyze children’s attention and its effect on engagement or comprehension to inform the digital support design [40], [44]–[46]. For example, Skibbe et al. showed that children paid the most attention to text when it was both read aloud and visually highlighted [44]. Eng et al. found that the design of picture books with removed extraneous details, leaving only illustrations relevant to the text, can significantly help improve child attention and reading comprehension [47]. Similarly, Liu et al. found that animations in digital picture books improved comprehension only when they were highly relevant to key story elements [46].
With the rise of LLMs, AI systems have become increasingly capable of supporting children’s recreational and learning activities [14], [15], [17], [48]. Still, without a careful understanding of children’s needs and thoughtful interaction design, these systems may provide support that is poorly timed, poorly grounded, or misaligned with children’s ongoing cognitive states [48]–[50]. For example, Sun et al. found that although AI storytelling systems can provide immersive interactions, they still struggle to respond when children show implicit signs of confusion [51]. Kurian similarly argues that conversational AI may simulate empathy, but it lacks the ability to read young children’s subtle emotional and contextual cues, which can lead to responses that do not match children’s actual needs [49]. Jiao et al. further note that when LLMs do not adapt to children’s cognitive developmental levels, they may provide support that is too complex, increasing cognitive load and causing frustration or confusion [52].
These findings confirm that implicit behavioral signals, such as gaze, remain important for designing AI-based digital support for children. Rather than only asking which multimedia features should be added to a picturebook, the design problem in this AI age becomes how a system can infer what a child is currently attending to and decide what support would be useful at that moment. This motivates us to develop a gaze-informed AI assistance approach that adapts to children’s moment-to-moment visual attention during picture exploration.
The concept of proactive AI assistance has a long history [53]. One of the earliest approaches is the mixed-initiative interaction, where both the user and the system can take the initiative to guide a process or achieve a goal, particularly in productivity tool design [53], [54]. Recent advances in LLMs have made proactive AI assistance more feasible, as these models can interpret rich contextual information and provide support through natural language. However, one central challenge remains consistent: deciding when to intervene [1], [3], [4]. Liu et al. [1] proposed Inner Thoughts, where chatbots internally generated and assessed possible contributions, speaking only when there was a strong reason to intervene. Chen et al. [4] triggered programming assistance during idle, rather than active, coding states. Zhao et al. [3] timed data-visualization assistance using cues such as prolonged pauses and annotations that suggested data inaccuracies or inconsistencies.
Another challenging question is what assistance to give, when the optimal “when” is identified. Yang et al. [2] used an LLM to generate social advice from predefined contextual inputs, including conversation content, nonverbal cues, social factors, and user/partner personas. Chen et al. [4] generated programming support by conditioning the LLM on the current coding context and a predefined taxonomy of help types. Zhao et al. [3] first mapped user difficulties to predefined problem categories, then generated corresponding support such as onboarding tips, exploration guidance, or error reminders. Liu et al. [1] generated candidate contributions and selected among them by scoring their intrinsic motivation to participate.
Only recently, researchers started exploring how to combine gaze data with LLMs for adaptiing the assistance content[55]–[57]. Danry et al. developed an LLM-driven gaze-aware assistant to support reading comprehension, but the assistance was provided reactively, i.e. only after the users completed each passage rather than during reading in real time [55]. Bae et al. also focused on reading support, moving the granularity of gaze-based assistance to the sentence level through deviation-based detection over gaze features [56]. Liu et al. proposed an immersive learning system that generated quizzes based on students’ attention distribution across lecture sections, and showed that gaze-based personalized assistance could improve attention management, motivation, engagement, and comprehension [57]. These works have mainly focused on difficulty-centered or instructional settings, while less is known about how gaze can support interest-centered, open-ended visual exploration. Combining gaze information with proactive assistance is particularly relevant in interaction with children, whose attention and interests are often expressed non-verbally through gaze rather than articulated prompts.
This section describes the design of Ollie’s gaze-informed proactive assistance pipeline. In response to the two key challenges introduced in 1—when to provide assistance and what assistance to provide—the pipeline consists of three stages: (1) detecting the child’s attentional state from gaze behavior, that is whether they are interested in a specific part of an image (primary AOI) or still exploring the full picture, (2) selecting a secondary AOI that is relevant to the child’s current visual focus but has not yet been talked about, and (3) generating a short narrative that connects the current primary AOI and the selected secondary AOI.
To address the challenge of when to provide assistance, we want to infer from the gaze behavior of a child whether they are currently interested in a specific part of an image or exploring. Therefore, we model the child’s attentional engagement as a latent temporal process using a two-state hidden markov model (HMM) corresponding to the child’s attention states: interested and exploring. At each time step \(t\), corresponding to a 500 ms window, the model observes a gaze-derived feature \(x_t\) that captures how strongly the child’s gaze is focused on a single AOI.
Each hidden state is associated with a Gaussian Mixture Model (GMM) emission distribution with two components, \[p(x_t \mid z_t = i) = \sum_{m=1}^{2} w_{im}\,\mathcal{N}(x_t \mid \mu_{im}, \Sigma_{im}),\] where \(w_{im}\), \(\mu_{im}\), and \(\Sigma_{im}\) denote the mixture weight, mean, and covariance of component \(m\) under state \(i\), respectively.
To compute the gaze feature, \(x_t\), used as the observation in the HMM, we extend standard AOI-based gaze measures [58] by defining a dynamic, area-normalized fixation time. Specifically, for each AOI on the picture, we compute the fixation time, defined as the sum of fixation durations that fall inside that AOI within the time window (500ms). Then, to account for different sized AOIs, all fixation time values are normalized by the AOI’s pixel coverage [59], [60]. Intuitively, a high value of \(x_t\) indicates that the child is focusing primarily on one dominant region of the picture, whereas a low value indicates that gaze is more dispersed across AOIs or not clearly concentrated on any single region.
Next, we identify the dominant AOI in that window—the AOI with the highest fixation time. We then compute the Fixation Ratio on the Dominant AOI (FR-D) as the proportion of fixation time on the dominant AOI divided by the total fixation time across all AOIs in that window. This ratio reflects how strongly the child’s gaze is concentrated on a single AOI during the time window—i.e., how likely the dominant AOI captured and held the child’s visual attention, as indicated in Equation 1 .
\[\text{FR-D}_t = \frac{\tilde{F}_{\text{dominant},t} + 1}{\sum_{i \in \mathcal{A}} \tilde{F}_{i,t} + 2} \label{eq:frd}\tag{1}\] where \(\tilde{F}_{i,t}\) is the area-normalized fixation time for AOI \(i\) in window \(t\). The constants 1 and 2 are Laplace smoothing terms to avoid extreme values near 0 or 1. The resulting feature value \(x_t = \text{FR-D}_t \in (0,1)\) serves as the observation at time \(t\).
To ensure stable parameter initialization, we include a 5 s warm-start phase (10 consecutive time steps) during which no state inference is performed. A batch EM algorithm is applied to the warm-start data to estimate the initial parameters \(\Theta_0 = \{\pi, A, w, \mu, \Sigma\}\), where \(\pi\) denotes the initial state distribution and \(A\) the transition matrix. After initialization, the system transitions to an online EM procedure: at each subsequent 500 ms interval, the model (i) updates the filtered posterior probabilities \(P(z_t \mid x_{1:t})\) via the forward algorithm, (ii) computes local expected sufficient statistics for transitions and emissions, and (iii) performs incremental parameter updates using exponential moving averages. Different update rates are applied to ensure stability and responsiveness: 0.1 for emission parameters and 0.05 for transition and mixture parameters.
To address the what assistance to provide challenge, once Ollie determines that assistance should be provided, it generates gaze-informed narrative assistance that is grounded in the child’s current visual focus and extends the narration toward related, previously unexplored parts of the picture. The core idea is to use the gaze-based attention detection method described in 3.1 to identify the primay AOI that currently captures the child’s visual attention, and then proactively introduce a second, contextually relevant but unassisted AOI to encourage further exploration. Specifically, the selection system integrates two components:
Each AOI is represented using three complementary feature types: (1) a list of objects contained within the region (semantic features); (2) its on-screen coordinates (spatial features); and (3) a gaze-time log indicating when the region was last viewed (temporal features). These representations support evaluation of how each AOI relates to the child’s current visual focus.
For each AOI, the system computes three distances relative to the currently fixated AOI: semantic distance (cosine distance over TF–IDF vectors of object lists), spatial distance (Euclidean distance between AOI centers), and temporal distance (time elapsed since the AOI was last fixated). Distances are normalized and combined into a weighted one:
\[D = 0.2D_{\text{semantic}} + 0.4D_{\text{spatial}} + 0.4D_{\text{temporal}}. \label{eq:relevance95distance}\tag{2}\] The AOI with the lowest distance score among those that have not yet received storytelling assistance (the unassisted set) is selected as the secondary AOI to be described.
We generate narrative assistance using LLMs with three sources of visual context: the currently gazed primary AOI (Section 3.1), the selected secondary AOI (Section 3.2), and the full image. The model is instructed to treat the task as a story-chaining process, in which each new narration builds on the previously generated content.
For the first assistance instance, we provide the model with a short, predefined mini-context that specifies basic situational information (e.g., place or time). The model then generates a short narrative that (1) describes the main characters or events in the primary AOI, (2) introduces the secondary AOI using a simple spatial or contextual bridge (e.g., “nearby” or “not far away”), and (3) connects the two regions by describing a plausible relationship or interaction. To encourage factual consistency, we also provide the model with lists of pre-detected objects in both AOIs for cross-checking. The narration ends with a brief open-ended hook (e.g., a question) grounded in the secondary AOI to spark the child’s curiosity.
This hook is also used to connect subsequent assistance instances where the model first briefly addresses the hook from the previous narration and then repeats the same structure: describing the current primary AOI, introducing a new secondary AOI, and ending with a new hook to encourage continued exploration. In the user study, Ollie generated and delivered all spoken narrations in German, which was the language used with participating children. In this paper, we translate example narrations and prompt excerpts into English for readability. The complete prompt we used in the study is presented in the appendix.
2 gives an example to demonstrate how exactly the whole proactive assistance works: firstly, the system responds to the child’s visual attention by identifying the currently gazed-at primary AOI and selecting the most related secondary AOI that has not been talked about yet to encourage further exploration. The LLM then generates a short description that links these regions within the context of the full picture, illustrating how gaze information is used to guide attention and support visual sensemaking.
To evaluate the effects of Ollie’s gaze-informed proactive assistance, we conducted a within-subject experiment addressing the following research questions:
RQ1: How does gaze-based proactive assistance affect children’s visual attention during picture exploration compared to a random assistance?
RQ2: How does gaze-based proactive assistance influence children’s verbal description after exploration, in terms of their uptake of the assistance and the quality of their description, compared to a random assistance?
RQ3: How do children and their parents/guardians perceive gaze-based versus random proactive assistance in terms of its preference, usefulness, and overall experience?
Prior to the main within-subject experiment, we conducted a pilot study to test the feasibility of the system and refine the experimental protocol for the main study. The pilot explored key design parameters, including the frequency and number of assistance instances, the choice of image materials, and different candidate experimental conditions, including a no-assistance baseline.
For this purpose, we recruited and tested the proactive assistant system with ten children aged four to ten years from the local community. The pilot study ran on a desktop computer with a 27-inch screen, to which a Tobii Pro Fusion eye tracker was attached.
In the pilot, each child viewed between 6 and 15 picture-book–style illustrations created by different artists, each presenting a visually-rich scene containing multiple characters, objects, and ongoing events. We first asked children for brief feedback for each image and whether they found them interesting or suitable for storytelling. Based on this feedback, we kept images that were consistently perceived to be engaging among the pilot participants. We then set the total number of images in the main study to six, with each image viewed for about three minutes. We chose this duration because we had learned from the pilot sessions that longer sessions cause fatigue, which makes it hard for children to sit still in front of the eye tracker. In addition, we set the interval between successive proactive assistance instances to a minimum of three seconds. In other words, the AI assistance would only be triggered if at least three seconds had elapsed since the previous one.
We also included a no-assistance baseline condition in the pilot, in which children were asked to explore the picture freely without any support nor gaze tracking. We observed substantial individual differences in this condition: some children quickly skipped through the images, while others remained on a single image for a long time. This variability made it difficult to align this condition with the assisted conditions in a controlled manner. Given prior evidence that children can be actively engaged through AI-mediated storytelling and questioning during reading activities [14], [61], we focused our investigation on how proactive assistance should be triggered and grounded. Therefore, we decided to omit the no-assistance condition in the main study and instead to compare gaze-based assistance with a random-assistance baseline.
For the main study, we recruited 26 children through mailing lists, physical flyers, and word of mouth, and conducted the study in two settings: a laboratory and a local kindergarten. Of these, 22 children completed the study. Prior to scheduling sessions, we emailed parents to verify that their child already engaged in at least one picture‑book reading activity per week, either independently or with a parent. This simple check ensured that participants were familiar with picture‑book exploration. Of the 26 children, two declined to finish exploring all images with Ollie. Also, one session was terminated due to an unrecoverable technical issue, and one child had to leave early when a parent arrived for pickup, before completing the study.
The final sample consisted of 22 children (12 female, 10 male): Among them, we had 11 five-year-olds, 4 six-year-olds, 4 seven-year-olds, and 3 eight-year-olds. Of the final sample, 13 children read picture books once per week, whereas the remaining 9 did so more than once per week.
In addition, 16 parents participated in the study and remained present throughout the sessions conducted in the laboratory setting. Among them, three parents each brought two children and therefore attended two consecutive experimental sessions. One day-care teacher also participated in the study during the kindergarten sessions. No parents declined to participate in the interview.
For the main user study, we prepared six images in a visually rich Wimmelbook-style picture-book format.1 The images were selected from three picture books by three different creators, covering three themes: construction sites, humans and animals, and emergency incidents, each having two scenes. In addition, we used AI to generate one image that approximates the visual style and composition of the study materials (3 (a)) for the pilot study and for illustration in the paper (as other images are copyrighted). The six images were organized into three theme-matched pairs. For each participant, one image from each theme-matched pair was assigned to the gaze-based condition and the other to the random condition, resulting in three images per condition. This assignment was randomized across participants so that each condition included one image from each theme.
For each image, AOIs were manually defined to segment the scene into meaningful regions based on interior and exterior spaces, as shown in 3 (b). To maintain comparable visual complexity across images, each AOI contained two to five key objects, and each image has a total of 12 to 15 AOIs.
Figure 3: Image Example. a — Image used in the Experiment, b — AOIs on the Image
We employed two assistance conditions in this study: (1) a random assistance condition and (2) a gaze-based assistance condition.
In both conditions, the same prompt described in 3.3 was used to generate spoken assistance with a large language model (GPT-4o). The generated responses were delivered to the child via text-to-speech using the Azure service. Each assistance instance followed the same structure: the system first described a primary AOI (the starting region of the narration) and then introduced a secondary AOI to encourage further exploration of the picture.
The two conditions differed only in how these two AOIs were selected. In the gaze-based assistance condition, the primary AOI was determined based on the child’s current visual attention using the attentional state detection method described in 3.1. The secondary AOI was then selected using the relevance estimation procedure described in 3.2. In the random assistance condition, both the primary AOI and the secondary AOI were selected at random from the set of available AOIs in the picture. All the other features were kept equivalent.
To support children’s understanding of what the narration referred to, the primary AOI was visually highlighted for 3 seconds in both conditions when the assistance was delivered. The secondary AOI was introduced only through speech and was not highlighted, so that any attention shift toward it reflects the quality of the verbal suggestion rather than an explicit visual cue.
The study was conducted in two settings: a laboratory environment and a local kindergarten. In the laboratory setting, children participated together with their parent/guardian, who remained present during the session. In the kindergarten setting, sessions were conducted in the closed library with a kindergarten teacher present with each participating child. In both cases, the same experimental setup and instructions were used across settings. Specifically, we used a 27 inch desktop monitor screen with a Tobii Pro Fusion eye tracker operating at 250 Hz to collect data. Each session began with a 5-minute introduction in which we explained the task to the child and demonstrated how the AI assistant, Ollie, would provide spoken guidance during the picture exploration. We also explicitly explained that Ollie would use two different assistance strategies (random and gaze-based), and that we would need their help in evaluating the system by sharing their thoughts and feedback afterwards. During the study, we checked whether children noticed this difference, including whether Ollie seemed to follow their gaze or point out regions on its own. The study had been approved by the Ethics Review Board at the affiliated university prior to data collection.
Each child explored six picture-book–style images in total. The images were organized into two blocks by assistance condition: three images with random assistance and three with gaze-based assistance. Within each block, three images were presented consecutively, with one image randomly selected from each of the three theme-matched pairs. The order of the two blocks was counterbalanced across participants. For each image, Ollie provided three consecutive instances of assistance. After these, we paused the assistance and asked the child whether they wanted to continue exploring the same image or move on to the next one. If the child chose to continue, exploration resumed and Ollie provided another assistance, after which the child was again asked whether they wished to continue or switch to the next picture. This loop continued until the child chose to stop or a maximum time limit of five minutes was reached.
After finishing exploration of each image, and before moving on to the next one, we asked the child to briefly describe their favorite part of the picture. After completing all six images, we conducted a short post-study interview with the child and the accompanying adult (parent/guardian/teacher). We provide interview questions in the following section.
To address our research questions, we evaluated the proactive assistance from three aspects: (1) visual attention during exploration for RQ1, (2) verbal description after exploration for RQ2, and (3) subjective feedback from post-study interviews for RQ3.
To explore how Ollie’s assistance affect children’s visual attention, we analyzed attention to both the primary AOI and the secondary AOI during their respective narrative descriptions. For each assistance instance, we defined a temporal analysis window covering the full narration and divided it into a primary-AOI interval and a secondary-AOI interval. For each interval, we computed three gaze metrics for the corresponding assisted AOI: follow-up fixation rate, the proportion of instances in which the child fixated the AOI; fixation duration, the total time spent fixating that AOI within the interval; and time to first fixation, the latency from the onset of the AOI’s verbal mention to the first fixation on it. We further tracked the dominant AOI across the full assistance window, following the definition in 3.1, to examine how children’s attention shifted over time on the narrated AOIs compared to other regions of the image.
After each of the six presented images, we asked the child to briefly describe their favorite part of the picture. Following the child’s initial response, Ollie delivered a pre-programmed encouraging prompt (e.g., “Good job – did you notice anything else?”) to elicit additional description. This continued until the child indicated that they had nothing else to share. If the child did not provide an initial response, we proceeded directly to the next image.
We transcribed children’s responses and evaluated them regarding two dimensions. First, response to assistance captured the extent to which the child’s description reflected the content introduced by Ollie. This was calculated as the proportion of objects mentioned in the child’s description that overlapped with the objects referred to in the assistance. Second, description quality captured the richness of children’s descriptions. Drawing on prior work that distinguishes object, attribute, and relational/event information [62], [63], we used a three-point coding scheme: 1 for object-only descriptions (e.g., “dog”), 2 for descriptions including attributes (e.g., “a big ball and a small ball”), and 3 for descriptions including events or actions (e.g., “the truck fell over”). We averaged these scores across each child’s valid responses to obtain an overall description quality score.
After completing all six images, each child and the accompanying adult participated in a short post-study interview. The interview focused on three aspects: (1) participants’ overall experience with Ollie, including whether they liked the stories and wanted to use such an assistant again; (2) their preference between the random and gaze-based assistance conditions, together with any underlying reasons; and (3) the perceived usefulness of such an assistant in everyday settings, such as home reading or kindergarten classroom use, as well as possible improvements to the system.
For children, we asked about their impressions of Ollie, which stories or images they liked the most, whether they preferred one assistance strategy over the other, and whether they would like to use Ollie during picture exploration at home. For children who regularly read picture books with parents or teachers, we also asked how Ollie’s storytelling compared with the support they usually receive.
For accompanying adults, the interview focused on whether they would consider using such an assistant at home or in school, which assistance strategy they preferred, how Ollie compared with parent or teacher storytelling, and what changes would make the system more helpful in practice.
The full interview questions are provided in the Appendix.
In the following, we first describe our data pre-processing steps. We then report on the impact of Ollie’s gaze-based proactive assistance on three aspects of children’s picture reading: visual attention (RQ1), verbal description (RQ2), and subjective feedback (RQ3). Specifically, we aggregated the gaze data at the participant level to account for repeated measurements within each child. After checking for normality with the Shapiro-Wilk Test, we performed statistical hypothesis testing, either a paired t-test (in case of normality of continuous variables) or the Wilcoxon signed-rank Test (in case of non-normality or non-continuous variables).
For gaze analysis, we treated the data as time-series gaze records corresponding to individual Ollie assistance instances. In total, we had 405 gaze data (assistance instance) from 22 participants with 203 in gaze-based assistant condition, and 202 in random condition. Each gaze point was captured as an (x,y) screen coordinate, together with their corresponding timestamp. From the raw gaze points, fixations were identified based on specific criteria: low dispersion (25 px) and adequate duration (50 ms), using fixation detection function from the PyGaze package [64]. To remove data with poor tracking quality, we used the highlighted primary AOI shown during each assistance instance as a validation check. As described in the assistance conditions (4.4), the primary AOI was visually highlighted while the assistance was delivered. Following prior work [65], we used the highlighted primary AOI as a validation check. For each assistance instance, we examined whether the participant made at least one fixation on the highlighted AOI. Instances with no such fixation were excluded from the gaze analysis. If this occurred in more than half of the assistance instances for a given image, we excluded the entire image and its corresponding gaze data. If more than three images were excluded for a participant, we excluded the participant entirely, as this would eliminate all data from one assistance condition. After this filtering procedure, one participant was excluded, and 13.2% of the gaze data were removed.
For the verbal-description analysis, we recorded children’s responses to the image-description prompt described in 4.6.2. We transcribed each response and considered a response valid if it contained at least one relevant word that eferred to the image. To maintain balanced data across participants, we excluded a child’s transcription data if more than three image descriptions were invalid. After filtering, the final dataset contained 101 valid descriptions from 18 children.
As described in 4.6.1, we first examined fixations on the two assisted AOIs separately. For the primary AOI, both conditions yielded similarly high follow-up fixation rates (Figure 4 (a)), likely because the primary AOI was visually highlighted in both conditions: the random condition achieved \(M = 91.66\%\), \(SD = 10.27\%\), and the gaze-based condition achieved \(M = 90.81\%\), \(SD = 10.73\%\), with no significant difference between the conditions (\(p > .05\)). In contrast, fixation duration on the primary AOI differed substantially across conditions (Figure 4 (b)). Children with the proactive assistance spent significantly longer fixating on the primary AOI (\(M = 6.48\) s, \(SD = 2.53\)) than with the random assistance (\(M = 1.67\) s, \(SD = 0.75\); \(p < .001\)). These findings suggest that while children noticed the narrated primary AOI in both conditions, gaze-based proactive assistance was more effective at sustaining their attention on the currently relevant region.
Figure 4: Follow-up fixation rate and fixation duration on primary AOI across conditions.. a — Follow-up Fixation Rate on Primary AOI., b — Fixation Duration on Primary AOI.
Next, we examined the secondary AOI, which was selected based on the relevance measure and introduced only through spoken assistance, without any visual highlighting. As indicated in 5 (a), the results show that the gaze-based condition yields a significantly higher follow-up fixation rate on the secondary AOI (M = 54.06%, SD = 20.92%) than the random condition (M = 35.48%, SD = 28.34%), with a statistically significant difference (\(p < .05\)). To directly examine how the two selection strategies differed, we compared the relevance distance between the primary AOI and the selected secondary AOI in each condition, as defined in Equation 2 . In the gaze-based condition, relevance distance between the primary AOI and the secondary one (M = 0.43, SD = 0.10) is significantly lower than the random condition (M = 0.52, SD = 0.13, p < .001). These results suggest that the relevance distance-based selection helped children follow Ollie’s spoken guidance: by choosing a secondary AOI that was closer to the child’s current focus, the gaze-based condition made the next narrated region easier for children to attend to.
Figure 5: Follow-up fixation rate on secondary AOI and relevance distance between primary & secondary AOIs across conditions.. a — Follow-up Fixation Rate on Secondary AOI., b — Relevance Distance
For assistance instances in which children eventually fixated the secondary AOI, we next examined how much attention they allocated to that region. Overall, attention to the secondary AOI remained limited. For fixation duration on the secondary AOI (6 (a)), children spent only a short time looking at the region in both conditions, with a mean fixation duration of \(M = 1.08\) s (\(SD = 0.90\)) in the gaze-based condition and \(M = 1.02\) s (\(SD = 1.13\)) in the random condition; this difference was not statistically significant (\(p > .05\)). Even among children who attended to the secondary AOI the most, the average fixation duration did not exceed 3.5 s. We also found that children generally needed substantial time to locate the AOI referenced by Ollie. Time to first fixation was \(M = 6.65\) s (\(SD = 2.26\)) in the gaze-based condition and \(M = 5.65\) s (\(SD = 4.26\)) in the random condition, again with no significant difference between conditions (\(p > .05\)).
Figure 6: Follow-up fixation duration and first time to fixation on secondary AOI across conditions.. a — Fixation Duration on Secondary AOI., b — First Time to Fixation on Secondary AOI.
Beyond the aggregate gaze metrics computed over the full assistance window, the temporal distribution of dominant AOI roles provides insight into how children’s attention evolved as Ollie guided them across two AOIs selected under different strategies. Specifically, we conducted a temporal analysis of dominant AOI role based on \(\tilde{F}_{\text{dominant},t}\) in 1 . Because assistance duration varied across instances (\(M = 34.05\) s, \(SD = 3.98\) s), we normalized each assistance instance by dividing the primary-AOI narration phase and the secondary-AOI narration phase into five bins each. Within each bin, we identified the dominant AOI by comparing proportional fixation duration across AOIs and classified it as the primary AOI, the secondary AOI, or other AOIs.
As shown in 7, both conditions exhibited a shift in dominant attention from the primary AOI toward the secondary AOI after the onset of the secondary-AOI description. However, the transition patterns differed. In the gaze-based condition (7 (b)), the secondary AOI already accounted for a non-trivial share of dominant attention during the primary narration phase, suggesting that relevance-based selection increased the likelihood that children had already noticed the upcoming region. Moreover, attention to the primary AOI remained relatively sustained into the secondary narration phase, consistent with the longer fixation duration on the primary AOI reported in 4 (b). By contrast, in the random condition (7 (a)), dominant attention to the primary AOI dropped more sharply after the secondary AOI was introduced, accompanied by a rapid increase in attention to the secondary AOI. Taken together, these patterns suggest that gaze-based assistance supported a smoother and more gradual transition between the two narrated regions, whereas random assistance elicited a more abrupt reorientation of attention toward the newly mentioned AOI.
Figure 7: Temporal shift of dominant AOI role across the full assistance window under the two assistance conditions. For each assistance instance, the narration was divided into a primary-narration phase and a secondary-narration phase, and each phase was normalized into five equal temporal bins. Within each bin, the dominant AOI was classified as the primary AOI, secondary AOI, or other AOIs based on proportional fixation duration. The dashed line marks the onset of the secondary-AOI description.. a — Random assistance condition., b — Gaze-based assistance condition.
As described in 4.6.2, we coded children’s verbal descriptions using two measures: response-to-assistance and description quality. Response-to-assistance rate captured the extent to which a child’s description overlapped with Ollie’s narration, with higher values indicating greater overlap. We observed substantial individual variability on this measure: some children did not mention any content introduced by Ollie (min = 0.00), whereas others produced descriptions that fully overlapped with the assistance (max = 1.00). Across the full sample, the standard deviation was 0.26. However, there was no significant difference between the two assistance conditions, as shown in 8 (a): the mean response-to-assistance rate was 0.51 (SD = 0.30) in the random-assistance condition and 0.61 (SD = 0.21) in the gaze-based assistance condition (\(p > .05\)).
A similar pattern was observed for description quality, which measured how richly children described objects, attributes, and events. Again, we found substantial individual variability (SD = 0.68). However, description quality did not differ significantly between conditions. The mean score was 2.35 (SD = 0.66) in the random-assistance condition and 2.28 (SD = 0.72) in the gaze-based assistance condition (\(p > .05\)), as indicated in 8 (b). Overall, these findings indicate substantial individual differences in how children described the images. At the same time, the response-to-assistance measure suggests that children generally attended to and incorporated content from Ollie’s narration as the mean overlap was above 50% in both conditions, suggesting that the assistance was generally effective.
Figure 8: Response to assistance and description quality of verbal description across conditions.. a — Response to Assistance Rate., b — Description Quality Score.
We analyzed the post-study interview data using a combination of descriptive coding and exploratory thematic analysis. For questions with clear evaluative or preference-based responses, such as whether participants liked Ollie, which type of assistance they preferred, and whether they would use such a system at home or school, responses were coded into simple categories and summarized using counts. For open-ended responses, one author conducted an iterative thematic analysis to identify recurring reasons related to participants’ preferences, perceived usefulness, and suggestions for improvement. The coding scheme was refined through repeated review of the interview notes to ensure that categories remained consistent across responses. Children, parents, and teacher responses were analyzed separately. We describe high-level themes from this analysis for each of the stakeholders (children, parents, and teachers).
Preference for the gaze-based assistance. When asked directly which assistance condition they preferred, 15 of 22 children chose the gaze-based condition, 4 chose the random condition, and 3 reported no preference. The most common reasons for preferring the gaze-based condition were better alignment with their interests (8 mentions) and a greater sense of control over what Ollie talked about (4 mentions). In contrast, children who preferred the random condition referred to its greater surprise and lower sense of constraint, for example not having to keep looking at a particular part of the picture. When asked which version they would want to use at home, this pattern largely remained, as 11 of the 15 children who preferred the gaze-based condition also selected it for home use. However, the preference was weaker in this context, as several children switched to the random version or reported being unable to decide, suggesting that preferences expressed during the study did not always translate directly into everyday use intentions.
Overall experience when being assisted with Ollie was positive across conditions. 20 of the 22 children reported that they liked Ollie’s descriptions in the gaze-based condition, and 19 of 22 reported liking them in the random-assistance condition. Similarly, when asked whether there were parts of Ollie’s descriptions that they did not like, most children responded “no” in both conditions (18/22 in the gaze-based condition and 19/22 in the random condition), suggesting that there was no clear rejection of Ollie’s assistance.
Would-use intention was mostly positive. Fifteen of the 22 children said that they would want Ollie to accompany them while looking at pictures, whereas 3 said no and 4 were unsure. In explaining this choice, among the 15 children who said they would want Ollie, 4 children mentioned that Ollie added information beyond what was directly visible in the picture, and 3 described the experience as fun. At the same time, 11 of the 22 children were unable to provide a specific reason for their answer, indicating that while the majority of children are willing to use Ollie, they were often not able to articulate exactly why.
Comparison of the reading experience with parents. Fourteen of the 22 children reported that they did not usually read picture books with a parent, leaving only 8 children who could make a meaningful comparison between reading with parents and reading with Ollie. Among those 8 children, 3 preferred Ollie’s storytelling, 3 rated Ollie and their parent as comparable, and 2 preferred their parent’s storytelling. This finding suggests no clear pattern of preference.
Parents strongly preferred the gaze-based assistant. Fifteen of the 16 parents preferred the gaze-based version over the random-assistance version. The most common reason was that gaze-based assistance better reflected what the child was currently interested in. Several parents also felt that this version made the child more active in guiding the experience. The only parent who preferred the non-gaze-based version raised privacy and security concerns related to eye tracking.
Parents’ evaluations of Ollie’s performance were generally positive, although often with limitations around the quality of interactions. No parent explicitly rejected Ollie’s performance. Instead, most described it positively while also noting limitations, particularly regarding voice quality, interactivity, or linguistic naturalness. This suggests that parents generally saw promise in the system, but also viewed its current form as needing refinement before broader everyday use.
Parents were divided on whether they would use Ollie or not. When asked whether they would consider using such an AI assistant in children’s daily reading activities, 7 of the 16 parents said no, 6 said yes, and 3 were uncertain. Compared with the children’s responses, parents were more cautious overall. The most common concerns were screen time and the wish to preserve parent-child interaction during reading. For example, some parents emphasized that children already spend enough time with screens, and others stressed that shared reading should not lose its role as a bonding activity. Parents who responded positively often described Ollie as a useful supplement rather than a replacement, especially when a child is reading alone, when a parent is busy, or as a possible learning support such as for language learning.
Suggestions for improvement focused mainly on language quality and interactivity. When asked how the system could be improved, parents most often mentioned problems with language quality, including grammar issues, inappropriate word choice, unnatural pronunciation, delays in story generation, repetitive narration, and a voice that sounded insufficiently natural. A second recurring suggestion was to make Ollie more interactive. Parents proposed that the system should allow children to ask questions, explain unfamiliar words, respond more clearly to children’s input, and generally support a more conversational form of storytelling. Some parents also suggested specific use cases, such as bedtime listening or language-learning activities on a tablet.
The kindergarten teacher saw Ollie as a situational classroom support tool. The teacher suggested that it could be useful during rest periods, with older preschool children, or as part of specific projects, rather than in routine classroom use. Their preference between assistance types depended on the child and the activity: the random version could help guide children who fixate only on one part of a picture, whereas the gaze-based version seemed better suited for free exploration. The main limitation she identified was the lack of full interactivity, particularly Ollie’s inability to answer children’s questions. She also suggested possible use cases beyond picture exploration, including support for children with language barriers and for explaining classroom routines.
A key challenge in proactive AI support for children is determining how to provide help at the right moment without requiring children to explicitly ask for it. Our findings suggest that gaze is a promising signal for this purpose, as it captures children’s interests in a natural way without requiring much effort. For RQ1, we found that during picture exploration, gaze-based proactive assistance kept children’s attention on the primary AOI for longer and guided them more frequently towards the secondary AOI than random assistance. Gaze metrics also indicated longer, more sustained engagement with the visual content that was talked about. However, regarding RQ2, this increase in visual attention did not translate into significant improvements in children’s verbal descriptions, either in terms of the mentions of the objects introduced by the assistance or the overall quality of their description. For RQ3, while there were some reservations about the quality of interactions, particularly mentioned by parents, we found a general trend that gaze-based assistance was preferred by the majority of children and accompanying adults.
These finding suggests that instead of interrupting the activity with verbal input, the gaze-based proactive assistance system can effectively use the child’s ongoing visual attention to decide when to speak and what part of the picture to talk about to better engage the child in the storytelling. In this sense, gaze-based proactivity makes the assistance more grounded in the child’s own activity and better aligned with how children naturally engage with visual materials. More importantly, gaze-based proactivity can guide exploration from where the child is currently looking at in a meaningful way. By starting from the child’s current focus and then introducing another relevant part of the scene, the assistant can support a smoother and more connected exploration process. This makes the interaction feel more child-led, while still allowing the AI to take initiative. We therefore see the combination of eye tracking and LLM-based narration as a promising design direction for children’s picture exploration, because it allows proactive assistance to remain timely, relevant, and aligned with the child’s unfolding interests.
Consistent with prior work that incorporates human gaze as a behavioral signal to inform LLMs about users’ current visual focus within a task context [66]–[68], our study further demonstrates that gaze can encode information that is both important and necessary for effective AI assistance. Rather than treating gaze only as a way to extract a small crop of the user’s current focus, we continuously use gaze information to identify when proactive assistance should be provided and what the user is currently, or may soon become, interested in. Thus, gaze can function as a timely grounding signal for proactive AI assistance, enabling an implicit and low-effort interaction paradigm that is especially promising for AI systems operating in visually rich contexts.
What Ollie has achieved here can naturally extend to children’s recreational and learning activities, where proactive assistance may be especially beneficial because children often express curiosity, confusion, or interest through gaze before they can articulate these states verbally. For example, in drawing, puzzle solving, block building, museum visits, or AR-based exploration, a gaze-aware assistant could notice what a child is focusing on and provide a timely hint, question, explanation, or narrative prompt.
Beyond children’s scenarios, the same idea can extend to visually intensive adult activities, including data analysis, document reading, design review, medical-image inspection, AR-based maintenance, and so on. In these contexts, gaze does more than indicate the visual region to pass to an LLM; it can also reflect user’s real time states when doing the activity such as searching, comparing, hesitation, uncertainty, sustained interest, or missed information. A gaze-based proactive assistant can therefore use these attention patterns to decide when intervention is appropriate and what kind of support is needed, making AI assistance more timely, situated, and minimally disruptive.
Our findings in this study also suggest that gaze-based proactive assistance can support child agency by making children’s ongoing attention influence how the interaction unfolds [69], [70]. In child-centered interactions, agency is not always expressed through explicit choices; children may also communicate interest, uncertainty, or curiosity through gaze and other non-verbal behaviors. By using gaze as input, Ollie could respond to children’s emerging attentional focus rather than relying only on spoken requests or predetermined system choices. Although children did not directly choose which AOI Ollie would describe next, their visual attention still shaped the system’s support, creating a form of implicit agency during picture exploration.
At the same time, making children’s attention consequential represents only an initial form of agency-supportive design. To fully support child agency, proactive assistance should not only adapt to children’s current attentional focus, but also create opportunities for children to notice, evaluate, and redirect that support over time. This connects the present study to self-regulated learning, where learners gradually develop the capacity to monitor their own state, control their actions, and adjust strategies in response to task demands. Molenaar’s concept of hybrid human-AI regulation offers a useful framing for this progression: AI can initially scaffold regulation by detecting learners’ needs and providing timely support, but control should gradually shift back to learners as their self-regulated learning skills develop [71], [72]. In this sense, Ollie’s gaze-based assistance can be viewed as an early form of hybrid regulation, where the system helps guide children’s exploration based on their attention, but future designs should increasingly invite children to confirm, question, revise, or redirect the system’s suggestions.
Our study also contributes to the broader discussion of how AI can be used in children’s reading and storytelling activities. Much of prior work on AI use for children’s reading and storytelling (referred to asinteractive storytelling) has focused on conversational agents that act on children’s verbal input. While LLMs have been proposed as a solution to support natural, flexible, and potentially engaging conversations in interactive storytelling, it is often reported that conversational agents (with/without LLMs) for reading and storytelling exercise limited capabilities primarily due to children’s limited skills of questioning, prompting, and expressing interest verbally [51].
Our approach demonstrates that it is promising to use gaze as information to make inferences about children’s interest and provide proactive guidance in children’s picturebook reading activity. Specifically, gaze-based assistance successfully redirected children’s attention and facilitated broader picture exploration, suggesting that AI can help structure visual engagement without requiring explicit instruction from the child. Further, the proactive support using gaze information was perceived positively by most of our participants. This finding suggests that gaze is a reliable, valid, and ecologically natural signal that can be effectively used to support children during picturebook reading.
We, however, note an important finding from our study: some participants, particularly parents, expressed hesitation about using Ollie because it would introduce yet another screen-based activity on top of children’s already-enough screen time. Further, parents emphasized the value of parent–child reading interactions and were concerned that AI support might reduce opportunities for shared engagement [14]. This insight poses a challenging tension in exploring AI use for children’s reading; on one hand, AI-assisted reading and exploration has certain benefits, which were also highlighted in our study, such as enhancing interest and reducing parental caregiving load. On the other hand, as stated above, stakeholders, especially parents, have concerns about sustained use of such systems. Therefore, future studies should investigate real-world, longitudinal deployment to understand how children and parents make both strategic (weighing benefits and cost of using AI support) and spontaneous (e.g, when parents have to take care of their child while working from home) choices in using AI-based assistance.
We acknowledge several limitations of the present study. First, the final sample size was modest: 22 children completed the study, 21 contributed valid gaze data after preprocessing, and 18 contributed valid transcription data, which limits the generalizability of our findings across ages, reading habits, and settings. Second, the AOIs in our system were manually defined in advance; while this supported controlled analysis in the current study, it limits scalability and depends on annotators’ judgments about what counts as a meaningful region in an image. Third, as also noted by parents and the teacher, Ollie currently offers only limited interactivity. Although it can provide proactive narration, it cannot answer children’s questions, explain unfamiliar words, or sustain a more conversational exchange, which reduces its usefulness in realistic reading situations. Finally, the quality of the generated narration still needs improvement, particularly in grammar, pronunciation, repetition, delay, and voice naturalness, especially when used with languages other than English.
Targeting these limitations, our future work will focus on improving the gaze-based proactive assistant system in several concrete ways.
Firstly, we plan to reduce the need for manually defined AOIs by introducing automatic or semi-automatic region construction. For example, recent work combining eye tracking with state-of-the-art image segmentation suggests a promising direction in which users can collect target segmentation masks simply by looking at the AOI [73], [74]. Building on this idea, future versions of Ollie could replace hand-labeled AOIs with gaze-grounded image regions generated online, which would make the system more scalable and adaptable to new images.
Secondly, recent work suggests that proactive LLM assistance can be interactive and multi-round rather than one-way [1], [75], [76]. For Ollie, this points to a natural next step: extending it from a proactive narrator into a more conversational assistant that can combine gaze-based guidance with other child-centered AI supports, such as question asking [14], drawing creation [15], AR activities [17], etc..
Lastly, although the LLM generally followed the prompt and produced assistance in the expected format, the quality of its output can still be improved. A useful next step is to make the prompting more adaptive to children’s real-time responses. For example, Ollie could use gaze feedback to check whether the child actually looked at the region it just mentioned. If the child followed the narration, Ollie could continue from that region; if not, it could briefly rephrase, reconnect the ignored region to the child’s current focus, or move on without over-directing the child’s attention. Future versions could also incorporate better memory management across assistance rounds. Instead of treating each narration mostly as a local response, Ollie could remember what has already been described, what the child seemed interested in, and which regions were ignored. This would help reduce repetition, maintain coherence across multiple turns or images, and make the interaction feel more responsive to the child’s exploration.
Compared with explicit verbal output, children often express internal states, such as interest and curiosity, implicitly through their behavior during activities such as gaze like picture exploration. As AI systems become more capable of offering timely support, these behavioral signals can provide important cues for making assistance more relevant and less dependent on explicit requests. In this work, we advanced child-oriented proactive AI assistance by using children’s gaze behavior to generate narrative descriptions grounded in what they are currently looking at and to guide their attention toward contextually relevant areas of the picture. To examine the effectiveness of our gaze-informed proactive assistant, Ollie, we compared it with a random-assistance baseline in a user study. The results show that gaze-informed assistance more effectively sustained children’s attention on the narrated content and guided them toward related regions of the image. It was also preferred by both children and accompanying adults, suggesting that gaze can serve as a useful interaction signal for designing proactive AI support in open-ended children’s activities. This work shows that gaze-informed AI can make proactive assistance more responsive to children’s own ways of exploring, offering a path toward AI systems that support children without requiring them to first know what to ask.
Wimmelbooks, or Wimmelbilderbücher, are picture books characterized by dense, full-spread scenes containing many people, animals, objects, and small events. See: https://en.wikipedia.org/wiki/Wimmelbilderbuch.↩︎