June 25, 2026
Current AI-powered creativity support tools (AI-CSTs) primarily use text prompting to generate solution-oriented outputs. However, the potential value of multimodal prompting in designer-AI interaction, specifically the introduction of productive
friction to encourage iteration and reflection, has not been fully explored. To address this, we developed SketchifAI, a prototype AI-CST, and evaluated it with design students. In a mixed-methods, within-participants study, we examined
how different input modalities (text, sketch, and sketch-plus-tags) affected design students’ perceived ability to express their intent, their perception of creativity support, and their divergent thinking performance. Our preliminary findings suggest that
the sketch modality tended to enhance fluency, with inconclusive evidence for differences in variety, originality, or quality compared to text modality. Yet, paradoxically, participants showed a strong preference for text prompting. We discuss how AI tools
might be designed to reintroduce reflection-through-sketching, ensuring that designer-AI interaction supports, rather than erodes, essential design skills in students.
<ccs2012> <concept> <concept_id>10003120.10003121.10011748</concept_id> <concept_desc>Human-centered computing Empirical studies in HCI</concept_desc> <concept_significance>500</concept_significance> </concept> </ccs2012>
Imagine it is week three of a design studio — a peak time for rapid ideation. A few years ago the room would have been filled with the scent of graphite and the rhythmic scratching of pencils. Students would have struggled with the blank page, some feeling paralysed before making their first pen-stroke. Yet that friction was productive; it was the force that empowered students to translate abstract ideas into a physical form, enabling reflection.
Today, the sketchbook has been mostly replaced by screens, and the rhythmic scratch of pencils is faint, replaced by the staccato clatter of keyboards. Many students no longer try to wrestle with visualising their ideas, instead racing to construct prompts that flood their screens with highly refined, photorealistic images.
Researchers argue that sketching functions as a thinking tool and an external memory [1], [2], enabling designers to experiment, to discover new insights, and to rapidly produce a stream of ideas that keeps pace with their cognition. Thumbnail sketches [2] and Crazy 8s [3]1 are some of the popular techniques used for brainstorming and ideation. A successful ideation activity encourages divergent thinking, enabling designers to explore diverse concepts without being constrained by the need to refine details now [4], [5].
A plethora of Artificial Intelligence (AI) tools has recently emerged. Vendors claim that AI can be used as a co-creative assistant, helping designers to kick-start the creative process when feeling stuck [6], [7]. However, while some studies suggest that AI can augment human creativity and promote divergent thinking [8], [9], others indicate that AI tools could narrow creativity [10], [11], and should be appropriated with caution. The diffusion of AI tools into design workflows has obviated the struggle to make the first penstroke in favour of the challenge of: “how do I explain my idea in the right words?”. But professional designers report that text prompting is not necessarily the best way to describe their intent to AI, highlighting the risk that linguistic limitations might inhibit visual exploration [12]–[15].
In the context of the design classroom, the research discussed above raises the question: Could the use of Generative AI (GenAI) in the early stages of design jeopardise the fostering of designerly thinking? Students risk shifting from being active creators to passive curators. This could impose the steep cost of cognitive atrophy; the erosion of cognitive skills due to over-reliance and cognitive offloading on AI tools [16] . While AI might enhance efficiency in that students can visualise refined ideas at speed, such ‘frictionless’ design processes might erode the reflective thinking required to achieve true creative mastery.
The risk of cognitive atrophy is further exacerbated by the UI of current GenAI tools, which force users to express their visual intent through a text-prompt, before exposing users to high-fidelity images. This immediate leap from idea to refined aesthetic can lead to design fixation [10], defined as a “blind adherence to a set of ideas or concepts limiting the output of conceptual design” [17]. It can be difficult for users to deconstruct these polished visual outputs into the fluid, malleable concepts necessary for novel idea generation [17], [18].
Despite the rapid growth of investigation into Human-AI co-creativity, a gap remains in HCI literature regarding the tension between efficient automation and the cognitive effects of text-to-image workflows. For example, [19], [20] explored how sketching can be used to mitigate users’ difficulty in translating ideas into words; [21] incorporated stepwise textual scaffolding to specify user intent through progressive prompting; while [22] explored how annotation and scribbling can define creative intent. Collectively, these findings suggest that GenAI interfaces need to support diverse input modalities, allowing designers to adaptively switch between modalities across different stages of the process [20], [23].
While these offer valuable insights, there remains a need to empirically examine how productive friction — deliberate resistance embedded in an interaction that creates moments of reflection and clarification [24]–[27] — influences reflective ideation, and the possible gap between participant perceptions and expert assessment of design ideas. We address this gap by investigating how design students communicate their creative intent to AI tools via different interaction modalities — text input, pure sketching, and a hybrid of sketching with short descriptive tags. Through a within-participant study with nine design students (\(N=9\)), we explore how effortful sketching and the deliberate generation of low-mid fidelity outputs impact inspiration retrieval in rapid ideation. We used a purposive sample, prioritising subjective depth over broad generalisability, and employed reflexive thematic analysis [28] and Bayesian hierarchical modelling [29] to derive preliminary findings through rigorous mixed-methods in understanding how reintroducing sketching and imperfect stimuli as productive friction affects students’ perceptions of creativity support, divergent thinking, and alignment with creative intent. Our study investigated the following research questions:
RQ1: How do different input modalities (Text-only, Sketch-only, and SketchPlusTags) influence students’ perceptions of their ability to express creative intent and the relevance of the AI-generated outcomes?
RQ2: How do these input modalities influence the level of creativity support perceived by students2?
RQ3: To what extent do these modalities influence students’ divergent thinking performance during a rapid ideation task3?
We contribute to Designer-AI Interaction literature in three ways:
Empirical Contribution: We provide evidence that while the Sketch modality shows a trend towards enhanced fluency, the addition of tags (SketchPlusTags) may constrain Originality, despite serving as a clarity anchor for
intent. This reveals a paradox where students preferred text-based prompting despite the potential generative benefits of sketching.
Implications for Design: We propose dialogic interaction that prioritises critical reflection over seamlessness. This can ensure that GenAI supports designerly thinking by positioning the interface as a site for reflection rather than just an inspiration generator.
We introduce a research probe: a multimodal prompting interface (Text, Sketch, SketchPlusTags), to investigate the role of productive friction in AI-supported ideation. A key design decision implemented in this artefact is that the
Sketch condition allows no text-input from the user, requiring users to externalise their intent purely through sketching (see Section 3).
Sketching is widely practised in design and is considered essential to designerly activity, as it enables designers to rapidly experiment and generate insights [31]. Research indicates that sketching performs two primary roles in idea generation: supporting introspective inspiration searches [1], [31] and enabling feedback during the generative process. This cognitive process is best understood through Schön’s framework of reflection-in-action [32], [33]. In this view, design is not seen as a linear execution of a pre-formed idea but a “conversation with the materials.” As the designer sketches, the ambiguity of the stroke talks back, triggering new interpretations and unintended discoveries. This iterative exercise is foundational to divergent thinking [5], where fluency (the volume of ideas) serves as a precursor to variety and originality [34], [35]. By exploring a larger conceptual space through rapid, low-fidelity sketching, designers avoid premature closure and reach more innovative solutions [17], [36].
The broad uptake of digital design tools has led researchers to highlight a risk of decline in sketching ability and reflection – particularly among emerging designers, who may be more attuned to using graphics software than pencil and paper. More recently, the emergence of AI-powered design tools creates two additional risks. First, designers may find it challenging to communicate visual ideas through text prompts; designers are forced to translate visual intent into words – a process that bypasses the spatial reasoning inherent in sketching [15], [37], which remain the primary mode of interaction with image generation tools. Second, AI tools immediately produce high-fidelity visual outputs, narrowing the opportunity for designers to engage in a traditional design-thinking process with low-fidelity sketches [10].
Because current GenAI tools prioritise seamlessness, they introduce a risk of cognitive atrophy. [26] argue that productive friction can “disrupt ‘mindless’ automatic interactions, prompting moments of reflection and more ‘mindful’ interaction” [p. 1]. Furthermore, as [10] notes, the immediate exposure of high-fidelity GenAI outputs can lead to design fixation, when used in rapid ideation tasks, narrowing idea exploration prematurely and fixating designers on AI’s initial interpretation. This offloading of the ideation process to a black-box limits the designer’s opportunity to exercise reflective thinking, potentially leading to a long-term erosion of the skills required for independent ideation. Our work investigates how reintroducing sketching and imperfect GenAI images as deliberate forms of interaction friction can restore the reflective exercise essential to designerly thinking.
Currently, there is interest in whether AI can provide creativity support and facilitate meaningful collaboration in design. Early explorations, such as the Creative Sketching Partner (CSP) [38], demonstrated that AI could reduce design fixation by providing unexpected visual stimuli that disrupts a designer’s narrow focus. Since then, the field has moved toward mixed-initiative systems where the AI acts as a co-creator rather than a passive tool [39].
A growing literature explores interaction beyond prompting, such as through sketch-based AI interfaces. Such work includes, SketchAI [19], Inkspire [20] and ImaginationVellum [40] represent a shift toward allowing designers to maintain spatial control, with the latter providing an infinite canvas for progressive refinement – a process that resonates with the iterative discovery described by [32]. Similarly, Objective Portrait [41] emphasises meta-cognition by reflecting the designer’s process back to them to stimulate new divergent paths. Studies of 3D sketching [23] have shown that linguistic prompts are often insufficient for conveying complex volumes, necessitating a return to sketch modalities.
Most such research has used the approach of designing a prototype, and assessing its usability and user perceptions of creativity support [20], [21], [23], [42]. For example, Choi et al. [42] designed a GenAI pipeline that recombines sketches and texts, finding that participants perceived this approach as more supportive of creativity than a text-only baseline. In a study with 12 participants, Lin et al. [20] compared their prototype, Inkspire, with ControlNet (a neural network architecture that guides pre-trained text-to-image diffusion models), noting that Inkspire promoted inspiration and user satisfaction. Similarly, Lee et al. [23] highlighted the critical role of sketching in exploring ideas for 3D generative images.
However, these approaches primarily investigate user preferences and perceived support, often omitting objective measures of creative output. This leaves open the question of whether the tool materially impacts the design process, or whether the satisfaction reported by users merely reflects reduced effort, which we identify as a potential driver of cognitive atrophy. As discussed [26], [43], a tool that is too easy may bypass the productive friction required for deep designerly thinking. Without evaluating creative thinking performance through expert assessment, it remains unclear if these tools genuinely support the designer or simply automate the path to a high-fidelity, yet potentially fixated, result.
Our work contributes to research on AI-powered creativity support in two ways: first, by exploring how designers communicate intent through varied modalities (Text, Sketch, and SketchPlusTags); and second, by providing empirical evidence that bridges the gap between participant perceptions and expert assessment. By employing Bayesian hierarchical modelling, we move beyond simple satisfaction metrics to explore how interaction friction influences divergent thinking and creative outcomes of design students.
We created SketchifAI (see Figure 1), a prototype that supports designers by generating inspiration stimuli in response to prompts entered through three distinct interfaces: Text (baseline),
Sketch, and SketchPlusTags. We made a simple and consistent UI, with individual canvas frames tailored to support respective input modalities.
While sketch-based interfaces have been explored in prior work, such as SketchAI [19], Inkspire [20], and CreativeConnect [42], a key design decision is that neither Sketch nor SketchPlusTags condition requires a full text prompt. The back-end prompt automatically interprets the sketch outline, allowing the users to
externalise their intent purely through sketching. In the SketchPlusTags condition, users can add short descriptive tags to guide the AI output, without needing to compose a full prompt.
The Text interface (see Figure 1-1) represents existing text-to-image tools. It allows users to create and manage canvas frames via a side panel. Each frame consists of a prompt input field, a generate
button, and a display area for the resulting image. For example, a user might input: “Happy elephant mascot character in coloured pencil sketch style in a plain background.” Users could create multiple frames and keep generating images to generate
inspiration stimuli for their work.
The Sketch interface (see Figure 1-2) gives users a pen tool and drawing canvas with standard undo/redo functionality. After drawing a sketch, users can click “generate” to produce AI-interpreted variations
of their sketch. A toggle feature allows users to switch between their original sketch and the AI-generated output (‘my sketch’ vs. ‘AI sketch’). This allows for iterative refinement and the exploration of new ideas based on the user’s initial sketch
outline.
The SketchPlusTags see Figure 1-3) interface functions similarly to the sketch interface but with an additional text box for descriptive tags. Users generate images by combining visual sketches with short
text descriptors-Tags. This hybrid approach enables the AI to produce more contextually accurate stimuli.
SketchifAI was developed as a web application (Next.js/Python) designed for use with a drawing tablet and stylus; for this study, we used the Huion Inspiroy H430P. For image generation, we utilised Stable Diffusion v1.54 . To interpret user sketches we utilised ControlNet-Scribble5, which uses a neural network architecture to guide pre-trained text-to-image diffusion models using scribble images as additional input [44]. We selected this model over newer variants (such as SDXL) because it is lightweight, well-documented, and optimised for human-style scribbles. All sketches were preprocessed as \(512 \times 512\) binary maps.
To integrate productive friction and maintain consistency, we utilised a back-end prompt for the Sketch and SketchPlusTags conditions:
Sketchy outline colour pencil drawing of a mascot character with new additional details to the sketch in plain background, HD+, HQ.This prompt ensured that all AI outputs had low-to-mid fidelity. This was in line with prior research which has shown that partially-complete or imperfect visual representations can stimulate more effective idea generation than completed ones [45]–[47]. By enforcing a sketchy aesthetic, we aimed to force participants to actively interpret and build upon the AI’s suggestions rather than passively accepting a polished, final solution. The system was hosted on a Google Cloud Platform6 Linux-based virtual machine with an NVIDIA T4 GPU, achieving a generation latency of 7-10 seconds per image.
We conducted a within-participant study to understand how different input modalities to AI affect: 1) ability to express intent to AI, 2) creativity, and 3) divergent thinking performance. Participants completed one mascot-design task under each input modality. To mitigate reduce carry-over and learning effects, condition exposure was randomised and design briefs were counterbalanced across the sample (see supplementary material-A for our Latin square design). This approach isolated the impact of interface design while controlling for task-specific influences. Our independent variable–input modality–had three levels:
(Baseline condition) Text: Participants generated inspiration stimuli via text input; this mimics the current input interaction used in AI image generation.
Sketch: Participants generated inspiration stimuli via sketches.
SketchPlusTags: Participants generated inspiration stimuli via sketches and small descriptive tags.
We asked participants to create mascot characters to serve as the logo and a key visual element for an app. (see Appendix 10 - for full briefs given to the participants.) We drew inspiration for our briefs from Ward’s creature invention task [48], [49], which directed participants to imagine and create aliens that lived on a different planet, and Choi et al.’s [21] work on using mascots for design ideation tasks.
The three briefs were designed to be abstract yet equivalent in complexity. The study received human research ethics approval from the University of Melbourne.
We recruited nine design students undertaking a Bachelor of Design or Master’s degree in UX/UI, graphic, or a related design field, via university digital noticeboards, student clubs, social media, and word-of-mouth. Participants registered their interest in the study via an online form. Their responses were screened to ensure they were aged 18+ and had prior experience using AI tools, designing characters or avatars, and were able to sketch with a stylus on a drawing tablet. Eligible participants were invited via email. Participants were aged between 19 and 28 years (mean = 21, SD = 2.8). We ensured participants did not have a direct connection with the researchers involved in this study (e.g., student-teacher or supervisor-employee) to avoid performance anxiety and to maintain ethical integrity (see Appendix [appendix32A] for participant information).
Participants first completed a brief demographic questionnaire (age, gender, design experience, and prior AI-tool use). To evaluate support for creativity, we administered a post-task questionnaire that included custom items alongside the Creativity Support Index (CSI) [30] . Intent expression was measured specifically through two 7-point Likert scale questions (1 = Strongly disagree to 7 = Strongly agree) adapted from [21]:
I was able to accurately express the direction of exploration I wanted to the system.
I was able to clearly define the scope of exploration I wanted to the system.
To evaluate our prototype, we used UMUX-Lite [50], a short instrument that correlates with System Usability Scale [51].
The study sessions were conducted in person in a UX research laboratory. Upon arrival, participants were given time to thoroughly read the project information sheet, and gave informed consent to participate (see Figure 4). Each experiment lasted ~2hrs and 15mins. The first author conducted all sessions.
In the pre-study questionnaire7, participants were automatically assigned a random, unique ID. The questionnaire collected participants’ background information, including age, design experience, familiarity with AI tools, and confidence in sketching and prompting.
In the main experimental sessions, participants were first asked to watch a two-minute video tutorial about the prototype. They were then given three minutes to familiarise themselves with the prototype, during which they were encouraged to generate any images of their choice. Participants were given verbal instructions about the design briefs and allowed to ask questions to clarify any uncertainties.
Participants were asked to complete three structured rapid ideation tasks. Each experimental session was 20 minutes long, to minimise participant fatigue. Prior studies have found this to be ideal for maintaining focus [18], [52]. In each task, participants interacted with the prototype and then sketched their own ideas on paper. Hereafter we refer to these paper sketches as the participant’s “products” (see Figure 3).
Each participant then took part in a post-session questionnaire and short interview (see Figure 4). Participants took an eight-minute break between sessions. After the three design ideation sessions, we conducted a semi-structured interview, focusing on usability, creative engagement with the AI tools, and participants’ expectations for interacting with such devices in the future. The participants were then debriefed, given the opportunity to ask questions, and given a summary of the study and prototype implementation. We thanked each participant with a $50 AUD gift voucher.
We scanned all products created by participants and assigned a unique ID to each, before importing them to a Miro board8 for expert clustering and evaluations. To analyse
divergent thinking performance, experts rated four measures using standard methods of design research [10], [53], [54]. These measures
included fluency: the number of products created by each participant [10], [53], variety: coverage of the solution space during ideation [54],originality: infrequency of an idea relative to other participants’ responses [54], and quality: the usefulness and feasibility of the product rated by the experts.
Two expert raters, blind to experimental conditions, evaluated the products. They brought complementary perspectives: one was a university lecturer with 11 years’ experience in user-centred design, while the other had 8 years of industry experience in
product and service design. Their ratings exhibited almost perfect agreement (\(\kappa\) = 0.87). Clustering, ratings and all questionnaire responses were imported to R9 for statistical analysis.
Post-task questionnaires and expert ratings were analysed using Bayesian statistical methods [29], [55]. Unlike traditional significance testing, this approach allows for a nuanced interpretation of differences between conditions by estimating effect probabilities and uncertainty. Its flexibility and capacity for extensibility make it well-suited for the small sample sizes common in HCI research [55]–[57], ensuring statistically stable preliminary findings.
Following established practices [10], [57], [58], we implemented Bayesian regression via the brms package [59] in R10. We used multivariate models for questionnaire outcomes
and multilevel models for single outcomes, including participants as random effects to account for repeated measures. We employed weakly informative regularizing priors, \(Normal(0, 2)\), to ensure data-driven inference.
Model stability and convergence were confirmed via standard diagnostics: R hat (\(\hat{R} < 1.01\)) [61]
and Effective Sample Size (\(ESS > 1000\)) [59]. All our estimates fulfilled these criteria.
We report our results using posterior means, standard deviations, and 89% credible intervals, which differ from frequentist confidence intervals, consistent with McElreath’s [29] recommendations and recent HCI literature [10], [57], [58]. For hypothesis testing, we calculated Bayes factors11 and report them on the natural log scale (\(\ln(BF)\)). All hypotheses are framed directionally (e.g., Sketch > Text), based on our expectation that
sketching would outperform the text baseline. Positive \(\ln(BF)\) values indicate evidence in favour of the hypothesised direction (e.g., Sketch > Text), negative values indicate evidence in
favour of the opposing direction (e.g., Sketch < Text), and the magnitude reflects the strength of evidence in both cases. We interpret the evidence strength using Table [tab:bf], which provides \(\ln(BF)\) thresholds adapted from [62]. For clarity, we plotted Posterior Estimated Means and Highest Posterior Density (HPD) intervals back-transformed to the original scale of each variable (e.g., the 1–7 CSI scale). Full modelling reports are provided in
the supplemental material. Note that p-values are not used in Bayesian statistics, and no claims regarding “statistical significance” should be derived from our results.
We analysed qualitative data (interview-transcripts, sketches and AI-generated images) using Braun and Clarke’s six-phase reflexive thematic analysis approach (RTA) [28], [63], in constructing themes (meaning-based patterns) to report our interpretations of the data [28], [63]. Audio recordings were automatically transcribed using Microsoft Word and verified and corrected by the first author. The first author then coded the interview data with an inductive, data-driven approach where codes operated at both semantic and latent levels (semantic codes captured surface-level meanings close to participants’ language, while latent codes captured implicit meanings). Participant sketches, sketches and tags as well as corresponding AI-generated images were also used to support the coding process. These codes were grouped to develop 5 candidate themes: AI for explorations and for validation: Emerging behaviours, Different interfaces for different design thinking stages, Inspiring for some, but frustrating for others: contradicting statements about sketch-based interfaces, Preconceived expectations for what AI outcomes, and Participants’ expectations for future idea explorations. These themes were then refined through iterative discussions among all authors, resolving overlaps, removing redundancies, and interpreting insights beyond the surface-level meaning. These were further clustered into 3 overarching themes that synthesised and extended the initial themes. As RTA requires researchers to reflect on how their own experiences shape the analytical process [28], we have included our positionality statement in Appendix 8.
We present our results in two parts: descriptive trends from the Bayesian statistical analysis and qualitative insights from Reflexive Thematic Analysis.
Overall, participants created a total of 177 products, 55 in the Text condition, 68 in the Sketch condition, and 54 in the SketchPlusTags condition (see Figure 5). The variation in number of products across conditions was an emergent outcome of the ideation activity, which we explored further in the fluency analysis (see Section [sec:Fluency]).
To address RQ1, we operationalised perceived support for expressing Intent as the extent to which participants felt the input modality allowed them to define the direction and scope of their intent to
SketchifAI. We modelled responses to both questions in the same Bayesian multivariate cumulative-probit model (see Table 1).
| Sketch | SketchPlusTags | ||||
|---|---|---|---|---|---|
| 2-3 (lr)4-5 | Est.(SD) | 89% CI | Est.(SD) | 89% CI | |
| Defining the direction | \(\mathbf{-1.44\;(1.08)}\) | \(\mathbf{[-3.35,\;-0.03]}\) | \(-0.30\;(0.95)\) | \([-1.78,\;1.13]\) | |
| Defining the scope | \(\mathbf{-2.02\;(1.19)}\) | \(\mathbf{[-4.11,\;-0.43]}\) | \(-0.26\;(0.73)\) | \([-1.41,\;0.83]\) | |
Our model indicated that Sketch condition had a 95% probability of yielding lower ratings compared to Text when defining direction (\(M=-1.44\), 89%CI [-3.35, -.03]), (\(ln(BF)=-3.00\), \(BF=0.05\)), strong evidence suggesting participants found it credibly harder to express their intended direction via Sketching alone.
For defining scope, the deficit of sketching was even more pronounced. The model suggested a 98% probability that Sketch led to lower ratings relative to baseline (\(M=-2.02\), 89%CI [-4.11, -0.43]) and
(\(ln(BF)=-3.91\),\(BF=0.02\)) provided very strong evidence that sketching alone restricted expression of scope compared to the text baseline. However, we observed
SketchPlusTags condition did not differ reliably from Text baseline (\(M=-0.26\), 89%CI [-1.41, 0.83]). Notably, (\(ln(BF)= 3.24\),\(BF=25.49\)) provided strong evidence that there’s a 96% probability that SketchPlusTags reliably outperformed Sketch alone (\(M=1.76\), 89%CI [0.52, 3.14]). In summary, our
findings suggest that sketching alone created a credible hurdle for expressing intent, but this limitation was effectively mitigated by adding tags.
| Sketch | SketchPlusTags | ||||
|---|---|---|---|---|---|
| 2-3 (lr)4-5 | Est.(SD) | 89% CI | Est.(SD) | 89% CI | |
| Collaboration: | \(-0.59\;(1.37)\) | \([-2.78,\;1.54]\) | \(-0.15\;(1.47)\) | \([-2.45,\;2.20]\) | |
| Enjoyment: | \(\mathbf{-2.97\;(1.43)}\) | \(\mathbf{[-5.51,\;-1.03]}\) | \(0.39\;(0.79)\) | \([-1.61,\;0.82]\) | |
| Exploration: | \(\mathbf{-2.40\;(1.25)}\) | \(\mathbf{[-4.49,\;-0.67]}\) | \(1.05\;(0.85)\) | \([-2.44,\;0.16]\) | |
| Expressiveness: | \(-1.12\;(1.11)\) | \([-2.96,\;0.44]\) | \(0.34\;(0.81)\) | \([-0.85,\;1.64]\) | |
| Immersion: | \(0.67\;(1.10)\) | \([-0.91,\;2.50]\) | \(-1.24\;(1.44)\) | \([-3.64,\;0.82]\) | |
| Results Worth the Effort | \(0.69\;(1.03)\) | \([-0.79,\;2.41]\) | \(0.40\;(1.39)\) | \([-1.66,\;2.60]\) | |
To address RQ2, we utilised the Creativity Support Index (CSI) [30]. This instrument evaluates how effectively a tool supports the creative process across six-dimensions: Collaboration, Exploration, Expressiveness, Immersion, Enjoyment, and Results Worth Effort. We modelled these subscales using a Bayesian multivariate cumulative-probit model to compare the impact of input modality (see Table 2).
Our findings revealed that the Sketch condition resulted in reliably lower ratings for Enjoyment (\(M=-2.97\), 89%CI [-5.51, -1.03]) and Exploration (\(M=-2.40\), 89%CI
[-4.49, -0.67]) compared to the Text baseline. Our model indicated a 99% posterior probability that sketching alone reduces enjoyment, (\(ln(BF)=-4.61\), \(BF=0.01\)) suggesting
extreme evidence for this effect. Similarly, a 99% probability and (\(ln(BF)=-4.51\), \(BF=0.011\)) provided very strong evidence that sketching hindered exploration. The
SketchPlusTags condition mitigated these negative effects to varying degrees. For the Enjoyment subscale, (\(ln(BF)=-0.89\), \(BF=0.41\)) indicated only anecdotal evidence for a
difference from the Text baseline. For Exploration, moderate evidence indicated that a deficit remained for SketchPlusTags relative to Text (\(M= -1.05\), 89%CI [-2.02, -0.11]) (\(ln(BF)=-2.40\), \(BF=0.091\)). A direct comparison between the two sketching modalities confirmed that SketchPlusTags provided a reliable improvement over pure sketching (\(M=1.35\) , 89%CI [0.01, 2.83]), with a (\(ln(BF)=2.1\), \(BF=8.211\)) providing moderate support for this recovery.
For all other subscales – Collaboration, Expressiveness, Immersion, and Results Worth Effort – the credible intervals included zero, indicating that these qualities were perceived comparably across modalities. Thus, we cannot reliably claim that
Sketching offers a distinct advantage or disadvantage for these subscales of creativity support compared to Text. In summary, pure sketching created a distinct shortfall in enjoyment and exploration scores. This suggests that without
proper semantic grounding, the ambiguity of sketching introduces friction that slows down participants arriving at an AI-generated solution. However, by anchoring sketches with tags, the SketchPlusTags modality restored user enjoyment to
levels comparable to the Text baseline, though a moderate deficit in exploration remained, suggesting that tags partially mitigated the frustration that was inherent in sketching alone.
To answer RQ3, we calculated the fluency, Variety, Originality, and Quality of the products based on expert ratings to examine how inspiration stimuli generated through different modalities affect participants’ divergent thinking performance (see Table 3).
| Intercept | Sketch | SketchPlusTags | |||||
|---|---|---|---|---|---|---|---|
| 3-4(lr)5-6(lr)7-8 Metrics | Model Family | Est.(SD) | 89% CI | Est.(SD) | 89%CI | Est.(SD) | 89% CI |
| Fluency | Neg.Binomial | 1.80 (0.23) | [1.42, 2.17] | 0.20 (0.33) | [-0.33, 0.72] | -0.04 (0.33) | [-0.57, 0.49] |
| Variety | Gaussian | 0.22 (0.04) | [0.15, 0.28] | -0.02 (0.07) | [-0.13, 0.08] | 0.02 (0.09) | [-0.12, 0.16] |
| Originality | Gaussian | 0.79 (0.03) | [0.74, 0.83] | -0.02 (0.04) | [-0.08, 0.04] | -0.06 (0.04) | [-0.12, -0.00] |
| Quality | Cumulative | -* | -* | -0.24 (0.50) | [-1.06, 0.52] | 0.14 (0.37) | [-0.44, 0.73] |
To model Fluency, we considered the number of products created by each participant [10], [53]. We used a negative binomial distribution with a log link. This approach allows us to estimate the
extent to which input modality influences the mean frequency of products produced. The Bayesian model estimated Fluency of the Text condition at mean = 1.80 on the log scale (89%CI [1.43,2.19]). Relative to this
baseline, Sketch (\(M=0.21\), 89%CI [-0.33, 0.75]) and SketchPlusTags (\(M=-0.04\), 89%CI [-0.56, 0.49]) had wide credible intervals overlapping zero, indicating
little evidence for differences. Complementing this, the hypothesis tests suggested a 74% probability that Sketch has a positive effect relative to Text (\(ln(BF)= 1.03\), \(BF=2.80\)). In contrast, SketchPlusTags had only a 45% probability of being greater than zero, (\(ln(BF)= -0.20\), \(BF=0.82\)). Taken
together, the Sketch condition shows a modest tendency toward higher fluency, although the posterior predictions and hypothesis tests indicated substantial uncertainty, with no reliable differences between other conditions.
To compute the variety score we used the following method [10] that uses expert ratings:
\[\begin{align} \text{Variety} = \frac{\text{Number of clusters that a participant's products belong to - 1}}{\text{Number of clusters - 1}}\\ \end{align} \label{eq:variety}\tag{1}\]
We computed Variety using the variety score (1 ). We limited our analysis to simpler models, emphasising the estimation of direct relationships instead of pursuing full causal mediation due to our sample size constraints. We employed a Bayesian linear regression model.
We modelled Variety using a Gaussian family with varying intercepts and slopes by participant. The group-level standard deviations indicate modest variation across participants, with greater variability for the
SketchPlusTags condition (\(M=-0.02\), 89%CI [ -0.12, 0.15]) compared to Sketch (\(M=-0.02\), 89%CI [ -0.13, 0.08]) and the intercept (\(M= 0.22\), 89%CI [ 0.15, 0.28]). Posterior probabilities indicated a 34% likelihood that Sketch outperforms Text (\(ln(BF)= -0.65\), \(BF=0.52\), anecdotal evidence) and a 58% likelihood that SketchPlusTags outperforms Text, (\(ln(BF)= 0.34\), \(BF=1.41\), anecdotal
evidence), indicating neither result is sufficient to draw a firm conclusion in either direction. Overall, the data provide inconclusive evidence as to whether the modality reliably alters the breadth of the products across conditions.
We calculated originality scores using the methods suggested in prior literature on divergent thinking evaluation [10], [17], [64] The originality of an output depends on how many other participants created outputs in the same cluster.
\[\begin{align} \text{Originality} = 1-\frac{\text{Number of other participants with products in the cluster}}{\text{Number of other participants}}\\ \end{align}\]
To model Originality, we used expert ratings. However, we restricted the analysis to simpler models, focusing on estimating the direct relationships rather than attempting full causal mediation, because of our sample size. We used a Bayesian linear regression model.
We modelled the Originality scores with varying intercepts by participant. The average score in the Text condition was estimated at (\(M=0.79\), 89%CI [ 0.74, 0.83]). Relative
to this baseline, Sketch was almost unchanged, (\(M=-0.02\), 89%CI [ -0.08, 0.04]), while SketchPlusTags had a slightly lower effect (\(M=-0.06\), 89%CI [ -0.12,
0.00]), with posterior probability of 72% that Sketch reduces originality relative to Text(\(ln(BF)=-0.94\), \(BF=0.39\)) suggesting only anecdotal evidence. In
contrast, the SketchPlusTags condition exhibited a credible negative effect on originality relative to Text, with a posterior probability of 94%, (\(ln(BF)=-2.81\), \(BF=0.06\), moderate evidence). In summary, there was insufficient evidence that sketching alone altered originality, while the addition of semantic tags appeared to create a bottleneck that limited the originality of
the generated products.
To model Quality, experts evaluated whether the participants’ products met the design brief: Yes =1, Maybe =0.5, and No =0. Again, our explorations were limited to simple models because of our sample size
restrictions and we opted to explore only the direct effects. The expected value for each response is based on a cumulative probit model.
In the cumulative probit model, the three intercepts (-0.96, 1.21 and 1.57) mark thresholds on the latent scale that separate the observed response categories. Condition effects describe how responses shift relative to the Text condition. For
Sketch, the posterior effect was (\(M=-0.24\), 89%CI [ -1.06, 0.52]). For SketchPlusTags, the effect was (\(M=0.14\), (89%CI [ -0.44, 0.73]), both intervals are
wide and include zero. Posterior probabilities indicated a 30% likelihood that Sketch outperforms Text, (\(ln(BF)= -0.87\), \(BF=0.42\)) providing anecdotal
evidence. Similarly, adding tags showed only a 65% likelihood that SketchPlusTags outperforms Text(\(ln(BF)= 0.62\), \(BF=1.85\), anecdotal evidence). In
summary, the posterior estimates provided no credible evidence to suggest a reliable shift in quality of products across conditions.
We modelled UMUX-Lite \(\sim\) Condition (1 + Condition |id), in a Bayesian linear regression model. The Text condition was (\(M=4.15\) 89%CI [ 3.56, 4.72]). Sketch produced
lower scores than Text (\(M=-0.77\), (89%CI [-1.38, -0.17]), with 97% posterior probability that Sketch reduces usability relative to Text (\(ln(BF)= -3.50\), \(BF=0.03\)), strong evidence. SketchPlusTags was perceived as comparably usable to text (\(M=-0.20\), 89%CI [-0.71, 0.32]),with 74% posterior probability that
SketchPlusTags reduces usability relative to Text (\(ln(BF)= -1.05\), \(BF=0.35\)), anecdotal evidence. These results suggest that while pure sketching
introduces interaction friction, the addition of tags appeared to restore usability to near-baseline levels, though the evidence remains anecdotal.
To explain the nuances in our statistical findings, such as why participants preferred text over sketch and why the use of sketching showed a modest tendency towards improvement in Fluency, but lowered their enjoyment and expressiveness, we conducted a reflective thematic analysis of the interview data, the prompts used (text, sketches, and sketches plus tags) and the AI generated images to paint a holistic picture of this phenomenon. This helped us to identify three core themes:
Some participants entered the study with expectations that AI should be a state-of-the-art collaborator; consequently, they did not find the prototype helpful for their specific needs. For example, P01 had used a commercial web design product and noted:
“I’ve tried Lovable12 before... what I mean [by] powerful is, like it [AI] can somehow grab your idea and... come up with maybe [a] more optimal solution.” (P01 | W)
This response illustrates that some participants were more interested in bypassing the design process to arrive immediately at a production-ready outcome indicating their temptation to arrive at the final solution straight away, rather than engaging in an iterative design journey.
In addition, participants who reported low confidence in sketching found the interface “frustrating to use” (P08|w) and not “[their] cup of tea” (P09 |w). Others explained that their preference for text over sketches depended on the specific task goal:
“I tend to use text... [to] branch out into many different aspects... that helps me to broaden my perspective... when I’m first exploring a design” (P03 |W) .
This indicate that preference for an interaction modality was shaped by sketching confidence, familiarity with the tool, and the desire to follow path of least resistance, which means choosing the course of action requiring the least cognitive effort to perform a creative task [45], [49].
In addition, several participants mentioned that the lack of conversational ability in SketchifAI made the AI feel like a tool rather than a collaborator. We infer that these factors likely contributed to the low ratings in the Creativity
Support Index. Some participants expressed a preference for dialogic engagement where the AI could ask for clarification or offer its own feedback:
“I feel like if the AI could respond back to me... with its own ideas, that would be helpful as well. It’s like a two-way interaction.” (P09 | W)
The SketchPlusTags interface emerged as a successful middle ground, providing the semantic grounding that the our prototype otherwise lacked. This was illustrated by P07:
“Whereas the previous ones were... sort of a gamble... I like the consistency for the sketch plus the text. [It is] the consistency of getting what I wanted, plus extra ideas I didn’t even think about” (P07| S+T-3-B).
Alternatively, participants suggested audio as a complementary modality. They noted that “talking through” their thought process while sketching would help articulate intentions that are difficult to capture via drawing alone.
“Having an audio input would be a great way... because as you’re designing, sometimes to help voice out your ideas, they might not be really fleshed out, but you’re just sharing your thinking process with the AI” (P03| W).
This highlights participants’ expectation for fluid interfaces that allow them to move between thinking aloud their ideas, which will lead to reflection, while making it clearer for the AI to understand their ideas in a natural way.
Most participants approached the design task by noting down key ideas from the brief and identifying the overall goal. Once the problem was defined, they followed distinct interaction paths: some drew directly on the digital canvas, while others began with analogue pencil sketches on paper, redrew them with alterations on the digital canvas, and used the AI to extend their initial explorations.
For instance, P03 used the tool to find “...relevant objects or ideas for [a] certain topic”** (P03| S+T21-1-A), while P06 adopted an iterative approach: “[I] sketched some elements... and [saw] how AI goes, and then I try to add more elements and see how it works”** (P06| S-1-A). Other participants focused on semantic keywords from the brief to guide their prompts:
“My overall thinking process was to focus on the keywords... like eco-friendly, hills, and hiking... the mascot is somewhat based around that idea” (P08| T-3-B).
In contrast, some participants sought to use the AI solely for verifying existing ideas–a behaviour that deviates from the intended purpose of a divergent thinking task. These participants sought reassurance that their concepts were “correct” or viable by looking at the AI generated image as a simulator of their mental imagery. P08 explained:
“It [AI] helped me... understand what I’m doing wrong. So, I... [don’t] make the same mistake on paper and don’t use the paper again and again.” (P08| S-1-C)
Similarly, P01 noted: “I came up with an idea first [and] I drew, and then I kind of validated using [the] system” (P01| S+T- C)
These behaviours indicate a split in user intent: while some prioritised exploration by asking the AI to generate new symbols and concepts, others utilised the AI for reassurance and visualisation of pre-determined ideas. Some participants were often in a validation loop, repeatedly generating images until the AI-output correctly resembled their mental imagery. Because the AI frequently did not deliver the exact idea they wished for based on their sketch outlines, as they tried to keep sketching and generating to get a “correct” or “validated” result.
Some participants discussed that distinct interfaces supported different stages of the design process, suggesting there is no single “silver bullet.” For example, participants noted that the text modality was best suited to “help brainstorm” (P05) and for initial exploration of the problem space, using AI to expand their conceptual solution space. P06 explained:
“My favourite [modality] is the one where I just add prompts... I feel like it’s very helpful if you’re brain blocked... you just type random stuff and it will generate the things that might help you” (P06| W).
For idea expansion, participants preferred the SketchPlusTags interface, as it enabled them to maintain creative agency while allowing the AI to produce serendipitous outcomes to spark ideation.Here, creative agency refers to the degree of
control a designer retains over the ideation process and its outcomes [65], [66].
“I think it [
SketchPlusTags] allows me to build on my own, brainstorming a little more... I could take more opportunities with the tags, because sometimes... it will generate something unexpected and then I would take that into consideration” (P06| W).
Participants preferred to use the Sketch interface after initial exploration, for refining their ideas, detailed adjustments, visualisation, and concept finalisation. P06 summarised this workflow:
“It was easier for you to break through blocks when you have prompts, but if you want refinements, it’s better to have sketching options” (P06| W).
These responses indicate that participants intuitively adopted different modalities at different stages of the design process. text-prompting was preferred for early divergent exploration, where low effort input allowed rapid idea generation, while
sketching was preferred for later refinement stages, where greater commitment to a spatial representation supported more focused, evaluative thinking. The addition of tags in SketchPlusTags appeared to offer a middle ground, supporting idea
expansion while maintaining creative agency.
Through a within participants mixed-methods study, we explored how different input modalities –text-prompting, sketching, and sketching plus tags– support designers’ ability to express their intent to AI tools, their perception of AI’s support for
creativity, and their divergent thinking performance. Our quantitative analysis indicated that the Sketch condition tended to enhance ideational fluency, with no credible differences observed in variety, originality, or quality, compared with
the Text baseline, when used to generate inspiration stimuli via SketchifAI. Yet despite the apparent advantages of sketching, students exhibited a strong preference for text-prompting, presenting a paradox for implementing
sketch-based interfaces. Our qualitative analysis offers insights into possible causes of these findings. Below, we reflect on the outcomes of this study and discuss implications for design, the limitations of this work, and future directions.
Most of our participants were undergraduate design students, from an “AI-savvy” generation, a demographic accustomed to GenAI tools with text-prompt interfaces. Their experience with publicly available AI tools varied (min = 0.5, max = 2 years; mean = 1.3, SD = 0.56). Our findings suggest that design students may be becoming less inclined to spend time sketching and experimenting with ideas, preferring instead to write prompts that drive the AI straight to polished outcomes. This reflects a preference to choose a path of least resistance, a behaviour reinforced by current efficiency-driven AI tools.
We speculate our findings align with a broader shift: the Visual Thinker archetype in design education seems to be gradually displaced by an Efficiency Seeker mindset that seeks the immediate gratification of high-fidelity text-to-image generation. By offloading the conceptual heavy lifting to GenAI, students may lose the opportunity to exercise the skills vital to building their creative muscle, which, in turn, may lead to cognitive atrophy, diminishing the skills necessary to generate designerly outcomes. In a literature review, [67] argue that one methodological approach in designerly thinking is to engage in reflective practices. We suggest that the ultimate goal of AI in design education should not be to make the design process easier, but to cultivate a deeper understanding of design thinking and knowledge structures [68].
Therefore, instead of allowing students to use AI that reinforces the temptation to seek instant solutions, we propose the design of a new genre of AI tools that can serve as scaffolds for students, by introducing reflective pauses [26]. These may prompt students to stop, look, and reflect before solutions are rendered, ensuring the machine’s speed does not outpace the student’s development of critical, independent creative judgment. Our suggestions here extend the discourse of slow and reflective AI [40], [41], [69]. For example, rather than auto-correcting a rough, triangular sketch into a polished image, the AI might highlight the shape and ask the user: “You’ve used sharp triangular forms here. In character design, this typically signals ‘danger’ ‘or speed.’ Is this intended, or should we explore rounder shapes to convey friendliness?” This could potentially move the student from a Curator Mindset, where they merely select from a menu of AI-generated possibilities, to a Creator Mindset, where they actively design. Similar implementations of productive friction have been explored to promote creativity and reflection in programming education [43], [70].
Prior research suggests that exposure to high-fidelity AI images causes design fixation [10]. In contrast, [45] have demonstrated that ambiguous imagery can disrupt the “path of least resistance” and reduce fixation. Recent studies such as Inkspire [20] and Reframer [71], both of which use sketching as an interaction medium, characterise unexpected AI outputs as serendipitous. This aligns with the literature on design by analogy [72], which suggests that misinterpretations can act as creative stimuli leading to novel ideas [73].
Therefore, to encourage divergence [45], we introduced sketching and deliberately restricted SketchifAI to produce
imperfect, low-to-mid-fidelity outputs. For instance, when participant P02 sketched a snail, SketchifAI reinterpreted the spiral shell as a stylised horn and the body as a hybrid of tortoise and bird, morphing the snail into a bipedal
character (see Figure 12), a transformation with the potential to inspire mascot characters, illustrating how productive friction may prompt unexpected creative directions.
While this resistance was intentional, it created a disconnect for users unaware of the constraint. Some participants perceived this as a lack of state-of-the-art capabilities, and interpreted unexpected outputs not as creative sparks, but as “technical glitches.” This mismatch between pedagogical intent (using low-to-mid fidelity to prevent fixation) and user expectations (of immediate, high-fidelity output) likely contributed to lower enjoyment scores. Without explicit scaffolding, productive friction can be easily misconstrued as a system error. It is important to note that the rationale behind this friction (implemented via mid-fidelity outcomes) was withheld until the debrief in order to observe authentic user reactions. These preliminary findings suggest a critical caveat: intentionally making AI outputs low-fidelity (rough or messy) to induce productive friction may be less effective if the pedagogical reasoning is opaque to the user.
Therefore, we propose that embedding productive friction into the system architecture alone is insufficient; explicitly communicating the pedagogical rationale to the user may be required. Intentional friction may be more effective when transparent, and future interfaces might benefit from explicitly labelling resistance mechanisms (e.g., “Ambiguity Mode Active”) to nudge user frustration towards reflective practice.
Recent work by [23] proposed that text supports divergent exploration, while sketching supports convergent refinement. Our findings
suggest a more nuanced relationship; specifically, we observed that the Sketch condition showed a trend towards enhancing Fluency, a key aspect of divergent thinking. We interpret this to mean that when the
sketch interface did not provide an immediate, polished solution, users were denied a cognitive shortcut. Once participants began to sketch, it became cognitively more efficient to iterate visually than to stop, switch contexts, and formulate a text
prompt. Consistent with Schön’s reflection-in-action [32], the ambiguity of the medium may have compelled users to think through the
sketching, reinforcing Goldschmidt’s view of sketching as a fundamental thinking tool [1].
In addition, the SketchPlusTags modality, which mitigated negative effects on expressing intent, actually reduced the Originality of the products. We attribute this to a Semantic Anchor effect. In the Text-only condition, the AI
had the freedom to hallucinate visual details within semantic bounds. In the Sketch-only condition, it interpreted semantics within visual bounds. However, when providing both a specific visual structure (sketch) and a rigid semantic definition (tags),
participants inadvertently over-constrained the generative model (see Figure 13). It is also possible that cognitive offloading played a role: the presence of textual tags may have encouraged users to produce simpler,
less detailed sketches, relying on the text to fill-in-the-gaps, which resulted in more generic outputs.
This finding complicates Shneiderman’s ideal of “Low Threshold, High Ceiling” interfaces in creativity support tools [74] and prompts a rethink of how we design for AI-powered creativity. While the SketchPlusTags hybrid mode successfully lowered the threshold (restoring usability), the semantic anchors simultaneously lowered
the ceiling for creativity. This highlights a potential trade-off: the interface that felt the safest (SketchPlusTags) resulted in a low threshold, low ceiling environment, potentially limiting unique ideas.
Some participants treated the AI as a “black box” [75], leading to trial-and-error frustration. They expressed a desire for the AI
to seek clarification, asking, “Is this what you mean?” Without the ability to explain the intent behind a specific shape or correct the AI’s misunderstandings in real time, Sketch Interface felt like a one-way command rather than a
co-creative partnership. To address this, we propose shifting to mixed-initiative co-creative interfaces [39].
We suggest that future tools incorporate a conversational layer that goes beyond simple chat. For example, the AI could pause to query intent (e.g., “I see a circular shape. Is this a face or a wheel?”) before rendering. Furthermore, the AI could provide reasoning for its creative choices. This conversational grounding could break the silent validation loop, transforming it into a constructive dialogue that builds mutual understanding. This dialogue need not be limited to text but can leverage visual interactions, such as:
Highlighting specific regions of ambiguity.
Extending sketches to propose completions or include analogous ideas.
Ghost interactions (faint overlays) that visualise potential design directions before they are finalised.
In addition, interfaces could also support non-linear exploration. By maintaining a version history similar to Git13, tools can enable participants to branch ideas and toggle between problem-framing and solution-finding without the fear of losing previous ideas.
Finally, we propose that future AI design tools be task-adaptable, prompting users to reason about their modality choices. If an interface detects that a user repeatedly iterates on the same idea, the AI could suggest: “You seem to be refining details. Would you like to switch to text?” Such meta-cognitive prompting can help students learn to deploy modalities as distinct cognitive tools.
We acknowledge several limitations that should be considered when interpreting our findings.
First, our sample size was small (\(N=9\)) which is a major limitation for deriving stable quantitative findings. For this reason, we employed Bayesian analysis, which is particularly well-suited for small-sample studies, as it provides probabilistic estimates of effects without relying on the large-sample assumptions required by frequentist statistics. However, the posteriors carry substantial uncertainty and estimates may shift with additional data. We therefore acknowledge that the quantitative findings should be treated as preliminary, and future studies with larger samples are required. In addition, our participants were drawn from a single institution - although, one with a diverse student cohort. Our findings should be interpreted as signals of emerging behaviours in “AI-native” design students rather than as indicative of the entire population.
Second, our study utilised a rapid ideation task, and therefore can support only initial insights; we recognise longitudinal research is needed to evaluate sustained tool appropriation.
Third, we employed ControlNet, which requires semantic prompts to interpret sketch outlines – we used a back-end prompt to maintain the scope and implement productive friction. We acknowledge that AI-models with native image-recognition capabilities may function differently; thus future research should evaluate how evolving architectures impact the intentional implementation of friction in creative workflows.
Through this study, we investigated how design students communicate intent to AI via Text, Sketches, and SketchPlusTags modalities, and how they affect creativity support and divergent thinking. Our findings point
to a critical paradox: While it is intuitive to assume that design students are visual thinkers, in practice they gravitated toward text-prompts. However, our preliminary evidence suggests that the productive friction of sketching may increase Fluency,
whereas the hybrid SketchPlusTags modality appeared to create a semantic anchor limiting the Originality of generated products. Taken together, these preliminary findings suggest that the challenge for the HCI and Design communities
is not simply to maximise efficiency, but to develop AI-CSTs that promote reflection-in-action. We must explore transitioning from making black-box content generators towards co-creative partners that provoke reflective thought, ensuring that AI-driven
speed does not come at the cost of the design students’ cognition.
This research is supported by Melbourne Research Scholarship, the Diane Lemaire Scholarship and the Rowden White Scholarship offered by the University of Melbourne. We would also like to thank Oshan Wisumperuma for his assistance with the back-end
development of SketchifAI, and Jeremy Silver at the Melbourne Statistical Consulting Platform for their support.
The first author has over two years of professional experience in UX/UI design and holds a bachelor’s degree in design, alongside over three years of experience teaching UX/UI design and conducting HCI research at university level. She conducted the interviews and performed the initial coding and theme development. This background enabled her to empathise with participants’ accounts of design practice, while remaining attentive to the risk of over-identifying with participant experiences given her shared background in design. The second and third authors contributed to the iterative team discussions, theme refinement, and finalising the analysis. Both hold doctorates and have over 20 years of combined experience teaching and researching UX/UI design. The collective expertise of the authorship team facilitated critical reflection on participants’ accounts and contributed to constructing meaning from the data beyond surface-level description.
| ln(\(BF\)) | \(BF\) | Evidence category |
|---|---|---|
| \(> 4.61\) | \(> 100\) | Extreme evidence for \(H_1\) |
| \(3.40 - 4.61\) | \(30 - 100\) | Very strong evidence for \(H_1\) |
| \(2.30 - 3.40\) | \(10 - 30\) | Strong evidence for \(H_1\) |
| \(1.10 - 2.30\) | \(3 - 10\) | Moderate evidence for \(H_1\) |
| \(0 - 1.10\) | \(1 - 3\) | Anecdotal evidence for \(H_1\) |
| \(0\) | \(1\) | No evidence |
| \(-1.10 - 0\) | \(1/3 - 1\) | Anecdotal evidence for \(H_0\) |
| \(-2.30 - -1.10\) | \(1/10 - 1/3\) | Moderate evidence for \(H_0\) |
| \(-3.40 - -2.30\) | \(1/30 - 1/10\) | Strong evidence for \(H_0\) |
| \(-4.61 - -3.40\) | \(1/100 - 1/30\) | Very strong evidence for \(H_0\) |
| \(< -4.61\) | \(< 1/100\) | Extreme evidence for \(H_0\) |
Read the instructions given below and come up with as many design ideas as possible for a mascot character.
Project overview: Design a mascot character to serve as a key visual element and the logo for an app designed for older adults to maintain their social connectedness with family and loved ones who live remotely.
Purpose of the character: The image will represent the character offering comfort, guidance, and emotional connection. It should make the mascot feel more personal and approachable.
Target audience: Individuals aged 65 or older. Things to consider: The mascot will appear in the logo, loading screens, onboarding flows, and occasional UI interactions.
Be creative and come up with as many ideas as possible, and sketch on the provided papers. Remember, you can annotate the sketch if you need to explain more about your design. Also, please remember always to give a number for each sketch you draw.
Read the instructions given below and come up with as many design ideas as possible for a mascot character. Project overview: Design a mascot character to serve as a key visual element and the logo for a nature exploration and hiking app that promotes eco-friendly outdoor exploration and environmental awareness.
Purpose of the character: The mascot should embody an element of nature, a sense of adventure, and a gentle guiding presence. It will support users on trails, share conservation tips, and celebrate sustainable actions like picking up litter or choosing low-impact routes.
Target audience: Individuals aged 18-44, who are more likely to prioritise sustainability when making travel decisions.
Things to consider: The mascot will appear in the logo, loading screens, onboarding flows, and occasional UI interactions.
Be creative and come up with as many ideas as possible, and sketch on the provided papers. Remember, you can annotate the sketch if you need to explain more about your design. Also, please remember always to give a number for each sketch you draw.
Read the instructions given below and come up with as many design ideas as possible for a mascot character. Project overview: Design a mascot character to serve as a key visual element and the logo for a music creation app that contains inspirational tones from folk music traditions.
Purpose of the character: The character should reflect the app’s mission to honour cultural roots while promoting creativity and modern expression. It must feel soulful, respectful, and vibrant, bridging tradition and technology.
Target audience: Individuals those aged 16 or above, who love to create music while honouring and respecting culture and traditions.
Things to consider: The mascot will appear in the logo, loading screens, onboarding flows, and occasional UI interactions.
Be creative and come up with as many ideas as possible, and sketch on the provided papers. Remember, you can annotate the sketch if you need to explain more about your design. Also, please remember always to give a number for each sketch you draw.
| PID | Degree | Major | Yr. | AI tools used | AI Exp. | Prompt. Conf. (/10) | Sketch. Conf. (/10) | Age | Gender |
|---|---|---|---|---|---|---|---|---|---|
| P01 | B.Des | UX Des. & Comp. | 3 | ChatGPT, Figma AI, Perplexity, Google Gemini | 1 | 7 | 6 | 22 | M |
| P02 | B.Des | UX Des. & Perform. Des. | 1 | ChatGPT, CanvaAI, AdobeAI | 1 | 7 | 8 | 19 | F |
| P03 | B.Des | UX Des. | 1 | Cursor, MagicAnimator, ChatGPT | 1 | 6 | 8 | 20 | F |
| P04 | M.Des | Arch. & Graphics | 2 | ChatGPT, Midjourney, Finch | 1 | 6 | 8 | 28 | M |
| P05 | B.Des | UX Des. & Comp. | 1 | Midjourney, Canva, ChatGPT, Adobe AI, Grok | 2 | 7 | 8 | 19 | F |
| P06 | B.Des | UX Des. | 2 | ChatGPT, Framer, Figma, Adobe, Canva | 2 | 7 | 5 | 19 | F |
| P07 | B.Des | UX Des. | 1 | ChatGPT, Dall-E | 2 | 8 | 5 | 20 | M |
| P08 | B.Des | Humanit. & UX Des. | 3 | Mid journey, Sora, ChatGPT, Adobe Firefly | 0.5 | 5 | 4 | 21 | M |
| P09 | B.Des | UX Des. | 2 | ChatGPT, Canva AI | 1 | 6 | 4 | 21 | F |
https://designsprintkit.withgoogle.com/methodology/phase3-sketch/crazy-8s↩︎
Measured using expert ratings for fluency, variety, originality, and quality.↩︎
https://huggingface.co/stable-diffusion-v1-5/stable-diffusion-v1-5↩︎
https://cloud.google.com↩︎
(deployed via Qualtrics https://www.qualtrics.com)↩︎
https://www.r-project.org↩︎
An interface to the Stan probabilistic programming language [60]↩︎
Also known as Evidence Ratios.↩︎
https://git-scm.com↩︎