July 16, 2026
Data visualizations are increasingly accessed in all facets of everyday life: business, news, and education. For people who are blind or have low vision (BLV), accessing these data visualizations poses significant barriers. Existing approaches, like textual descriptions [1], [2], sonification [3], [4], screen readers [5], and static tactile graphics, each address aspects of accessible data visualization, but none provide the interactive, multimodal experience that lets BLV users independently explore, query, and verify data during sensemaking.
Refreshable tactile displays (RTDs) render tactile graphics dynamically on pin-based displays, offering a more interactive alternative to static printed graphics, which require a new graphic to be produced for each change. However, RTDs in isolation have limitations: they are currently low resolution, which constrains labeling, and tactile graphics are difficult to comprehend without descriptions [6]. Combining an RTD with a conversational AI agent could address many barriers BLV people face when analyzing data [7]: the agent provides verbal context and analytical support, while the rendered chart provides interactive spatial grounding and independent tactile access to the data.
In our previous work, we conducted a Wizard-of-Oz (WOz) study revealing that BLV participants strongly preferred multimodal interaction combining touch with conversational queries [8]. However, it remains an open question how these modalities should work together in a fully functional implementation. The wizard simulated an idealized system, seamlessly bridging touch and speech in real time, and abstracting away design challenges that a real system must solve, such as how to communicate feedback to users whose fingers must actively search for it, what role the agent should play relative to what touch already provides, and how to coordinate touch and speech to support chart orientation, exploration, sensemaking, and deictic interaction.
To address these questions, we co-designed Graphywith three BLV co-designers. Graphyis a conversational tactile data interface (CTDI)—a system combining an RTD with an LLM-powered conversational agent—that enables BLV users to explore data visualizations through touch and speech. Through four co-design workshops, the work presents an initial exploration of the design space for CTDIs. Our contributions are:
**Graphy—the first conversational tactile data interface, combining a multi-line RTD with an LLM-powered agent, created with three BLV co-designers. Graphyextends our prior WOz work into an instantiated system, revealing challenges WOz abstracted away and capabilities that emerged through co-design.
Design knowledge for CTDIs. Key findings include: a layered presentation that scaffolds chart exploration through interactive layers; a feedback grammar that distinguishes user- and agent-initiated tactile highlights; and a sequential interaction pattern—select, confirm, ask, verify—grounding each step in the data. We distill these into design recommendations.
Accessible data visualization has attracted a growing research interest [9]–[12]. Printed tactile graphics have long been the traditional approach to making data visualizations accessible to BLV people, and are recommended by transcription guidelines for the production of accessible charts, diagrams, and maps, where spatial relationships are important [6], [13]. However, they are time-consuming to produce and costly to update, as any change requires producing a new graphic, making them unsuitable for interactive data exploration. Additionally, He et al. [14] designed 3D printed tactile charts but found that existing design guidelines were inadequate. Tools like Tactile Vega-Lite [15] exist to address production of tactile charts, but the resulting graphics remain static.
Researchers have explored alternative formats. This includes descriptions of charts, where researchers have investigated what makes an effective description [1], [2] and how to automatically generate them [16]–[18]. However, these do not support spatial exploration. Another approach is using sonification to convey data, which can be used alongside or as a replacement for data visualizations [3]. Recent work has explored combining sonification with speech [4], [19], natural language queries [20], and tactile graphics [21]. Other systems support BLV data analysis through screen reader navigation of chart and data structure [5], [22], [23], and authoring environments that combine visualization, sonification, and textual representations [24]. However, none of these approaches offer interactive tactile exploration of data visualization through touch.
Refreshable tactile displays (RTDs), pin-based devices that render tactile graphics, represent a promising option for supporting BLV users in consuming data visualizations. Consisting of pins raised and lowered using electromechanical actuators [25], RTDs offer an advantage over traditional printed tactile graphics: they can display new graphics in seconds. More affordable devices have recently entered the market, including the Dot Pad [26] (2,400 pins, 60\(\times\)40, under $5,000 USD), Monarch [27] (3,840 pins, 96\(\times\)40), and the Graphiti [28] (2,400 pins, 60\(\times\)40). Research prototypes such as MagnePins [29] are also exploring affordable, open-source alternatives. While RTDs remain expensive up-front, there is significant value in their interactivity and refreshability.
Research using RTDs has mostly focused on graphics spanning art [30], illustrations [31]–[33], accessible maps for navigation [34]–[36], and automatic translation of graphics using AI [37]. Recent work has begun exploring more interactive and dynamic content on RTDs, including animations [38] and sports visualizations [39], [40]. Jiao et al. [41] devised tactile data comics, a step-by-step presentation of tactile graphics with verbal narration on RTDs for educational content.
Relatively little work has explored how RTDs can support data visualization. Elavsky [23] co-designed tools for exploring set diagrams, and Holloway et al. [42] surveyed stakeholder perspectives on how RTDs could improve access to data visualization. Seo et al. [43], [44] created MAIDR, which integrates Braille Unicode characters in single-line RTDs for data visualization, though it lacks support for 2D charts. Tactually encoding charts on RTDs remains an open challenge: visual encoding standards do not directly translate to the tactile dimension [45]–[47], and whether standards for printed tactile charts [14], [15] transfer to RTDs is an open question.
Chundury et al. [7] compared an RTD with two other accessible data systems: audio-only (screen reader) and audio-tactile (sonification), finding that each modality had limitations in isolation, and recommended combining conversational speech with tactile representation. In our previous work [8], we used Wizard-of-Oz (WOz) to explore how BLV users interact with data visualization using a system that combined RTDs with a conversational agent, finding preference for multimodal interaction combining speech with touch. The present work builds on these findings by co-designing a fully instantiated system.
Advances in Large Language Models (LLMs) have enabled conversational agents that can answer questions about data visualization [48], [49]. VizAbility [50] allows users to query visual data trends using natural language, and Kim et al. [51] mapped queries from BLV users exploring charts to analytical task taxonomies. Choe et al. [52] employed an LLM with visual charts; sighted users with higher data literacy engaged more with the charts and utilized the LLM as a specialist, whereas those with lower literacy relied more on the agent. Seo et al. [44] added an LLM to MAIDR [43], and found that BLV users developed strategies to verify LLM-mediated responses. Combining speech with tactile representations has been explored with interactive 3D printed models [53]–[55] and deictic querying on tactile maps [56], and our previous WOz work [8] began to explore this for data visualization on RTDs. However, no system exists that tightly integrates multimodal touch input, speech, and tactile feedback for interactive data exploration.
Graphyis a CTDI that integrates touch, speech, and tactile feedback for accessible data visualization. To our knowledge, it is the first to combine a multi-line RTD with an LLM-powered conversational agent, supporting line charts (single- and multi-series), bar and stacked bar charts, and scatterplots. Graphyis characterized by:
Charts are split into interactive layers, each narrated by the agent and rendered on the RTD, so users can explore and build their understanding through touch and speech together: one component at a time. Users can skip layers and navigate back and forth at their own pace.
Users explore charts freely by touch, selecting and inspecting data points through double-tap gestures or stepping through values using the RTD’s buttons. Each selection is grounded across three feedback channels: an audio label, a Braille label, and a tactile highlight that marks the point on the RTD.
Users can query the agent for trends, calculations, and comparisons, fusing touch and speech by referencing gestured data points directly in their requests (e.g.,selecting two points and asking, “What is the difference between these two?”). Responses from the agent are segmented into navigable one-sentence chunks, each synchronized with animated tactile highlights that draw attention to the referenced data on the RTD, allowing users to pace the agent’s response alongside their own exploration.
Users can isolate individual data series through voice commands to the agent, reducing chart complexity for more targeted exploration and interaction.
Together, these capabilities support a fluid interplay of touch and speech. The following scenario, drawn from WS4 (§ 4.6), illustrates these in practice, combining tactile exploration with conversational engagement to explore a data visualization. The interaction flow is shown in Fig. 2, and the manner in which they were co-designed is detailed in § 4:
D1 holds the push-to-talk button and asks “Load the computer component data line chart.” A multi-series line chart loads on the RTD. The full chart is visible and the layered presentation begins, announcing “Memory, Storage, and GPU price multi-series line chart, February 2025 to February 2026.”
They begin tracing the line chart, then advance through the x-axis, y-axis, and series one (Memory) presentation layers using the pan right button. Each layer is rendered in isolation on the RTD, and D1 explores by touch as Graphynarrates each component, e.g.,“Plotted using the plus symbol, Memory prices rose sharply, with two major surges across the period.”
In the Memory layer, D1 double-taps the first data point with their left index finger; Graphyproduces a static bounding box around the value and announces “February 2025, Memory, $190.” They confirm the selection by touching the box, and proceed to double-tap the last data point with their right index finger; accordingly, a bounding box is rendered and an audio label played.
D1 asks “What was the average price of Memory between these points?”; Graphyresponds with synchronized speech and animated highlighting on the RTD: “Your left hand selected Feb 2025, where the price of Memory was $190, and your right hand selected Feb 2026, where it was $1100. The average price of Memory between the periods was $503.” D1 traces the Memory line between the points to verify the agent’s response.
D1 advances through the series two (Storage) and three (GPU) layers using the pan right button, reaching the summary layer with the full chart. Graphyannounces “Storage rose steadily, GPU stayed flat, and Memory surged dramatically. Memory intersects with Storage at $360 in Aug 2025, and meets GPU at $850 in Dec 2025, before pulling sharply away from both.”
D1 asks “Filter everything apart from Storage.” The other series are removed. D1 steps through Storage’s data using the buttons, each confirmed with audio and a tactile highlight. They ask “What is the trend of Storage?”; Graphyresponds “Storage prices showed steady growth from $330 in Feb 2025 to $600 in Feb 2026...” D1 traces the Storage line to verify the agent’s response.
Graphyruns on a host computer that drives a Dot Pad RTD (the most affordable multi-line RTD available) over USB serial, and tracks a user’s index fingers via an Ultraleap Leap Motion Controller mounted overhead. Speech interaction is invoked by a Picovoice wake word or a push-to-talk button; Google Cloud Speech-to-Text and Text-to-Speech handle transcription and synthesis. Charts are rendered using our Vega-Lite-to-RTD renderer. The agent, built using ChatGPT and LangChain, fuses touch context with queries and answers questions by running pandas code over a chart’s dataset, such that values are computed programmatically rather than produced by the LLM. Full architectural detail is in [57], and open-source code is available.1
We co-designed Graphywith three BLV co-designers. Our aim was to generate design knowledge and deliver a fully functioning implementation through iterative co-design, not to formally evaluate the system. Following Thompson et al.’s [22] ‘parallel co-design’, we conducted one-on-one rather than group sessions, obtaining unique design insights and personal preferences. Across four rounds of workshops (twelve individual sessions over eight months), we prioritized longitudinal engagement over single-session studies with more participants, consistent with co-design practice in accessible visualization [5], [22], [23], [43], [58].
We recruited three co-designers2 from our lab-managed participant pool, each with prior design experience from unrelated research projects. The project was approved by our institutional ethics committee (ID: 37871), and all co-designers provided informed consent. Their tactile-graphics experience ranged from some (D3) to substantial (D1, D2), and all had used RTDs, voice assistants (Alexa, Google Assistant, Siri), and the LLM-based agent ChatGPT (Tab. 1). All were comfortable working with data for basic numerical tasks; D1 and D2 were confident in performing more complex tasks (e.g.,statistical analysis and interpreting trends), whereas D3 reported limited confidence beyond interpreting line charts and was neutral on more advanced data analysis.
| D# | Age | Blindness | Onset | TG Experience | RTD Experience |
|---|---|---|---|---|---|
| D1 | 24 | Legally Blind | Congenital | Substantial | Some |
| D2 | 27 | Totally Blind | Congenital | Substantial | Substantial |
| D3 | 55 | Totally Blind | Congenital | Some | Some |
5pt
Co-designers participated in all four workshops and were compensated after each session with a gift card, pro-rated at $50 AUD per hour. Successive rounds were spaced two months (WS1–WS2), three months (WS2–WS3), and one week (WS3–WS4) apart. These intervals reflected the time required to discuss, implement and stabilize design changes between rounds. Datasets were drawn from real-world sources, and included annual rainfall, interest rates, quarterly product sales and vehicle efficiency data (Fig. 3). Across workshops, co-designers were asked to use the ‘think-aloud’ method [60], verbally describing intended interactions, expectations, and reasoning.
We used design probes to ground design discussions in interactive artifacts rather than abstract feature descriptions. Probes included both researcher-proposed designs informed by our prior work [8] and alternatives proposed by co-designers, organized around five areas:
Orientation & Overview: How Graphyintroduces and orients users to new charts.
Chart Tactile Encoding: How data is represented tactually across chart types on the RTD.
Navigation, Selection & Interaction: How users want to navigate, select, and interact with the data.
Tactile Feedback: How the system provides tactile feedback to confirm user actions and draw attention to agent-referenced elements on the RTD.
Agent Interaction: How users engage with Graphy’s conversational agent and the role it plays.
These probes were integrated into workshops WS1–WS3, each of which consisted of three parts:
(Re)-Familiarization & Changes: In WS1, co-designers were introduced to the system’s functionality and asked to try each interaction. In WS2, they were re-introduced to the system and shown all changes made since WS1. In WS3, co-designers were instead asked to recall how each interaction worked from memory, with researchers filling in gaps as needed, before being shown the changes made since WS2.
Hands-On Activity: Co-designers explored datasets using Graphy, with each round introducing new chart types: single-series line charts (WS1), bar charts (WS2), and multi-series line charts, stacked bar charts, and scatterplots (WS3). Activities progressed from independent exploration (WS1) to hands-on engagement structured around design probes (WS2–WS3).
Design Feedback & Ideation: Open discussions, forming the bulk of each session, used design probes to prompt reflection across tactile, conversational, and multimodal aspects of Graphy. Probes were initially based on design challenges identified by the research team (WS1) and increasingly shaped by co-designer feedback (WS2–WS3).
WS4 followed a different structure, aiming to observe how memorable the co-designed system was through independent use with co-designers’ own data. Co-designers were asked to form a summary of their data that they could share with a friend. Afterward, they completed a NASA-TLX workload assessment—used as a reflective prompt rather than an evaluative measure—and a question assessing perceived naturalness, followed by a semi-structured reflection on how the system had evolved throughout the co-design process.
All sessions were video and audio-recorded and transcribed. WS1 and WS2 were attended by two researchers, and WS3 and WS4 by one. After each round, the attending researchers reviewed the transcripts, categorizing comments into five areas: new feature requests, system bugs, system improvements, and positive and negative experiences. Proposed changes were discussed with the broader research team and implemented before the next workshop round.
Our findings (§ 4) were derived from a design synthesis across workshop rounds, tracing recurring patterns and changes in preferences and workflows over time. For WS4, we also report workload and naturalness ratings, along with co-designers’ reflections.
The co-design process began with an initial prototype (v1) informed by our WOz findings [8]. It supported the display of single-series line charts with basic touch and speech interaction, including: an LLM-generated audio overview upon chart loading, double-tap gesture selection with audio and Braille labeling, free touch during speech for deictic references to the agent, wake-word agent activation, and single-pin blinking for highlights (selected and agent-referenced values). The agent supported data analysis backed by Python data libraries, including trend analysis, comparisons, and statistical calculations.
In this section, we trace how Graphytransformed from this v1 prototype into the system described in § 2. In doing so, we expand the design space for CTDIs. We present findings across five design areas, followed by observations from free-form use and retrospective reflections. Within each area, we describe how co-designers shaped the design through feedback and ideation across the workshops (Tab. [tab:features]).
Sighted users can orient themselves in a new chart by instantly surveying its shape, peaks, and trends. In contrast, when BLV users read tactile charts, touch exploration is a sequential experience: they move their hands across the graphic to incrementally build a mental model, often focusing on local features before understanding the whole. The v1 prototype attempted to support this by automatically presenting an overview on chart load: the chart title, x-axis range, y-axis range, and a brief interpretive summary of the data narrated as a single continuous description, akin to alt text or a transcriber’s note. While exploring the chart through touch simultaneously, co-designers found it difficult to follow due to its length, and often replayed it.
D1 proposed replacing the single narration with a layered, sequential introduction of chart components (WS1). We implemented this: each component – title, x-axis, y-axis, data, and interpretive summary – was introduced one at a time, rendered cumulatively on the RTD, with button press to advance (WS2). D3 described this as follows: “It’s one thing at a time, and you only move on to the next complication when you’re comfortable\(\ldots\) it’s a scaffolding thing, we just start gently and then build up.”
The v1 overview supported touch exploration but not interaction; co-designers insisted on being able to select data points and query the agent during the overview, to support active sensemaking rather than passive listening (WS2). D3 spoke of the overview as replacing the role of a sighted guide: “The [need for a] human is replaced by what it [the system] says.” Co-designers also agreed that experienced users could skip the overview (D2: “If it was a data set I had already looked at a lot\(\ldots\) I would want an easy way to skip the thing”). D2 distinguished between structural and interpretive content, wanting to know “the domain and the range and what the symbols are,” but not the system’s interpretive summary: “I’ll interpret it.” Despite this, the included components and their ordering (title \(\rightarrow\) x-axis \(\rightarrow\) y-axis \(\rightarrow\) data \(\rightarrow\) summary) were consistent across co-designers.
As more complex chart types were introduced, e.g.,multi-series charts (WS3), the overview became more valuable (D1: “[It was] more useful now than it was with a single line”), but multiple overlapping series make cumulative layering tactually cluttered. The co-designers suggested presenting components in isolation: showing the x-axis alone, then the y-axis alone, then each data series individually (with the axes for spatial context) to improve focus (Fig. 1). As D2 explained, “Show only the first one [series], then show only [the second]\(\ldots\) you build the mental model of each line in isolation and then you put that together”. The co-designers felt this approach could apply to all chart types and complexity levels. The summary layer returned the complete chart with all components visible, allowing co-designers to explore the full chart while listening to the summary.
When sighted users read a chart, they leverage multiple visual channels (e.g.,position, shape, size, color) simultaneously. On the majority of RTDs, tactile output is binary — each pin is either raised or lowered — and constrained by low resolution (60\(\times\)40), requiring encoding decisions that balance tactile readability against clutter. The v1 prototype rendered simple line charts using single raised pins for data points. With densely sampled data, co-designers could trace without explicit connecting lines, but noted that sparser data would require guessing the direction of change between points. The prototype also lacked symbols to distinguish series or encodings for other chart types.
Connecting lines were introduced between data points (WS2). All co-designers responded positively, with D2 describing them as “filling in the truth\(\ldots\) you don’t need to guess the direction [of the data].” However, single-pin markers blended into the line (something we did not anticipate), making it difficult to distinguish data points from the line itself. We introduced point symbols to mark data points, and tested four: arrow down, arrow up, cross, and the plus symbol (Fig. 5). Co-designers drew on their experience with printed tactile graphics, suggesting familiar symbols (e.g.,circles and 2\(\times\)2 squares), but these proved harder to distinguish on the RTD. Co-designers ranked the plus as the most readable, with the cross ranked second. The arrows were easy to locate through touch, but ranked last. Co-designers suggested a one-pin clearance gap between symbols and connecting lines to enhance readability (Fig. 6). D3 said, “You can [now] determine the symbols. They’re nice and separate from everything else,” while D1 preferred an adjustable setting.
As multi-series line charts were introduced, new encoding challenges emerged around overlapping and adjacent lines (WS3). For intersections, D3 suggested a 3\(\times\)3 square symbol. For adjacent lines, line thickness was tested, but co-designers were divided: D2 felt that thicker lines made “things way more cluttered”, while D3 described them as “100 times better.” The difficulty of reading more complex charts motivated D1 to suggest filtering (§ 4.3).
Bar charts were introduced in WS2, and stacked bar charts in WS3. D3 had never encountered a bar chart before the co-design sessions, initially reacting with dislike (“It’s never the way that I [have] ever had data represented for me\(\ldots\) yuck”). By the end of WS3, they were independently reading and comparing stacked segments within a single session. For stacked bar charts, co-designers emphasized that fill textures and clear segment delineation were essential. We explored five textures adapted from tactile graphic guidelines (Fig. 7). One (horizontal bars) was rejected as it could not render in segments smaller than three rows; the remaining four — solid infill, vertical bars, a checkerboard pattern, and a hollow perimeter — were all distinguishable by touch without being named. Segments were separated by a one-pin gap with a solid boundary line (D1: “With textures and a gap you’ve got two kind of tactile indicators”), and adjacent bars by a two-pin gap. Co-designers could clearly feel where one bar ended and the next began, but felt it should be configurable. When textures were replaced with a solid fill during a design probe, readability dropped sharply (D1: “Way less readable, actually not even readable at all”).
Co-designers initially had difficulty interpreting scatterplots, needing to traverse the layered overview more than once to build a conceptual understanding (WS3). They raised concerns about symbol overlap: D2 felt that slight overlap (1 pin) was manageable from a readability perspective, but greater overlap would require jittering or some form of aggregation, while D1 suggested reusing the same intersection symbol from multi-series line charts.
With a tactile chart rendered on an RTD, BLV users can only perceive what is directly under their fingertips: touch is sequential and local, so reaching a data point means first locating it—and risking loss of spatial context along the way. The v1 prototype supported double-tap gestures for selecting data points and retrieving audio labels, and deictic references to the agent through gestures or free touch during speech.
The double-tap gesture for selecting data points was well received, with co-designers noting its familiarity from mobile screen readers (WS1). As finger tracker reliability improved across sessions (tuned to co-designers’ touch-reading patterns), gestural input became increasingly valued; by WS3, D1 rated it as Graphy’s “most important feature,” noting, “So cool, this actually [makes me feel] like a child on Christmas Day.” Initially, selecting a data point triggered audio and Braille labels, but co-designers wanted a tactile cue to confirm correct selection, leading to the introduction of tactile highlighting (WS2, §4.4). All three channels would be activated simultaneously: audio for immediate confirmation, Braille for verification, and a tactile highlight for spatial confirmation.
Co-designers shaped how selection and querying should be conducted, preferring a sequential flow: exploring by touch, selecting values using gestures, confirming them via an audio or Braille label, and then asking the agent questions (WS1). They wanted to verify they had selected the intended point before querying the agent; without that confirmation, they risked asking about the wrong data point. Once a point of interest was found through touch exploration, each step confirmed the previous: select, confirm, ask. This contrasted with our WOz study [8], where participants preferred free touch during speech, which offered no confirmation step between selection and query; we subsequently focused on gestural input.
D1 suggested that the RTD’s physical buttons could be used to step through data points, noting that it would “allow [them] to get to the next value without having to find it” (WS2). Implemented in a probe (WS3), stepping is analogous to a screen reader, allowing users to move through values one at a time. Each stepped value is confirmed with the same triple feedback as gesture selection: audio and Braille labels, and a tactile highlight. D2 used stepping to improve selection precision and for recovery when gestures overshot, while D3 used the narrated values to guide directional touch exploration, predicting the shape of the data before tracing it by hand, calling it “voice-guided exploration.”
For multi-series and stacked bar charts, stepping order proved task-dependent (D3: “Am I wanting to look at software only [same segment across bars]? Or quarter per quarter [all segments within a bar]?”). For scatterplots, D2 noted that users must choose which dimension to traverse: “You kind of have to select the [dimensions] to fix, which means you’re traversing one.”
As chart complexity increased, co-designers found it difficult to parse overlapping lines for interaction (D2: “More cluttered\(\ldots\) this feels very noisy to me”). D1 proposed filtering to isolate individual series, describing it as “critical to reading charts, particularly more complicated plots.” Filtering was seen as a way to reduce tactile clutter and sharpen focus for touch reading and interaction by narrowing the data points available for gestures and stepping.
Designing effective tactile feedback on RTDs is challenging: feedback must be easy to locate by touch and meaningful. The v1 prototype used single-pin blinking, at an interval similar to a blinking cursor, for confirming gesture selections and agent-referenced highlights. However, co-designers found this feedback too subtle, and often missed it; D1 described the v1 blinking as “easy to miss unless you were already touching that area.” This problem was compounded by a hardware limitation: the electromagnetic actuators of the Dot Pad (and similar RTDs) lack sufficient force to raise pins against finger pressure, meaning pins do not reliably actuate when being touched.
We suggested a static bounding box as an alternative way of drawing attention (WS2), which co-designers found easy to distinguish and locate. D2 described their approach as “You familiarize with the pattern [shape], and then scan to find it.” A crosshair pattern was also suggested (D2: “X marks the spot”), but was too large and similar to the cross symbol. Co-designers found that the bounding box also served as a spatial anchor, allowing them to return to a previous point after removing their hands (D3: “I want to know where I am\(\ldots\) then I go across the chart [exploring] and find [return to] the box”). Each finger can maintain one highlight at a time, allowing co-designers to mark and compare two points simultaneously. D3 also suggested an isolation mode, which would temporarily lower and hide nearby chart elements to increase contrast around highlighted points (WS2).
Both user- and agent-initiated highlights initially used the same static bounding box (WS3). This became problematic because gesture selection, stepping, and agent references rendered the same pattern: co-designers had to infer whether a highlight was in response to their own action or something the agent was drawing attention to. All co-designers arrived at the same solution: static highlights work for user selections because they already know where they selected, whereas agent-initiated highlights require animation to draw attention. Co-designers explored animating the data point symbols, but found them harder to interpret than the bounding box: they had to first locate the animation, then determine which symbol was animating and what it represented, adding cognitive overhead.
D1 refined this into a feedback grammar: gesture selection = static box; stepping = animation that settles to static; agent reference = animated box. Co-designers transferred this grammar to bar charts (D3: “For consistency, let’s still use animation for when we ask questions”), though they proposed using the bar’s infill textures rather than a bounding box: gesture selection would remove infill, stepping would animate and then settle, and agent references would continuously animate. Co-designers agreed that all highlights should persist until dismissed.
Combining a conversational agent with an RTD raises the question of how it will be used when users can already feel the data. The v1 agent was invoked via a wake word (“Hey Graphy”), with responses generated without constraints on format or length. Co-designers found the wake word unreliable (WS1), and indicated it was difficult to determine when the agent was listening and when it was processing their request.
All three co-designers suggested a push-to-talk button (WS1). Although this requires briefly lifting a hand off the chart, which can cause a momentary loss of spatial context, they felt that deterministic control outweighed this cost (D2: “Walkie-talkie with a robot”). The wake word was retained for situations requiring both hands on the chart, but seldom used. Co-designers also replaced v1’s uniform start/stop listening earcon tone (based on Siri) with distinct ascending and descending tones (based on Google Assistant): “Low to high, you speak, high to low closes the loop” (D3) — valued as confirmation that their input had registered, even with push-to-talk.
D1 expected the agent to support automatic follow-ups when a query lacked specificity, prompting the user for clarification by using the same earcon cues rather than requiring re-invocation (“That’s cue to give more information\(\ldots\) that’s what LLMs do”), noting this would be important for complex analytical queries involving multiple steps.
Without being told what the agent could or could not do, co-designers directed queries to the agent that touch alone could not resolve, including: calculations and comparisons (e.g.,“Is this point lower than the average?”), trend analysis (e.g.,“Describe the overall trend”), visualization literacy (e.g.,“How does a scatterplot work?”), and operations (e.g.,“Filter everything”). The complexity of queries grew across workshops, from basic value retrieval (WS1) to queries about trends and queries with deictic references (WS2), and requests for comparative analysis (WS3). D2 also expected the agent to explain its capabilities to improve discoverability: “What does this button do\(\ldots\) what can I ask you?”
Co-designers noted the chart already provides spatial context through touch. Hence, they wanted concise, answer-first responses (e.g.,“$120 in 2010” rather than “The stock price in 2010 was $120”), “Not all the fluff” (D3). However, overly terse responses (e.g.,“$120”) were rejected, as they lacked context to verify correctness (D3: “What if it misheard me? I think you need to include my parameters in your answer”). Response length was seen as task-dependent, with D2 likening length to “skim-reading versus depth-reading.” They also insisted on consistent structure (D1: “The information has to be in the same format”), which requires precise prompting of the LLM.
In deictic queries, users fuse touch and speech by referencing selected data points directly in their query. Co-designers wanted the agent to acknowledge the touch context before answering, so they could confirm that the correct point(s) were selected. Without this, D1 noted “I just assume [that the agent] has picked a random spot, and I don’t know where.” D3 similarly valued this as confirmation the agent had correctly interpreted their interaction, preferring the agent to describe hand positions explicitly (e.g.,“Your left hand is touching 2010, where the average was X, and your right is touching 2015\(\ldots\)”).
Co-designers found it difficult to follow long agent responses when engaging in touch exploration, with D3 noting that “longer responses can be overwhelming”, as audio and touch compete for attention. D1 and D3 suggested breaking responses into sentence-sized chunks navigable via the RTD’s buttons, allowing users to pace responses to match their exploration speed (WS1). Each segment was synchronized with corresponding tactile highlights, linking the agent’s spoken response to spatial locations on the chart. D1 described this as a “game changer for keeping track” (WS2).
In WS4, co-designers explored a multi-series line chart rendered from data that they had suggested at the end of WS3: computer component pricing (D1), computer storage format pricing (D2), and census language data (D3). Each chart contained three series. With no familiarization component and WS3 sessions held one week prior, we wanted to assess how memorable and intuitive the system was. Tasked with forming a summary of the data that they could share with a friend or colleague, co-designers were free to use Graphyhowever they chose. The free-form activity lasted between 28 and 39 minutes (average 32).
All co-designers followed a similar interaction pattern centered on the layered overview: they navigated back and forth between layers throughout their exploration, selecting and inspecting values via gestures and stepping, querying the agent for what touch could not resolve, and building their understanding of the data within and across layers (Tab. ¿tbl:tab:ws4?). D2, who had initially described the layered overview as “gimmicky” (WS2), had become its strongest advocate by WS4.
There was a clear division of labor: touch was the primary sensemaking channel, gesture/button selection supported precision inspection, and the agent was used for targeted queries such as data retrieval (D3: “How many Mandarin speakers were there in 1991?”), trend analysis (D1: “What is the trend of Memory in this time period?”), and cross-series comparison (D2: “When is the cost of all three storage media projected to be equal?”). Filtering was performed through speech (D1: “Filter everything apart from Memory”), functioning as a focusing strategy to reduce chart complexity and allow targeted exploration; filtering commands were mediated by the researcher behind the scenes.
However, there were individual differences between co-designers. D2 was the most Braille-reliant, reading labels after nearly every selection to verify the agent’s responses; D3 was the most touch-dominant, with deliberate, repeated tracing and comparing along and across series; and D1 made the most use of filtering, progressively isolating individual series to increase focus, before comparing across them.
All the co-designers were able to produce substantive summaries of their dataset that went beyond the system’s interpretive summary. D1 identified that the price of Memory surged from cheapest to most expensive, noting that “in October something happened to Memory where it just went ballistic, I wonder what it was\(\ldots\) data centers?”; D2 described SSD and Hard Disk reaching price parity and asked the agent to project when/if all formats would converge; and D3 traced three languages, describing the steady decline of Italian against Mandarin’s rise over 30 years, personifying the two series (“Previously I was more than you, and then we became the same, and then you overtook me”).
WS4 interaction patterns closely matched strategies co-designers had articulated during WS3, suggesting the co-designed approaches to reading and consuming charts were being internalized. Notably, co-designers consistently returned to touch after receiving the agent’s response, tracing the referenced data to verify it — extending the select, confirm, ask pattern with a fourth step — verify (Fig. 8).
2pt
Following the free-form activity, co-designers reflected on their experience using the NASA-TLX dimensions and a question assessing the perceived naturalness of the system (“The way of interacting with the system felt natural”, 7-point Likert scale) (Tab. ¿tbl:tab:ws4?).
NASA-TLX: Physical demand was deemed low by all the co-designers. Mental demand was deemed low by D1 and D2, but D1 noted that additional effort stemmed from hardware constraints, rather than interaction complexity: “The display resolution is the limit, not the system’s capabilities.” In contrast, D3 reported higher mental demand, noting that they were still learning the system’s capabilities: “I still had to make deliberate interaction choices, to remember [for instance] ‘oh I can actually maybe ask this thing,’ \(\ldots\) but this will improve with use.” Performance was deemed high by all the co-designers, with D2 noting that “using the system, I got to understand the trends and shape of the data well.” Finally, D3 singled out latency as a source of friction.
Naturalness was rated highly by all co-designers. D3 valued the breadth of multimodal interaction options: “I like the possibilities \(\ldots\) use your hands and sometimes speech to confirm or verify; a very patient tutor.” Even after the three-month gap between WS2 and WS3, they recalled how to ask the agent questions, select data points via gestures (though not the exact timing), and navigate the layered overview. By WS4, the co-designers used the full range of interactions unassisted.
When reflecting on why they continued touch exploration after receiving an agent response, they described three main motivations: verification, sensemaking, and independence.
Verification: D1 described doing “all three at once”, but reflected that without the RTD, verification would not be possible: “Without the tactile display, I have to trust the agent, there’s no certain way of verifying.” D2 explained that they were “verifying the agent’s accuracy and building [their own] understanding [through touch]”, and noted, “I don’t immediately trust anything with an LLM. I want to look at the actual data.” Similarly, D3 spoke of prior experiences with AI: “I’ve experienced too many instances of AI not giving accurate information, so I don’t trust it completely, and often verify it for accuracy.” Even so, the co-designers emphasized that they would still want to explore by touch, even if the agent’s responses were guaranteed to be correct (e.g.,computed from a spreadsheet, not an LLM).
Sensemaking: Touch was also described as central to building an overall understanding. D2 noted that “the richness of the information you get [through touch] is greater than a verbal response”, and D1 added that “exploring the tactile representation of the data gives me a better overall understanding than a verbal summary, even if the exact details of particular data points are less clear.”
Autonomy and independence: D3 described the agent as a last resort, preferring to resolve questions themselves through touch first: “Asking is the last resort, because I like to figure things out myself.” D1 noted that “[touch] enables me to be independent, allowing me to consume data at my own pace.”
Touch and the agent took on distinct, complementary roles. Following Chundury et al.’s recommendation [7], our work combined conversational speech interaction with tactile representations to address the limitations of each modality in isolation. Touch served as the primary sensemaking channel: co-designers traced to understand shape and trend, comparing and cross-referencing values against axes and other chart features. This aligns with our WOz study, where participants with tactile experience exhibited similar behavior. Our co-design process reveals why, showing that touch serves a broader role than independence.
In contrast, the conversational agent served as a data question-answering specialist, responding to queries that touch alone could not resolve. Its role evolved as familiarity grew: WS1 centered on basic value retrieval and extrema; WS2 saw use of trend analysis with deictic queries; WS3 shifted toward richer comparative analysis; and by WS4’s more open-ended task (§4.6), co-designers were attempting to push the agent beyond its current capabilities (e.g.,compound queries). In summary, touch supported the spatial understanding that the agent lacked, while the agent provided analytical depth that touch could not.
Our co-designers articulated reasons why they chose to anchor their exploration in touch rather than the agent: independence, verification, and sensemaking (§4.6.3). Co-designers valued independence as a form of agency: the ability to build their own understanding through touch, at their own pace, before (or instead of) asking the agent. This echoes Choe et al.’s [52] finding that, when interpreting charts, sighted users with higher data literacy engaged more with charts and used LLMs as more of a specialist, whereas those with lower data literacy over-relied on the agent. Our BLV co-designers, who had both data and tactile experience, showed a similar pattern. Whether this holds for BLV users with less tactile and data experience remains an open question.
Touch also served as a method of verifying the agent’s responses. Driven by a distrust of LLMs, co-designers reflected that without the plotted chart on the RTD, verification would not be possible (§ 4.6.3). Seo et al. [44] illustrate this gap: although MAIDR supports tactile output in the form of Braille Unicode characters on single-line RTDs, without 2D spatial representations of charts, their BLV participants relied on other verification strategies, such as cross-referencing against other LLMs, or relying on background knowledge. A chart spatially rendered on an RTD offers something these strategies cannot: the ability to feel whether the agent’s response matches the shape of the data.
At the same time, framing touch as only a verification method or means of maintaining independence undermines its broader role. Co-designers emphasized that they would still explore by touch even if the agent’s responses were guaranteed to be correct, suggesting that the RTD provides something that speech fundamentally cannot: not just answers about the data, but a spatial understanding of it.
When a sighted user opens a chart, they can instantly survey its shape, peaks, and trends; a BLV user must move their hands and touch-read to build up an equivalent understanding, one feature at a time. On an RTD, this becomes even more difficult, due to its low resolution. This raises a fundamental design question: how to present charts to users on RTDs. A single continuous narration, akin to alt text or a transcriber’s note [6], proved mismatched: co-designers found it too long to follow while exploring by touch, with two channels competing for attention rather than reinforcing one another.
Their solution was a layered overview: chart components introduced one at a time, each rendered on the RTD with a spoken description. During workshops, co-designers insisted that the chart remain fully interactive within each layer, supporting touch exploration, gestures, and agent queries, and that layers be navigable, such that they could independently move back and forth through a chart. By WS4, the layered overview had evolved into what we now term a layered presentation: no longer just orientation and overview, but the primary framework in which co-designers explored and consumed their charts.
Each layer is displayed in isolation, creating a tight coupling between what is being heard and what is being felt: the spoken description narrates the component rendered on the RTD, and nothing else. Co-designers used this to focus their touch exploration, building spatial understanding without competing with tactile clutter. Zong et al.’s [5] Olli breaks charts into individual components via a hierarchical tree navigable by screen readers; our layered presentation takes a similar approach in the tactile domain: components are introduced in a guided sequence that grounds the chart’s structure in something that users can physically feel. D2’s reversal, from “gimmicky” to essential, illustrates how this coupling between touch and speech transformed a linear walkthrough into an interactive exploration structure.
The same challenge arises during agent interaction: longer responses overwhelm when users engage in touch exploration. Response segmentation addresses this by breaking agent responses into navigable one-sentence chunks, each synchronized with animated tactile highlights that draw attention to referenced data on the chart. This allows users to feel the data the agent is describing and pace the interaction themselves, verifying the agent’s claims against the spatial layout of the chart and building their understanding as they go. Both the layered presentation and response segmentation serve the same purpose: structuring multimodal output so that touch and speech reinforce one another rather than compete for attention.
Jiao et al.’s [41] tactile data comics share a step-by-step multimodal presentation on RTDs, but are strictly presentational; our layered presentation extends this into scaffolding for interactive data exploration, where co-designers could freely navigate layers or skip them entirely: D2 chose to skip the interpretive summary while still receiving the foundation they needed [2].
The layered presentation was particularly valuable with unfamiliar chart types: all co-designers traversed back and forth through the scatterplot, building their conceptual understanding. D3 suggested the technique could generalize beyond charts to other tactile graphics, supported by Jiao’s work [41], where step-by-step presentation improved comprehension across mathematics, maps, and literature.
In our earlier study [8], we used WOz to simulate an idealized system where a human wizard seamlessly coordinated touch and speech. Participants with more tactile experience preferred touch, with speech as a complementary channel. Our co-designers, each of whom had tactile experience, exhibited the same pattern, preferring to build understanding through touch before turning to the agent (§ 5.1). Our v1 prototype carried forward capabilities from the WOz study: gesture selection, free touch during speech, deictic queries, wake word activation, and a narrated chart overview (Tab. [tab:features]). Our co-designers retained some of these, replaced others, and introduced additional capabilities as new challenges emerged during the process of building a fully instantiated system that WOz had abstracted away:
In the WOz study, risk was invisible: as participants verbally articulated their intended interactions using think-aloud, the Wizard always correctly interpreted touch and speech, so the system never misunderstood. With a real conversational agent, infallibility is no longer guaranteed: speech transcriptions can misinterpret the query, and the LLM can hallucinate or otherwise produce incorrect results. With Graphy, co-designers developed their own verification strategies through touch, driven by a distrust of AI-generated responses (§ 4.6.3). They also rejected the free touch method identified in WOz, where participants would simultaneously touch points when asking their question, in favor of a sequential pattern: select, confirm, ask, ensuring each step was grounded before querying the agent. This evolved into select, confirm, ask, verify (§ 4.6), as co-designers used touch to verify the agent’s responses against the chart.
The Wizard served as a real-time guide, narrating chart overviews from prepared scripts. When instantiated, a single continuous narration (v1) proved difficult to follow, with co-designers unable to retain the information while simultaneously exploring the chart by touch. The layered presentation (§ 5.2) emerged as the co-designed solution, evolving beyond the Wizard’s orientation mechanism into the primary mode of exploring and consuming charts (§ 4.6).
WOz participants extracted values by gesturing specific points, or by asking the Wizard directly, but both methods required already knowing where a point was located. Stepping—moving through values one at a time using buttons (analogous to navigating with a screen reader)—was proposed as a faster, self-service alternative. It was particularly useful for precise interaction and recovery if gestures overshot, and for guided exploration where narrated values helped predict the shape of the data.
The WOz study used the Graphiti RTD, which supports raising pins to four heights. The current work used the Dot Pad RTD, which offers only binary output (raised or lowered). Without pin height as an encoding channel, co-designers shaped encodings within these constraints: designing distinct point symbols, connecting lines with clearance gaps, and fill textures for bar charts. However, this constraint also drove exploration of the encoding design space, including the intersection symbol and a feedback grammar that uses pin animation to distinguish user- and agent-initiated actions. This knowledge is specific to binary-pin RTDs, a constraint shared by the majority of the RTDs on the market (e.g.,the Monarch).
These challenges and solutions were also shaped by the longitudinal nature of the co-design process. Unlike our WOz study, which comprised single 90–120 minute sessions, this work spanned four workshops per co-designer over eight months, allowing co-designers to internalize interactions and refine preferences through repeated use—a depth our WOz study could not produce.
Based on our findings, we contribute design guidance for researchers and developers building CTDIs:
Rather than a single continuous narration, introduce chart components sequentially [41], each rendered in isolation on the RTD with a spoken description, to reduce clutter. The same principle should extend to how the agent responds to questions: segment responses into navigable chunks, each synchronized with attention-drawing tactile feedback, so that users can feel the data being described at their own pace. Charts should always remain interactive, so that users can extend their sensemaking by selecting values and querying the agent.
Touch and speech complement one another: touch provides spatial grounding, the agent calculation and analytical depth [7]. When users combine the two in deictic queries, the agent should acknowledge the touch context before answering so that users can confirm their input was correctly interpreted. Agent responses should be concise and answer-first, enabling verification grounded in the chart.
Each interaction should confirm the previous before moving on: select (via gesture or button), confirm the selection (via audio, Braille or tactile feedback), query the agent, then verify the response through touch. The pattern select, confirm, ask, verify (Fig. 8) grounds each step in the data, ensuring queries have the correct context, and enabling users to check the agent’s response against the chart.
Different exploration goals require different navigation methods. Support gestures for selecting known points, and button stepping for traversing unknown points. Filtering can further aid exploration by reducing tactile clutter.
Blinking pins can draw attention on RTDs [38], but when both user and agent can trigger highlights, users need to distinguish the source. Use distinct patterns for each: static highlights for gesture selections, animated highlights for agent references, and transitional animations for stepping that settle to static. Highlights should persist until dismissed by the user, affording them time to locate the highlighted component. The grammar should be consistent across chart types in order to reduce cognitive load.
Visual encoding conventions do not reliably transfer to tactile encoding [45], and conventions for printed tactile graphics do not necessarily transfer to RTDs due to their low resolution [46]. Encodings should be designed for touch discriminability with RTD pin arrays and be user-configurable, akin to how screen readers allow adjustment of speech rate and verbosity.
CTDIs could be a game changer for accessible data visualization. At their core, though, they are constrained by RTD hardware itself. Our co-design process surfaced three hardware limitations of current pin-grid RTDs that manufacturers and researchers should consider:
The Dot Pad and the Monarch RTD support only binary output, constraining encodings to a single pin height. The Graphiti RTD [28] supports four heights, enabling data series and chart elements to be distinguished by elevation rather than symbols alone, but remains inaccessible at $25,000 USD. As RTDs become more affordable, broader support for multiple heights would expand the design space of tactile encoding for data visualization on RTDs.
The Dot Pad lacks built-in touch sensing entirely; the Graphiti and Monarch support only single-finger touch detection, with the Monarch requiring users to lift other fingers when gesturing. We used an external finger tracker supporting both index fingers, but built-in multi-touch would enable interaction techniques such as pinch-to-zoom and multi-finger gestures for interactive data analysis.
The Dot Pad and Monarch use electromagnetic actuators [61] that cannot reliably raise pins against finger pressure, affecting Braille reading and tactile feedback alike. Reliable actuation under the finger would allow users to feel tactile feedback in place, making the search for highlighted elements easier.
Beyond these improvements, we see the need for tactile displays that move beyond binary, braille-pitch pin grids. Higher-resolution technologies or deformable surfaces would better support complex, denser charts. Further, we envision AI-native devices with the agent, speech processing, and rendering built into the RTD rather than distributed across a host computer and cloud services (as ours required), making multimodal interaction more portable and self-contained.
Several limitations of our work point to clear directions for future work. Our co-design involved three congenitally blind co-designers with varying levels of experience using tactile graphics and RTDs. This number is consistent with co-design practice in accessible visualization [5], [23], [43], [58]. Our findings may not generalize to BLV users with acquired blindness, less tactile experience, or residual vision, who may exhibit different usage patterns, and may rely more on the agent and less on independent touch verification. The layered presentation could give such users the confidence needed to engage with data through touch, while investigating interaction techniques better supporting agent-led exploration remains a clear opportunity.
We aimed to generate design knowledge, not to evaluate performance. While WS4 included free-form exploration with co-designers’ own data, it was not a formal usability study. Future work should also include controlled and in-situ evaluation to assess how effectively users can consume and interpret data visualizations, and to measure the agent’s accuracy and robustness, which we did not quantify.
Our co-design workshops utilized traditional RTDs, comprising binary pins on a fixed, braille-pitch grid. Hence some of our design knowledge, such as tactile encodings, may not transfer to future tactile-display paradigms. However, higher-level interaction principles, e.g.,the layered presentation and the select, confirm, ask, verify pattern, are likely to be device-independent, and so carry over to future displays.
We explored a limited set of chart types (line, bar, stacked bar, scatterplot), all relatively low in data density; of these, scatterplots were the least resolved. More complex or denser charts will strain the low resolution of current RTDs, where adjacent or overlapping marks might be difficult to distinguish. We expect the layered presentation to be even more valuable in these cases, isolating components to keep each readable. Future work should expand our tactile encoding guidance to additional chart types. By WS4, co-designers were also pushing the agent’s boundaries with compound queries requiring multiple operations, which Kim et al. [51] found prevalent among BLV users; support for compound queries grounded in touch would be a clear next step.
Co-designers also identified the need for interactive chart manipulation: charts that exceed the RTD’s display require zooming and panning, but refreshing the entire display risks disorienting users who may lose spatial context [8], [62]. Maintaining spatial context during chart manipulation remains an open challenge for our future work, as does integration with existing tools like Power BI and Excel.