Ouvia: A User-centered Framework for Measuring Usability of
Speech Translation in Real-World Communication Scenarios
June 04, 2026
Speech translation (ST) is increasingly adopted in user applications, yet its evaluation largely focuses on decontextualized testbeds and holistic quality, rather than end users’ communication needs. We introduce Ouvia, an evaluation framework for measuring user-perceived usability of speech translation outputs in real-world settings. Ouvia focuses on one-to-one communication: an English speaker needs to convey a request to a Portuguese speaker, and the message is automatically translated. Through a custom web app and multi-phase study design, we collect more than \(1{,}750\) such interactions in healthcare and everyday situations, mediated by four ST systems, involving speakers from three English dialects and two genders. We find that modern ST serves people only to a limited extent—only around half of interactions are rated as usable—with significant gaps in reported usability across demographic groups. Moreover, among quality metrics, we find that QA-based evaluation is a substantially stronger predictor of real-world usability than standard approaches. Together, these findings stress the importance of situated, user-centered evaluation frameworks that go beyond holistic quality scores and attend to who the technology serves—and how well.1
Speech translation now underpins a broad range of user-facing applications, from cross-lingual real-time conferencing to software for multilingual education2 or clinical healthcare triage.3 A key catalyst of this progress is the field’s ability to evaluate translation systems rigorously, whether through targeted evaluation campaigns [1] or scalable automatic quality metrics [2]–[4].
While laudable, these efforts address only a fraction of what evaluation demands. Campaigns typically run in vitro: they present decontextualized (non-situated) test segments to participants and treat quality as a holistic concept—detached from function and purpose [5], [6] and primarily serving system performance tracking and ranking rather than actual real-world use. Moreover, automatic quality assessments tell little about ramifications—e.g., whether bad translations imply clinical risk [7], or quality gaps reflect real-world allocative harms [8]—and risks amplifying social biases [9]. Purely automatic, decontextualized, holistic evaluation cannot be the sole end. It should, instead, be supported by new forms of inquiries, centered around end-users— today, largely laypeople [10]—their communication needs, and real-world uses and situations [11], [12].
This paper introduces Ouvia,4 an evaluation framework for capturing user-perceived usability of speech translation systems in real-world communication scenarios. We pursue three overarching goals: (1) capture lay users’ perceptions of the usability of translation outputs across different high- and low-stakes situations; (2) measure whether speaker dialect and gender identity give rise to gaps in translation quality and usability, and how these relate; (3) investigate the predictive power of (automatic) quality metrics for usability, and characterize the practical implications of metric choice. Ouvia bundles the following contributions:
A conceptual evaluation framework that moves ST evaluation from decontextualized testbeds to situated, user-centered assessment. We devise a multi-phase protocol that mimics ST-mediated one-to-one interactions (see Figure 1): a sender initiates a conversation with an English request, which is automatically translated for a Portuguese receiver who answers comprehension questions based on the translation. A third bilingual validator annotates the exchange with quality signals. Finally, the sender is surveyed on several dimensions of usability. This design captures what in-vitro evaluation cannot: whether a translation is usable for a specific person, in a specific situation, to accomplish a specific communicative goal.
A user study, data, and findings. We implement the framework through a user study grounded in English-to-Portuguese translation in Portugal—a highly multicultural country where large segments of the immigrant population rely on English as a lingua franca [13]. We recruit native English speakers across different dialects [14] and Hindi speakers as a non-native group—the largest migrant population in Portugal after Brazilians and the fastest-growing, for whom mediated translation is particularly relevant. We collect \(N=1{,}738\) subjective judgments from \(174\) unique speakers on translations from four state-of-the-art ST systems across healthcare and everyday scenarios. Our findings show that current systems serve real-world users only to a limited extent, that usability gaps exist across speaker demographics, and that QA-based evaluation predicts usability substantially better than standard coarse-grained metrics—reinforcing the need for situated, user-centered evaluation of ST.
A custom web platform and speech corpus. To run the study, we develop a web platform supporting speech recordings and multi-user interactions. To support future research, we release \(14.6\) hours of prompted speech recordings, \(13.8\)K+ human annotations for QA-based quality evaluation, and \(1.7\)K+ translation quality scores.
We focus on communication scenarios [15]: a speaker wants to correctly convey a message to someone in a different language. In this setting, correctness might not necessarily mean generic translation quality. Rather, the exchange might be successful “as long as people get the information they want, […] or manage to convey their intentions” [16]. To study this at scale, we approximate such an interaction through an online study with a novel multi-phase design. Besides the core participants in the exchange, we additionally enlist a third-party bilingual validator to independently assess the correctness and quality of the translation mediating the exchange.
Each study unit consists of four sequential phases.
Phase . A sender records a passage we collected (§2.2). This conversation starter is 40-60 words long and contains key information, such as named entities or quantities, that a system should translate correctly (see bold phrases in Figure 1).
Phase . We translate the sender’s recording automatically and show the outcome to a receiver, asking them to reply to a set of open-ended questions grounded in the original passage. (§2.3).
Phase . We recruit a bilingual validator to verify the translation in two ways. (i) A direct assessment score [17]—Direct Assessment with Scalar Quality Metrics [18]—where we provide a slider from 0 to 100, marked with seven labeled tick marks indicating different quality labels that capture accuracy and grammatical correctness (details in Appendix 9.3). (ii) An assessment for each question indicating if the receiver answered it correctly.
Phase . We give the automatic translation to the initial sender and task them with a survey to collect their subjective judgments about the translation. Specifically, they rate their agreement on a 5-point scale (1 = Strongly Disagree, 5 = Strongly Agree) with statements aimed to assess: i) satisfaction perceived with the translation output (adapted from [19]); ii) trust, the belief that the output will be beneficial (adapted from [20]); and iii) reliance, the self-reported behavioral intention to depend on the output in practice (adapted from [21]).5 Crucially, we encourage them to answer as if they were in a real-life situation needing to communicate with someone who speaks Portuguese. Together, these judgments provide a situated indication of three dimensions of a translation usability (§3).6
Since senders do not speak Portuguese, we administer the survey at two time points. First, we capture their baseline—their unconditioned tendency toward the translation usability based solely on the source passage, their recording, and the Portuguese translation they cannot linguistically verify. Upon completion, we disclose the outcomes of Phases and , enabling senders to make another, but informed, final judgment.
For a realistic condition in , we aim for conversation starters that are truthful in content and style to real-life interactions, are of high quality, and cover both high- and low-stake scenarios. Moreover, we prioritize information-dense starters, i.e., passages that include multiple key information. To increase diversity, starter data are sampled from existing corpora as well as automatically generated. We confirm their quality via a manual validation, as described in Appendix 9.2. 7
Healthcare Scenarios. We target the healthcare domain, where automatic translation finds frequent use [23] and errors have profound ramifications [24]. We sample entries from MED-MT [25]. The dataset contains transcripts of simulated patient-physician interviews, primarily focusing on respiratory cases. Since the original transcripts contain filled pauses and disfluencies, we use Gemini 2.5 Pro to strip them and improve fluency without changing the core content. To increase content and style diversity, prompted Qwen 3 32B and Gemini Pro 2.5 to create over-the-counter pharmacy conversation starters, which mentions to symptoms, their duration, medication names, and allergies.
Everyday Scenarios. We source customer requests from BConTrasT [26], which consists of customer support dialogues. As in the healthcare case, we collect synthetic LLM-generated starters expressing requests about everyday scenarios, such as booking a taxi, banking requests, or hotel check-in.
Overall, we sample 50 items each from both MT-MED and BConTrasT and generate 100 samples per domain (50 from each LLM), collecting a total of 300 unique conversation starters. Appendix Table 3 reports examples and statistics.
Largely inspired by translation evaluation via question answering [27], we automatically generate up to 10 unique questions for each starter. Following [28], who show improved question relevance when generating in English, we generate questions in English before translating them into Portuguese. For both generation and translation steps, we use Gemini 2.5 Pro, prompting the model to ground the question in the starter’s main information (see an example in Figure 1, ). See Appendix Figure 8 and 7 for the prompt used. We validate quality, as reported in Appendix 9.2. Out of 200 questions, only one was answerable. Errors are thus extremely rare, but we account for them in the validation stage of our study: receivers can flag questions as unanswerable and validators can accept this as correct.
We experiment with four translation systems based on open-weight models that are currently state of the art [29]. We include three direct speech translation models, Phi 4 Multimodal [30], Voxtral Small [31], and DesTA2 [32], and one cascaded system that uses Whisper large-v3 [33] to transcribe the recordings followed by Tower+ 9B [34], a strong language model specialized for translation in European languages. After a sender completes , we pick randomly one translation system using default inference parameters and prompts to avoid introducing confounders. See generation details in Appendix 9.4.
We recruited study participants on Prolific.8. We require the first language of receivers to be Portuguese and proficiency in both English and Portuguese for validators. To explore the impact of native English dialects and non-native accented speech, we recruited senders from three language groups based on (self-declared) nationality and ethnic group:9 (1) US White and (2) US Black speakers, whose first language is English, and (3) Hindi native speakers, who are proficient in English as a second language. Within each group, we balanced participants by gender identity. Using the platform’s filters, we select “Man” and “Woman.”10
Since 21% of the senders dropped out of the study (Phases and were, on average, six days apart), we recruited replacement participants only to complete the survey in . We saw no significant difference between third- and first-person survey scores (see Appendix 8.1 for details).
For senders, we use a between-subjects design with respect to group membership [35]: each participant belongs to exactly one language-by-gender combination, with 30 participants recruited per cell. Cell sizes were determined to support adequately powered comparisons [36]. The stimuli (starters) are shared across all groups, enabling both within- and across-group comparisons. Each participant is assigned 10 starters, stratified across domain (healthcare, mundane), source type (synthetic, natural), and model. Stratified assignment preserves internal validity by ensuring no individual participant is disproportionately exposed to any domain or model, thus preventing individual biases. For receivers, we recruit 20 participants per language group and randomly assign them 30 translations each. Similarly, we recruit 20 participants who each validate 30 exchanges. To prevent practice effects, we randomize the order of items displayed to each participant.
4pt
| S | T | R | \(\mu(u)\) | \(\sigma(u)\) | \(\Delta_u\) | |
|---|---|---|---|---|---|---|
| USW | 3.91 | 3.84 | 3.86 | 3.87 | 1.15 | -0.19 |
| \(\text{b}\rfloor\) | 4.11 | 4.05 | 4.01 | 4.06 | 0.86 | |
| USB | 3.50 | 3.48 | 3.46 | 3.48 | 1.18 | -0.20 |
| \(\text{b}\rfloor\) | 3.7 | 3.66 | 3.69 | 3.68 | 0.91 | |
| Hindi | 3.40 | 3.36 | 3.31 | 3.35 | 1.20 | -0.40 |
| \(\text{b}\rfloor\) | 3.78 | 3.74 | 3.75 | 3.76 | 0.88 |



Figure 2: (a) Survival curves of usability scores by language group. Each curve shows the fraction of the group whose usability rating reaches or exceeds a given score on the x-axis. Annotations at \(u=4\). (b) Fixed-effect estimates (95% CI) from a LMM predicting usability (\(u\)) (§4.1). Filled circles \(p < 0.05\), diamonds \(0.05 \leq p < 0.1\), hollow non-significant. Estimates annotated right of each interval. (c) Estimated marginal means (\(\pm\)95% CI) of usability by language group and gender, marginalized over topic and model..
Our study results can be analyzed across several cohorts, identified by three sender language groups (US White: USW, US Black: USB, Hindi speakers), two genders (Woman, Man), conversation topic (Health, Everyday), source (Natural, Synthetic), and translation model. To formalize the relationships between these factors and their impact on usability, we describe them in a causal graph: the sender, starter’s source and topic, and translation model all contribute to the translation quality, etc.11 The graph is shown in Appendix Figure 5.
Prompted speech collected through online platforms is prone to recording artifacts due to (device or environment) noise, microphone activation delays, or unfamiliarity with recording software that can cause the first milliseconds of an utterance to be clipped [37]. We controlled for these issues with a semi-automatic procedure. We first transcribed all recordings with Whisper Large v3 and flagged entries meeting either of two conditions: the first or last three words (after normalization) do not match the prompt, or WER exceeds \(0.6\). One author then manually reviewed all flagged entries to eliminate false positives, and entries that failed verification were discarded. We excluded participants with more than three faulty entries entirely. These checks led to the exclusion of two women and two men from USW, one woman and one man from USB plus two individual records from USB, and two women and three men from Hindi. The final corpus counts \(N=1{,}738\) observations across \(174\) senders, \(60\) receivers, and \(69\) validators.
Prior literature establishes that trust and satisfaction can predict reliance [21], [38], which led us to expect that senders would rate these three dimensions similarly within a given translation. To verify this, we run a standard factor analysis, confirming that more than 90% of the variance in the scores is explained by a single underlying factor (see Appendix 10.1). For practical purposes, we average the three scores at the instance level, creating a single compound usability score \(u\) based on the informed final judgments (§2.1), which we use as the target variable of interest in the study. We compare it to the baseline \(u\) in an upcoming analysis.
When studying the correlation between translation quality and usability, we rely on the following set of automatic metrics, established and widely used in the NLP literature for their high agreement with human assessment of quality: COMET [39], COMET Kiwi Base and XL [2], XCOMET-XL [40], and MetricX 24 XL [41]. We also include the judgments from : Translation Score and QA Score, defined as the ratio of correct answers given in .
We begin by investigating our main interest: whether our participants have found the translations usable (§4.1). Upon noting differences across groups, we analyze the effect of demographic factors (§4.2) and translation quality (§4.3). Finally, we study how different “quality” signals relate to user-perceived usability (§4.4).



Figure 3: (a) Fixed-effect estimates (95% CI) from a linear mixed model predicting validator translation score. Reference: Hindi, Man, DeSTA2, Everyday, MED-MT. Filled circles \(p<0.05\), diamonds \(0.05 \leq p<0.1\), hollow non-significant. (b) Model-adjusted marginal means (\(\pm\)95% CI) by language group and gender, marginalized over topic, model, and source. (c) Descriptive means (\(\pm\)95% CI) by translation model and conversation topic..
Not fully, with most usability scores falling between 3 and 4 in a 5-point Likert scale. Table 1 reports descriptive statistics across language groups, including the (final) usability \(u\) scores compared to the baseline scores. Figure 2 (a) shows the \(u\) score distributions.
USW score usability the highest in both the baseline (\(u_b=4.06\)) and final conditions (\(u=3.87\)) while Hindi the lowest (\(3.76\) and \(3.35\), respectively). When comparing baseline and final assessments, all groups decrease their scores, and the distribution of their scores also increases in variance. Therefore, the quality signals mitigate prior assumptions and increase the polarization of judgment. This result is expected, as this signal likely surfaces errors more evidently to them. However, such an effect varies across language groups. Hindi lower their scores (\(-0.40\)) more than twice as much as USW do (\(-0.19\)), an effect likely mediated by receiving poorer quality translations (discussed in §4.2). The two factors, hence, compound: USW and Hindi start from uneven baselines, and this gap is exacerbated when becoming aware of external quality signals. We do not record the same phenomenon between USW and USB.
Both contextual factors and demographics affect usability significantly. We fit a linear mixed regression model (LMM) to predict usability (\(u\)) using the translation model, starter topic, source, and baseline \(u\) as fixed effects, and sender, validator, and receiver as random effects. Figure 2 (b) shows some of the resulting coefficients (full details in Appendix 10.2). The largest variance is explained by the translation model, with Voxtral and Tower+ yielding the best scores (\(p<0.05\); see per-model \(\mu(u)\) in Appendix Figure 10). Scores on Health-related starters are significantly higher than Everyday (+\(0.33\), \(p<0.05\)).
Interestingly, the LMM discloses that language variant-based differences (Table 1) are not significant. However, we observe a negative effect of Gender: Woman (\(-0.2\), \(p<0.1\)) and a positive intersectional effect for the intersection of USW x Woman (\(+0.36\) over the reference Hindi x Man, \(p<0.05\)). As this finding suggests that within-language group differences might exist, we plot the marginal means estimated by the LMM in Figure 2 (c). The figure offers one additional core insight: gaps in usability across groups are dampened by similar results on men; women, instead, scored translation unevenly. The biggest difference is between USW and Hindi, where usability is \(4.18\) and \(3.69\), respectively—a significant gap (\(p<0.05\) if computing pairwise differences, after Holm correction). Together, these findings suggest that a combination of factors—led by the choice of the translation model and the intersectional demographic group—determines self-perceived usability. Moreover, these usability gaps across demographic groups raise a natural question:
To address this question, we focus on the validator’s SQM quality scores—one of the two human judgments of translation quality our study provides (more in §4.4). We fit a new LMM to predict it and inspect the model coefficients (see Figure 3). Compared to Hindi, validators score USW and USB higher by \(13.71\) (\(p<0.05\)) and \(7.43\) (\(p<0.1\)) points, respectively. Marginal means (Figure 3 (b)) highlight these gaps better, where USW men score \(76.4\), against Hindi men’s \(62.7\). Interestingly, within Hindi—the group with the worst quality according to validators—women score higher than men, despite later expressing lower usability scores (Figure 2 (c)). Moreover, Figure 3 (c) reports the means by model and topic. As expected, translation quality also varies significantly by translation model, with Voxtral and Tower+ leading the board and Health receiving higher scores overall.
Yet validator SQM scores are just one lens, and an expensive one. The question is whether cheaper, automated signals tell the same story.
| Metric | e | elow | emed | ehigh |
|---|---|---|---|---|
| MetricX 24 | 2.35 | 1.35 | 0.74 | 1.03 |
| XCOMET XL | 1.94 | 1.13 | 0.79 | 0.82 |
| COMET | 2.85 | 1.27 | *0.58 | 0.65 |
| COMET Kiwi | 2.43 | 1.37 | *0.47 | 1.35 |
| COMET Kiwi XL | 2.69 | 1.44 | 0.77 | 1.33 |
| Translation Score | 2.11 | 0.69 | 0.55 | 0.32 |
| QA Score | 2.94 | 1.61 | 2.64 | 3.06 |
Partially, and QA is superior to coarse-grained overall judgments. We compute simple Spearman rank correlation against \(u\) and note a variable degree across metrics (see Appendix Figure 11). Among automatic metrics, XCOMET XL is the highest (\(\rho=0.49\)), and COMET is the lowest (\(\rho=0.42\)), whereas Human judgments have reasonably higher correlation. However, QA Score has \(\rho=0.63\), which is interestingly higher than translation Score (\(\rho=0.56\)).
To further investigate these differences, we conduct two more measurements. (1) We fit one LMM per metric using the same fixed- and random-effects setup as in §4.1 adding the metric to covariates, and storing the resulting coefficient—this number measures the metric’s effect on usability. (2) We split the records following the Translation Score’s tertiles, identifying three quality regimes; for each, we fit one Ordinary Least Squares (OLS) regression on \(u\) (more robust than LMM for setups with fewer data points), again rotating the metrics and recording the coefficient. These numbers tell us how closely each metric tracks usability at varying quality. Table 2 reports these coefficients. We find that automatic metric effects decline as quality increases. Translation Score shows a similar trend, whereas QA confirms the stronger effect, outscoring all metrics across all regimes. This finding suggests that fine-grained content-based evaluation tracks usability more consistently than a single overall quality score at varying quality levels.
Our results call for renewed interest in human-centered evaluation of ST. We demonstrate that user-perceived usability and coarse-grained translation quality are related but distinct constructs—and that this distinction matters. This is where Ouvia adds value: by situating evaluation in a real communicative exchange, it surfaces nuances that aggregate quality scores obscure—from which systems and demographic groups benefit, to where standard metrics fail as deployment signals.
For realistic use cases, direct ST is not always the right solution: both Phi 4 and DeSTA2, despite being recently released, lag behind a simpler cascade pipeline that combines a smaller foundation speech recognition model (Whisper v3) with a larger, highly specialized translation model (Tower+) (see Figure 3 (c) for details). This finding underscores that “new” does not guarantee “usable,” and that practitioners should evaluate systems on ecologically valid tasks rather than standard benchmarks alone.
We find significant usability gaps along both language variety and gender. This finding aligns with prior work documenting inequitable ST performance on standard benchmarks [42], but our results extend this concern to systems in actual deployment scenarios. As shown in Figure 2a, more than \(66\%\) of USW speakers find the system usable (\(u\geq4\)), compared to only \(49\%\) for USB speakers and \(43\%\) for Hindi speakers. These are not abstract disparities: because our usability measure explicitly captures willingness to rely on the translation in real-world situations, what is at stake is whether a speaker will use ST technology for consequential communication, depending on the variety of English they speak.
Translation quality predicts usability, but not all quality signals are equally informative. Standard quality estimation metrics rank highly on leaderboards for overall quality prediction [43], yet in communicative settings, they explain only a fraction of variation in user-perceived usability. QA score, which measures whether key information, such as named entities and quantities, is translated correctly, is a substantially better predictor. Crucially, it is also more robust: in the high-quality regime, differences in standard automatic metrics poorly distinguish between translations that users find more or less usable, whereas differences in QA score do. Figure 2 illustrates this contrast directly.
Translation quality metrics are not designed to capture usability, so their limited predictive power is unsurprising. They also remain opaque: providing a generic overall score with mostly comparative power—model X outperforms model Y—but offering little sense of whether, and how effectively, a translation serves a user’s communicative needs [11]. Yet usability human assessment is not scalable. We thus leverage our study to anchor existing metrics to a concrete usability threshold. In Figure 4, we map QA Score—a potentially automatable evaluation signal—and COMET—the metric with the highest impact on usability (Table 2)—onto the usability scores they correspond to. At \(u\) = \(4\), a reasonable threshold for a translation a lay user would satisfactorily rely on in practice, a QA Score of \(\sim\) \(0.91\) and a COMET score of \(\sim\) \(0.82\) can serve as practical deployment benchmarks. Notably, unlike COMET, QA score remains more meaningful beyond this point, thus adding further evidence in favor of its discriminative power. When human judgments are unavailable, these thresholds can offer a principled, usability-grounded basis for system selection and deployment decisions.
Our work sits at the intersection of user-centered approaches to machine translation (MT) [12] and demographic disparities in speech technologies [44]–[46]. For MT, prior work has explored user needs in realistic settings across medical and migration domains [47], [48]. [49] shows that users overly trust MT outputs with mistranslations, and [50] focuses on quality feedback that helps users calibrate MT reliance in communication settings. [8] further demonstrates that automatic metrics fail to capture the real-world human and economic costs of gender bias in MT. Calls for grounding evaluation in user needs and societal relevance are growing [11], [51], [52], yet remain largely unaddressed in ST—the emerging frontier for human-to-human mediated communication. Our work is the first to bring these concerns together under that modality, revealing the limits of ST for communication, across accents, as well as showing the limits of coarse-grained assessments for real-world usability. Our findings point toward QA-based assessment [27], [28]—which targets accurate transmission of key information rather than closeness to a human reference—as a stronger proxy for the real-world usability that lay users need in situated, communicative scenarios.
We have introduced Ouvia, a new evaluation framework to study the user-perceived usability of several state-of-the-art ST systems. We described how we designed a multi-phase user study that underpins our research questions, collecting \(N=1,738\) judgments from \(174\) speakers across diverse demographic groups. Our results show that current systems serve real-world users only to a limited extent, that usability gaps persist across speaker demographics, and that QA-based evaluation is a substantially stronger predictor of usability than standard automatic metrics. Together, these findings reinforce the need for human-centered, situated evaluation of ST—one that goes beyond holistic quality scores and attends to who the technology serves, in what context, and to what end.
Reliable assessments require sizable data samples and participant pools, which we prioritize during budget allocation. For this reason, we limit our study to one language pair (English - European Portuguese), while including speakers across genders and language variants motivated by realistic usage scenarios of immigrants in Portugal. We remain cautious about generalizing our findings to language pairs that are underserved by current speech technologies, where usability gaps may be more pronounced.
Our conversation starters are grammatically fluent, and senders are instructed to read them aloud without fillers, repetitions, or disfluencies. While this allows us to control for content quality, it departs from naturalistic spoken communication, where such features could affect downstream translation. Similarly, we ask senders to record in noise-free environments, resulting in clean, read speech that differs from the acoustic conditions typical of real-world use—such as background noise. Thus, these conditions likely favor current speech translation systems, meaning our usability estimates may be optimistic relative to deployment scenarios. While—to the best of our knowledge—our study is the first to approximate a realistic communication setting for evaluating speech technologies, future work should examine usability in more spontaneous settings.
Our validators are bilingual crowdworkers rather than professional translators. Still, non-expert can provide reliable assessment using standardized scales [18]. Furthermore, our binary QA validation task requires comprehension rather than translation expertise.
QA-based evaluation is a strong predictor of usability in our data, particularly in high-quality regimes, pointing to its potential as a future evaluation proxy. Still, further work is needed to understand this relationship across other communicative settings and with fully automated pipelines. Notably, our QA assessment is manual and disclosed to the sender—yet so is the validator’s holistic quality score, which correlates substantially less with usability. This suggests it is the fine-grained, information-targeted nature of QA that drives its predictive power, not the human involvement or disclosure alone. Moreover, QA-based evaluation in Ouvia is simplified by (i) the existence of ground truth conversation starters, which we use to extract the questions, but that might not be generally available, and (ii) the starters’ content itself, which is, by design, information-rich in our study.
Ouvia is designed for evaluating the usability of ST systems across diverse demographic groups. All participants were recruited through Prolific and financially compensated according to the platform’s standards. The study was approved by the Ethics Committee of Instituto Superior Técnico. All participants provided informed consent prior to participation.
Voice data carries inherent re-identification risks: a speaker may be recognized even absent explicit metadata, and recordings can in principle be used for voice cloning. Our study design minimizes these risks by collecting only few segments for each participant (10), with short recordings (30s or less), which hinder faithful voice cloning. The released dataset will require users to explicitly agree not to use it for voice cloning or individual identification purposes.
Gender is among the most salient perceptual traits of one’s identity [53], [54], and gendered differences represent a central concern in sociophonetic research [55] and speech technologies [56], [57]. Following Prolific’s screening filters, we operationalize gender as a self-declared attribute using Man and Woman as categories, each subsuming both cisgender and transgender speakers. We initially included a Non-Binary option (as available in Prolific) for our two largest cohorts (US White and US Black). However, we ultimately excluded it: the category is an overly broad umbrella term that conflates distinct gender identities, making group-level comparisons theoretically unsound. This was further reflected in low uptake (four participants). Recruiting through queer and allied communities to improve participation and gender diversity beyond the binary remains an important direction for future work.
We recruited all participants through https://www.prolific.com/, an established platform to recruit crowdworkers. To join the study, each participant had to read and accept the Terms and Conditions, which had been approved in advance by an Ethics Committee. Participants could join the study only once and in a single role. All participants were paid 9 GBP/h. All participants were allowed to join the study using a smartphone, tablet, or notebook. The following is a list of criteria to join the study:
USW Sender: Nationality: USA, First Language: English, Gender: Man/Woman (including transgender), Broad Ethnic Group: White, Approval Rate: \(75\)–\(100\).
USB Sender: Nationality: USA, First Language: English, Gender: Man/Woman (including transgender), Broad Ethnic Group: Black/Afroamerican, Approval Rate: \(75\)–\(100\).
Hindi Sender: First Language: Hindi, Fluent Language: English, Gender: Man/Woman (including transgender), Approval Rate: \(75\)–\(100\).
Receiver: First Language: Portuguese, Gender: Man/Woman (including transgender), Approval Rate: \(75\)–\(100\).
Validator: First Language: Portuguese, Fluent Language: English, Gender: Man/Woman (including transgender), Approval Rate: \(75\)–\(100\).
To validate the use of replacement participants in cases where original senders dropped out of the longitudinal study, we examined whether subjective usability scores differed depending on whether the assessment was given from a first-person perspective (senders evaluating their own recorded interactions, \(N=1{,}358\)) or a third-person perspective (replacement participants evaluating interactions they had not produced, \(N=380\)). We fitted a linear mixed-effects model (REML) predicting the average survey score from perspective alongside translation quality, comprehension rate, translation model, gender, and language group as fixed effects, with random intercepts for participant and conversation starter. The perspective coefficient was negligible and non-significant (\(\beta = 0.009\), \(p = .873\)), indicating no meaningful difference in subjective ratings between the two perspectives. These results justify the replacement strategy: crowdworkers assessing conversations they did not themselves produce provided ratings statistically indistinguishable from those of the original senders.
None
Figure 6: Example prompt for synthetic data generation. Over-the-counter pharmacy scenario..
None
Figure 7: Prompt for Question Translation. We use it to prepare Portuguese questions in ..
None
Figure 8: Prompt for Question Generation. We use the generated questions in ..
None
Figure 9: Paraphrasing Prompts. Used to prepare the conversation starters from MT-MED (top) and BConTrasT (bottom). “seed” is the original data point from the respective dataset..
Our starters sourced from existing sources required some rephrasing to be transformed into a plausible conversation started. In practice, we used Gemini 2.5 Pro to conduct the rephrasing. In all cases, we sampled with temperature of 0.9, top_p of 0.95, and top_k of 20, and used the following system prompt: “You are a language model specialized in generating high-quality synthetic data.”. Figure 9 reports the prompts used for both datasets.
To assess the quality of our starters and associated questions, we conducted a manual validation on a sample of 20 starters associated with 10 sets of unique English questions (i.e. 7% of the dataset, 20 starters and 200 questions). The analysis was conducted by one author with a background in translation studies. We verified the starter content and the grounding of such questions in the starter content. No issues were found with the starters. Out of 200 questions, only one raised an ambiguity, making it arguably unanswerable. Such issues are thus rare and do not undermine the overall quality of the dataset. Moreover, they are accounted for in the validation process: receivers can explicitly flag a question as unanswerable, and validators can accept this response as correct.
For the quality assessment, we use the SQM scale that features seven labeled tick marks indicating different quality labels combining accuracy and grammatical correctness described as follows:
6: Perfect Meaning and Grammar: The meaning of the translation is completely consistent with the source and the surrounding context (if applicable). The grammar is also correct.
4: Most Meaning Preserved and Few Grammar Mistakes: The translation retains most of the meaning of the source. It may have some grammar mistakes or minor contextual inconsistencies.
2: Some Meaning Preserved: The translation preserves some of the meaning of the source but misses significant parts. The narrative is hard to follow due to fundamental errors. Grammar may be poor.
0: Nonsense/No meaning preserved: Nearly all information is lost between the translation and source. Grammar is irrelevant.
We use Hugging Face’s transformers classes and models to run all automatic translations [58]. For Phi
4 (HF ID: microsoft/Phi-4-multimodal-instruct) and Voxtral Small (HF ID: mistralai/Voxtral-Small-24B-2507), we use the prompt “Translate the audio to Portuguese.”, whereas for DeSTA2 we use the slightly more complex “Translate this audio to Portuguese.
Produce only the Portuguese translation, without any additional explanations or commentary.” as the model showed the tendency to generate additional content besides the translation. We load Whisper v3 and Voxtral in bfloat16. In all runs, we used a batch
size of 1 and standard decoding parameters as implemented in transformers. To translate the transcript with Tower+, we use the following prompt: “Translate the following English source text to Portuguese (Portugal):\nEnglish:
{text}\nPortuguese (Portugal):” and greedy decoding.
4pt
| Source | Topic | Example | Count |
|---|---|---|---|
| MED-MT | Patient-Physician | Hello, I came in today because I have been feeling unwell for about a week now. My primary symptom is a very persistent and sore throat that hasn’t improved. In addition to that, I have also started experiencing chills over the last few nights, so I thought it would be best to get it checked out professionally. | 50 |
| Synthetic | Over-the-counter Pharmacy | Hi there, I’m hoping you can help. For the past five days, I’ve had terrible seasonal allergies—constant sneezing, a runny nose, and really itchy, watery eyes. I’ve been taking one 10mg Cetirizine tablet every morning, but it’s barely making a dent. Is there a more effective over-the-counter option you’d recommend for these symptoms? | 100 |
| BConTrasT | Customer Support | Hello, I’d like to place an order for pickup from Bella Luna Pizzeria. I need four large pizzas, all with extra cheese, please. The first one should have pepperoni and pineapple, the second one should be the five-cheese blend, the third a taco pizza, and the last one should just have tomatoes, olives, and green peas. | 50 |
| Synthetic | Everyday | Excuse me, good morning. I’m trying to get to the Belém Tower from here at Rossio station. I believe I need to take the 15E tram, is that correct? I’ll need to buy a Viva Viagem card for two people for a single trip; could you tell me the total cost and the approximate travel time around 11 AM? | 100 |
Since each sender reports three scores—satisfaction, trust, and reliance—we conduct a factor analysis to verify whether they represent distinct constructs or can be treated as a single composite measure. Given the nested structure of the data—10 interaction-level observations per participant—we decompose variance into a within-user component (person-mean-centered scores, \(n = 1{,}200\)) and a between-user component (per-user means, \(n = 120\)), and run exploratory factor analysis (EFA) independently on each. Prior to fitting, Bartlett’s test of sphericity confirmed that the correlation matrices were significantly different from identity at both levels (within: \(\chi^2 = 5927.9\), \(p < .0001\); between: \(\chi^2(3) = 733.6\), \(p < .001\)), and Kaiser-Meyer-Olkin (KMO) values indicated adequate sampling adequacy (within: \(\text{KMO} = 0.781\); between: \(\text{KMO} = 0.774\)). Parallel analysis (500 iterations, 95th percentile threshold) retained a single factor at both levels. The one-factor solution explained \(89.1\%\) of within-user variance and \(92.7\%\) of between-user variance. These results indicate that satisfaction, trust, and reliance do not function as empirically distinct constructs in our data: participants rated them nearly interchangeably across interactions and across individuals. We therefore average the three scores into a single composite measure, which serves as the dependent variable in all subsequent analyses.
We use pymer4 [59] to fit Linear Mixed Effects Models in our analysis, using REML. Table 4 reports the coefficients, confidence intervals, and significance level (\(p\)) of the model used to predict the usability variable \(u\).
| term | Estimate | Std. Error | t | p | 2.5% | 97.5% |
|---|---|---|---|---|---|---|
| (Intercept) | 1.564 | 0.164 | 9.535 | 0.000 | 1.242 | 1.886 |
| Group: USB | 0.136 | 0.133 | 1.022 | 0.308 | -0.127 | 0.399 |
| Group: USW | 0.131 | 0.135 | 0.975 | 0.331 | -0.135 | 0.397 |
| Gender: Woman | -0.202 | 0.120 | -1.677 | 0.096 | -0.440 | 0.036 |
| Model: Phi 4 | 0.069 | 0.067 | 1.035 | 0.301 | -0.062 | 0.200 |
| Model: Tower+ | 0.644 | 0.069 | 9.378 | 0.000 | 0.509 | 0.779 |
| Model: Voxtral | 0.733 | 0.068 | 10.860 | 0.000 | 0.601 | 0.865 |
| \(u\) (Baseline) | 0.466 | 0.030 | 15.504 | 0.000 | 0.407 | 0.525 |
| Topic: Health | 0.332 | 0.064 | 5.177 | 0.000 | 0.206 | 0.459 |
| Source: Gemini | -0.541 | 0.084 | -6.483 | 0.000 | -0.706 | -0.377 |
| Source: MED-MT | -0.343 | 0.109 | -3.147 | 0.002 | -0.557 | -0.128 |
| Source: Qwen 3 | -0.452 | 0.083 | -5.413 | 0.000 | -0.616 | -0.287 |
| Group: USB × Gender: Woman | 0.070 | 0.172 | 0.408 | 0.684 | -0.269 | 0.410 |
| Group: USW × Gender: Woman | 0.363 | 0.173 | 2.098 | 0.038 | 0.021 | 0.706 |
We used AI coding tools to streamline the generation of visual artifacts of the paper, and writing assistants to polish parts of this manuscript.
Study platform and data under CC-BY 4.0 at https://github.com/g8a9/ouvia.↩︎
Past tense of Portuguese ouvir, to hear.↩︎
The exact statements are: i) “I am satisfied with the quality of the AI translation”; ii) “I trust the AI translation to convey my message”; iii): “I would use this AI translation in a real-world situation”.↩︎
Although usability is traditionally defined in HCI as the effectiveness, efficiency, and satisfaction with which users achieve goals with a specific system [22], here we use it more broadly to capture how useful a user finds a translation in a given communicative context.↩︎
See Appendix 9.1 for full details on automatic data cleaning and generation of both high- and low-stakes sets.↩︎
https://www.prolific.com/. See Appendix 8 for details on recruitment.↩︎
We acknowledge “ethnicity” is a fraught concept, but use it to match the screening filters available on Prolific.↩︎
We discuss the motivations and implications of this binary setup in Ethical Considerations.↩︎
We acknowledge the existence of other contingent factors—e.g., the age of participants, the recording device, background noise—which we do not control in this study.↩︎