ChildVox: A Speech, Audio, and Large Audio-Language Model Benchmark in Understanding and Characterizing Sound across Childhood

Tiantian Feng1 Anfeng Xu1 Xuan Shi1 Aditya Kommineni1 Shakhrul Iman Siam2
Megan Micheletti3 Zhonghao Shi4 Helen Tager-Flusberg5 Mi Zhang2
Lynn K. Perry6 Catherine Lord3 Daniel Messinger6 Shrikanth Narayanan1
1University of Southern California 2The Ohio State University
3University of California, Los Angeles 4Harvard University
5Boston University 6University of Miami
tiantiaf@usc.edu


Abstract

We present ChildVox, a novel benchmark for characterizing the diverse acoustic signals through which children communicate. Specifically, ChildVox follows the full developmental trajectory from birth through school age, covering physiological sounds, non-linguistic vocalizations, canonical syllables, and spoken language. ChildVox integrates more than 20 sub-tasks across 17 child-centered audio and speech datasets, enabling systematic cross-corpus and cross-domain comparison. We evaluate a representative range of audio and speech foundation models, including self-supervised, ASR-oriented, and large audio-language models, on tasks including physiological sound classification, vocalization and canonical syllables modeling, and speech quality assessment and recognition. Benchmark results show that ChildVox provides a suite of high-performance models in recognizing a wide range of acoustic signals from children, supporting downstream applications such as characterizing children’s language levels and tracking speech production with age.

1 Introduction↩︎

Deep learning and Artificial Intelligence for child-centered speech or voice processing and modeling have long focused on automatic speech recognition (ASR) [1][3], a general-purpose task for populations with established spoken language skills. However, children “express” themselves through a wide range of verbal and non-verbal skills at different stages of childhood [4], shaped by their physical [5] and neuro-developmental status [6] and educational backgrounds [7]. Therefore, tasks like ASR cover only a subset of children’s communicative skills and life experiences, and do not fully capture signals critical for understanding development, language level, health, and behaviors across childhood.

This limitation is amplified for children who have not yet fully acquired socially recognized spoken language, or who experience speech sound disorders (SSD) and other conditions that impact speech production. In early childhood, for example, communication is often dominated by non-spoken forms, including vocalized events (e.g., cries). Moreover, physiological sounds (e.g., cardiac sounds [8]) can carry critical information when verbal or non-verbal cues are insufficient for understanding a child’s condition.

In this work, we therefore rethink what constitutes speech, voice, or language in the context of child development. We introduce ChildVox, an audio benchmark that extends conventional ASR modeling to broader forms of children’s voices”, enabling the characterization of sounds across childhood experiences for child-centered applications. We broadly view “voice” in children as embodied communication that includes not only spoken language, but also vocalization, canonical syllables, and physiological sounds, all of which work together to convey meaning and express internal states. This definition calibrates communication skills to the developmental stage, recognizing that how children interact with the world evolves as their linguistic, motor, and cognitive skills develop.

Figure 1: Overview of the wide range of ways that children “express” themselves. ChildVox benchmarks audio, speech, and large audio-language models on physiological sounds, vocalizations, canonical syllables, and speech.

Our proposed ChildVox benchmark offers unique contributions compared to prior works: (1) Unlike previous studies that primarily focus on children’s ASR, ChildVox considers the full developmental trajectory from birth through school age, covering physiological sounds, vocalizations, canonical syllables, and speech within a single unified benchmark. (2) ChildVox integrates more than 20 sub-tasks across 17 child-centered audio and speech datasets into a consistent evaluation protocol. (3) We evaluate a representative range of self-supervised, ASR-oriented, and large audio-language models, showing that different pre-training setups offer complementary measures of child-centered acoustic signals. Moreover, models provided in ChildVox consistently outperform the frontier proprietary models like Gemini 3.5 Flash. Extensive experiments show that ChildVox provides a robust platform for supporting downstream applications in understanding children’s audio and spoken language across developmental contexts.

2 Related Works↩︎

2.1 Child ASR Benchmarks↩︎

Several recent studies have proposed benchmark-style evaluations for children’s ASR, where many of these works consistently show substantial degradation when adult-trained ASR systems are applied to children’s speech, due to developmental acoustic variability, pronunciation differences, and conversational dynamics. For example, [1], [2] evaluate supervised and self-supervised speech foundation models for child ASR, analyzing the effects of model scaling, training paradigms, and datasets. Other studies examine ASR robustness of speech from children with developmental delays, highlighting challenges from atypical vocal behaviors [9]. Recently, causal analyses of children’s ASR errors further quantify the contributions of physiological, cognitive, and environmental factors to recognition failures [3].

Table 1: Summary of the audio and speech data used in the ChildVox benchmark. The proposed benchmark includes more than 20 sub-tasks related to physiological sound classification, vocalization classification, canonical syllables classification, and speech recognition and classification.
Dataset Evaluation Tasks Labels Data License
CirCor [8] Murmur Detection Absent, Present, Unknown Open Database
2-4 Crackle Detection Crackles, No crackles Open Database
Wheeze Detection Wheezes, No wheezes
Respiratory Condition Healthy, COPD, Other condition
2-4 Respiratory Sound Normal, CAS, DAS, Not Specified
CAS & DAS, Poor Quality
Donate-a-cry Cry Classification Hunger, Other Open Database
2-4 Cry Classification Hunger, Loneliness, Discomfort Not Specified
2-4 Child speech, Babbling, Giggle, CC-BY-4.0
Child Sound Children shouting, Baby laughter,
Classification Baby cry, Whimper, Child singing,
Children playing, Music for children
ReCANVo [10] Affective Status Delighted, Dysregulated, Frustrated, Not Specified
Request, Selftalk, Social
2-4 Vocal Development Crying, Laughing, Canonical, Customized License
Classification Non-Canonical, Junk
2-4 Vocal Development Crying, Laughing, Canonical, Customized License
Classification Non-Canonical, Junk
C-BESD [11] Emotion Recognition Anger, Happy, Neutral, Sad Not Specified
2-4 Pronunciation Rhotic, Derhotic PhonBank License
2-4 Prosody Poor, Nearly correct, Correct intonation CC-BY-4.0
Fluency Little influent, Fluent in general, Fluent
Pronunciation Poor or Understandable, Good, Excellent
2-4 Speech Articulator Bilabial, Labiodental, Dental, Glottal CC-BY-NC-4.0
Labiovelar, Alveolar, Palatal, Velar
2-4 Diarization Speaker labels and timestamps Private
Speech Intelligibility Intelligible, Unintelligible, Vocalization
2-4 Diarization Speaker labels and timestamps Private
ASR Speech Transcript

2-4

MyST [12]

ASR Speech Transcript Customized License
2-4 TinyVox [13] ASR Phoneme Transcript Not Specified

2.2 Child Speech Benchmarks beyond ASR↩︎

Apart from ASR, recent work has applied speech foundation models to a range of child-centered tasks, showing that self-supervised representations capture meaningful developmental information from children’s speech and infant vocalizations [14]. Other studies have explored speech foundation models for child-adult speaker diarization in naturalistic dyadic interactions [15], spoken language development analysis in children with ASD [16], and phoneme recognition for language learning applications [17]. Additionally, layer-wise analyses from a recent work further reveal that different model layers encode distinct developmental information [18]. Collectively, these studies show the growing potential of speech foundation models beyond ASR in children’s speech.

3 ChildVox Benchmark↩︎

Figure 1 presents the overview of benchmark design to characterize the range of ways that children “express” themselves from birth through school age.

3.1 Benchmark Overview↩︎

Unlike conventional ASR-centric benchmarks on children’s speech, ChildVox captures the diverse spectrum of child-produced and child-related acoustic signals that emerge throughout development: physiological sounds, such as cardiac sounds that provide markers of health conditions; vocalizations, such as crying in infancy; canonical syllables and emerging language behaviors in toddlers, such as proto-speech that reflect the progressive development of communicative skills; and the more complex social communicative behaviors of school-age children associated with spoken language interaction in everyday interactions. Collectively, this developmental framing across childhood positions the ChildVox as an ecologically grounded platform for advancing child-centered audio intelligence.

3.2 Benchmark Tasks and Datasets↩︎

ChildVox evaluates four broad categories of child-centered audio datasets: physiological sounds, vocalizations, canonical syllables, and speech. The information regarding evaluation tasks, labels, and data license is provided in Table 1. The details about each dataset are described in the Appendix 11.

3.2.0.1 Physiological Sound Classification

ChildVox includes three benchmark datasets for pediatric physiological sound classification covering both cardiac and respiratory assessment. The benchmark targets four tasks: murmur detection, crackle detection, wheeze detection, and respiratory condition classification. Murmur detection is evaluated on CirCor [8], a collection of 5,272 pediatric cardiac auscultation recordings, while crackle detection, wheeze detection, and respiratory condition classification are evaluated on ICBHI [19] and SPRSound [20], two respiratory sound datasets consisting of pediatric and general patient populations.

3.2.0.2 Vocalization Event Classification

ChildVox targets two vocalization-related tasks: generic child sound and infant cry classification. Child sound classification covers ten child-related audios: speech, babble, giggle, shout, laughter, cry, whimper, sing, play, and music for children. This is drawn from AudioSet [21], a large-scale ontology-driven collection of YouTube data. Cry cause classification is evaluated on two infant cry corpora: Donate-a-Cry [22], where we classify hunger versus other, given its imbalanced distribution, and CryBank [23] is with labels: hunger, loneliness, and discomfort.

3.2.0.3 Canonical Syllables Classification

We experimented with ReCANVo [10], a database of affective nonverbal sounds recorded in everyday interactions from minimally speaking children and adults with neurodevelopmental conditions. This corpus supports affective status classification (e.g., dysregulated). BabbleCor [24] and SpeechMaturity [25] target early vocal development in children, supporting vocal development classification across cry, laugh, canonical, non-canonical, and junk.

3.2.0.4 Speech Quality Assessment

We include three datasets evaluating speech production, pronunciation, and articulation. PERCEPT-R [26] (with PhonBank [27]) contains recordings from children with typical speech development and speech sound disorders, emphasizing rhotic production in American English. SpeechOcean762 [28] provides 5,000 English utterances from 250 non-native speakers (half children), each annotated by five expert raters for pronunciation accuracy, fluency, and prosody. UltraSuite [29] pairs synchronized ultrasound tongue imaging with acoustic recordings for research on child speech disorders. Inspired by [30], we use it for articulator classification across eight places of articulation (e.g., velar) on word-level audios.

Figure 2: Training data distribution of each sub-dataset in ChildVox-Balanced dataset for training LALMs.
image
image

3.2.0.5 Adult-child Speaker Diarization

A core component of analyzing child speech in interaction is speaker diarization, which detects speech segments and labels “who spoke when”. Here, we include two in-house datasets to benchmark the speaker diarization. Specifically, the Natural Language Sampling (NLS) [31] dataset includes child-parent conversations collected from remote dyadic interactions involving minimally verbal children with ASD. The corpus includes 73 English-speaking child-parent sessions with detailed child/ adult speaker and speech intelligibility labels. Moreover, the ADOS2-Mod3 dataset [32] consists of recordings from more than 100 children participating in Autism Diagnostic Observation Schedule-Module3 [33] assessments. Within this dataset, interactions related to emotion and social-difficulty segments are annotated with speaker labels.

3.2.0.6 Automatic Speech Recognition

We target both word-level and phoneme-level transcription of children’s speech. Word-level recognition is evaluated on MyST [12], a corpus of elementary school students interacting with a virtual science tutor, and on ADOS2-Mod3, whose transcripts are annotated with emotion and social-difficulty segments. Phoneme-level recognition is evaluated on TinyVox [13], which provides over half a million phoneme-level transcriptions across multiple languages. Here, we focus on the English subset with one item.

3.3 ChildVox-Balanced Dataset↩︎

Given the substantial aggregate scale of the combined datasets, we construct ChildVox-Balanced, a curated subset with labels balanced within each classification task, to support practical and reproducible supervised fine-tuning of Large Audio Language Models (LALMs) mentioned later. It draws only on publicly accessible resources and limits training and testing samples per label to 2,000 and 50, respectively. For ASR evaluation, we include 10,000/500 utterances from MyST and TinyVox for training/testing. Figure 2 shows the resulting per-dataset distribution, where is a total of 64,641 audio and speech recordings across 14 subtasks.

4 Benchmark Models↩︎

In ChildVox, we evaluate several representative audio, speech, and large audio-language models for benchmark experiments. The details of the experimented pre-trained models are listed in Table [tab:foundation95models].

4.1 Encoder-based Models↩︎

4.1.0.1 Self-supervised Models

We evaluate three representative self-supervised audio and speech foundation models: SSAST [34], voc2vec [35], and WavLM [36]. SSAST, voc2vec, and WavLM apply self-supervised learning to learn representations from generic audio events, non-verbal vocalizations, and speech, respectively. These pre-trained models have shown strong performance on audio event detection, speech recognition, and speaker diarization tasks.

4.1.0.2 ASR Models

We primarily evaluate models from the Whisper family, an encoder-decoder architecture trained on large-scale multilingual speech primarily for ASR. Apart from ASR, Whisper encoders have shown strong transferability across speech processing problems, such as child-adult speaker diarization [15]. To explore the effect of model scale, we evaluate encoder embeddings from three Whisper variants: Whisper-Base, Whisper-Small, and Whisper-Large-v3. Moreover, we include Parakeet-TDT [37], [38] for fine-tuned ASR tasks.

6pt

c l *6S[table-format=1.3] & Dataset & SSAST & voc2vec-H & WavLM-L & Whisper-B & Whisper-S & Whisper-L
& CirCor & 0.636 & 0.563 & 0.643 & 0.499 & 0.546 & 0.552
& ICBHI-crackles & 0.644 & 0.626 & 0.612 & 0.581 & 0.597 & 0.586
Physiological & ICBHI-wheezes & 0.638 & 0.592 & 0.614 & 0.614 & 0.621 & 0.609
Sounds & ICBHI-Condition & 0.466 & 0.445 & 0.392 & 0.473 & 0.483 & 0.508
& SPRSound & 0.448 & 0.368 & 0.439 & 0.345 & 0.362 & 0.414
& AudioSet-Child & 0.657 & 0.560 & 0.539 & 0.598 & 0.634 & 0.619
& CryBank & 0.300 & 0.303 & 0.321 & 0.333 & 0.326 & 0.323
& donate-a-cry & 0.473 & 0.486 & 0.468 & 0.467 & 0.528 & 0.502
& ReCANVo & 0.444 & 0.416 & 0.423 & 0.414 & 0.415 & 0.359
& Speech Maturity & 0.686 & 0.640 & 0.620 & 0.625 & 0.634 & 0.640
& BabbleCor & 0.569 & 0.514 & 0.485 & 0.546 & 0.596 & 0.594
& NLS-Speech Intelligibility & 0.588 & 0.548 & 0.603 & 0.635 & 0.658 & 0.661
& Percept-R & 0.789 & 0.788 & 0.813 & 0.794 & 0.807 & 0.811
& Ultrasuite & 0.850 & 0.635 & 0.882 & 0.923 & 0.926 & 0.888
& SpeechOcean-Fluency & 0.601 & 0.554 & 0.613 & 0.625 & 0.626 & 0.627
& SpeechOcean-Accuracy & 0.564 & 0.557 & 0.625 & 0.605 & 0.628 & 0.649
& SpeechOcean-Prosody & 0.613 & 0.613 & 0.705 & 0.701 & 0.704 & 0.715
& C-BESD & 0.753 & 0.566 & 0.892 & 0.596 & 0.593 & 0.532

4.2 Large Audio-Language Model↩︎

We include two widely used Large Audio-Language Models (LALMs) in our benchmark: Qwen2-Audio-Instruct [39] and Audio Flamingo 3 (AF3) [40]. Both models unify diverse audio understanding across speech, environmental sounds, and music. Specifically, Qwen2-Audio combines a Whisper-large-v3 audio encoder with the Qwen-7B language model backbone. Similar to Qwen2-Audio, the audio encoder of AF3 is initialized from Whisper-Large-v3.

5 Experiments↩︎

5.1 Data Preprocessing↩︎

For all datasets included in ChildVox, audio samples are resampled to 16 kHz to match the sampling rate of all pre-trained models. Across most experiments, we restrict a minimum input duration of 200 ms to maintain sufficient acoustic information for downstream evaluation. For datasets without predefined training, development, and test partitions, we adopt a 5-fold cross-validation setup in the experiment. More details, such as the prompt design for the SFT experiments with LALM and data filtering, are in the Appendix 11 and  13.

5.2 Modeling Details↩︎

The details of training settings, like learning rate and training epochs, are listed in the Appendix 12.

5.2.0.1 Model Parameters

We fine-tune SSAST using full-model optimization, whereas for voc2vec-HuBERT, WavLM-Large, and Whisper-family models, we keep the pretrained backbone weights frozen and adopt parameter-efficient fine-tuning through LoRA. Specifically, we apply LoRA with a rank of 64 to the feedforward layers for voc2vec, WavLM, and Whisper as shown in Figure [fig:lora]. For LALMs, LoRA modules with rank 64 are inserted into the query, key, value, down-projection, and up-projection layers of the language model, and feedforward layers of the audio encoder.

5.2.0.2 Evaluation

We report utterance-level Macro-F1 on the test set for most classification tasks in the benchmark. For speaker diarization, we report diarization error metrics, while ASR and phoneme recognition are evaluated using word error rate (WER) and phoneme error rate (PER), respectively. Results for encoder-based models are reported as averages across multi-folds in Appendix 11.

6 Benchmark Results↩︎

6.1 Encoder-based Model Results↩︎

Table [tab:benchmark95results] presents the results of the encoder-based model on the complete ChildVox benchmark data.

6.1.0.1 Physiological Sound Classification

For physiological sound classification tasks, SSAST and WavLM-Large generally outperform Whisper-family models or the voc2vec model. In particular, WavLM-Large achieves the highest performance on CirCor (0.643), while SSAST performs best on ICBHI-crackles, ICBHI-wheezes, and SPRSound. These results suggest that models learned with generic audio representation may better capture fine-grained spectral and temporal characteristics present in physiological signals. In contrast, Whisper-based models show comparatively lower performance on most physiological sound tasks, indicating that speech-oriented pretraining may not provide benefits for physiological sound classification, particularly in children.

6pt

l *4S[table-format=2.2] & &
(lr)2-3 (lr)4-5 Model & NLS & ADOS & MyST & ADOS
WavLM-Large & 22.10 & 45.60 & 16.56 & 59.34
Whisper-Base & 24.20 & 48.20 & 16.99 & 55.85
Whisper-Small & 18.70 & 44.00 & 15.93 & 59.91
Whisper-Large & 17.70 & 42.50 & 14.80 & 40.20
Parakeet-TDT & – & – & 15.82 & 45.60

6.1.0.2 Vocalization and Canonical Syllables

For vocalization and canonical syllables modeling, the results reveal a more balanced ranking between SSL-based and ASR-oriented models. SSAST achieves the best performance on AudioSet-Child (0.657), ReCANVo (0.444), and Speech Maturity (0.686), suggesting that audio-based pre-training is still beneficial for modeling non-linguistic and early developmental vocal behaviors. Unlike physiological sound classification, Whisper-based models achieve competitive performance on multiple tasks, such as BabbleCor and donate-a-cry.

3pt

Table 2: Benchmark performance on ChildVox-Balanced test set using encoder-based models and LALMs. We report PER and WER (lower the better) for TinyVox and MyST, respectively. We report the Macro-F1 score (higher the better) on the remaining sets. SO refers to the SpeechOcean762 dataset.
Model CirCor SPRSound AudioSet ReCANVo Speech PERCEPT-R SO-Prosody TinyVox MyST
(\(\uparrow\)) (\(\uparrow\)) (\(\uparrow\)) (\(\uparrow\)) Maturity (\(\uparrow\)) (\(\uparrow\)) (\(\uparrow\)) (\(\downarrow\)) (\(\downarrow\))
SSAST 0.543 0.508 0.638 0.500 0.735 0.800 0.644 - -
WavLM-Large 0.571 0.518 0.538 0.448 0.612 0.843 0.726 - 0.147
Whisper-Base 0.349 0.362 0.605 0.373 0.663 0.769 0.745 - 0.147
Whisper-Large 0.402 0.452 0.636 0.370 0.689 0.815 0.759 - 0.128
AudioFlamingo3 0.452 0.162 0.089 0.157 0.069 0.323 0.400 0.958 0.274
Qwen2-Audio 0.397 0.541 0.699 0.514 0.726 0.779 0.671 0.602 0.133

6.1.0.3 Speech Quality Assessment

For speech quality classification, Whisper models achieve consistently higher performance. Specifically, Whisper-Large obtains the best results on all tasks from the SpeechOcean762 dataset. This suggests the strong speech recognition capabilities acquired from large-scale multilingual pre-training. On the other hand, SSL-based models remain competitive, with WavLM-Large achieving the best score on speech emotion recognition. Overall, these results indicate that speech-based pre-training benefits the classification of children’s speech, despite such models being trained predominantly on adult data.

6.1.0.4 Speaker Diarization and Speech Recognition

Table [tab:diar95asr95results] reports diarization and ASR performance across the NLS, ADOS2-Mod3, and MyST datasets. For speaker diarization, Whisper-Large achieves the lowest error rates on both NLS (17.70 DER) and ADOS2-Mod3 (42.50 DER). For ASR, Whisper-Large also obtains the best results, reaching 14.80 WER on MyST and 40.20 WER on ADOS, with Parakeet-TDT ranking second on both ASR benchmarks. Overall, model scale is the dominant factor, with the largest Whisper model leading on the speech recognition and diarization tasks.

6.1.0.5 Findings

Overall, performance varies substantially across child-centered audio categories, highlighting the diverse acoustic properties of physiological sounds, vocalizations, canonical syllables, and speech. Meanwhile, no single model dominates across all tasks, whereas different encoders capture complementary aspects of these signals.

6.2 Large Audio Language Models Results↩︎

Table 2 compares the two LALMs against the encoder-based models on the ChildVox-Balanced subset, where two LALMs behave substantially differently in their performance. Qwen2-Audio is consistently competitive with the strongest encoder-based models, achieving the best scores on AudioSet and ReCANVo. It remains close to the best encoder-based models on SpeechMaturity and Percept-R, and obtains a low WER on MyST. In contrast, AudioFlamingo3 performs substantially worse on nearly every task, with particularly low scores on vocalization and canonical syllables classification, such as AudioSet and SpeechMaturity. It also returns an extremely large error rate on TinyVox phoneme recognition (0.958 PER).

Figure 3: Macro-F1 on ChildVox-Balanced test set. Zero-shot proprietary models (Gemini 2.5/3.5 Flash) underperform the ChildVox -trained encoder and Qwen2-Audio baselines across all five datasets.

6.2.0.1 Error Inspection

Manual inspection indicates that this performance difference is driven by inconsistent instruction-following from the response, where AudioFlamingo3 often returns free-form descriptions (e.g., the audio clip is of a baby laughing”) instead of the requested label. It also tends to summarize or hallucinate on transcription tasks. For example, given the reference transcript “Well, without electricity, an electromagnet would never be possible, because it gives it the magnetism,” AF3 hallucinates a summary rather than a transcription, producing “The electromagnet is not working.”

6.3 Comparison with Proprietary Baselines↩︎

We compare zero-shot Gemini 2.5 Flash and Gemini 3.5 Flash against our ChildVox-trained encoder (best one) and Qwen2-Audio on five public datasets in Figure 3. We limit this comparison to these five datasets to ensure that no restricted or license-protected data is transmitted to third-party APIs. Overall, specialized models from ChildVox outperform both proprietary baselines on every task. Specifically, both Gemini models perform particularly poorly on CirCor, SPRSound, and ReCANVo (Macro-F1 < 0.35), suggesting that general-purpose LALM struggle with fine-grained pediatric vocal and physiological audios. In summary, these results highlight the effectiveness of models trained on the ChildVox benchmark. Zero-shot Gemini settings are detailed in the Appendix 14.

7 Exemplar Applications with ChildVox↩︎

Inspired by the previous work [13], we design two experiments using the benchmark models developed from the ChildVox to measure the child’s speech production and language development in two different applications.

7.1 Characterizing the Language Level↩︎

7.1.0.1 Application Setup

Our first application is on the in-house NLS dataset described in Section 3. In this study, domain experts categorize each child’s spoken language level into pre-verbal communication, first words, and word combinations, following the definitions by [41]. Specifically, we refer to these three levels as LL-1, LL-2, and LL-3, respectively. We then apply our speaker diarization estimation models to compare language level differences across three groups.

7.1.0.2 Observations

Figure 4 reveals a clear monotonic relationship between language level and utterance rate: the median number of utterances per minute increases steadily from LL-1 (pre-verbal communication) through LL-2 (first words) to LL-3 (word combinations). This trend indicates that children at more advanced spoken language levels are more vocally active, producing speech at a higher rate. Moreover, the variability in utterance rate is increased with language level, with LL-3 showing the most varied distribution. Overall, these results show that diarization-derived utterance rate captures a meaningful developmental signal consistent with expert-assigned language levels.

Figure 4: Utterance rate (utterances per minute) derived from the speaker diarization model by language level (LL-1 to LL-3) in the NLS study. The violin plot shows that the utterance rate increases with the language levels.

7.2 Speech Production Development↩︎

7.2.0.1 Application Setup

Our second application uses the PERCEPT-R corpus (Section 3), which contains single-word rhotic-targeted speech samples from children with typical development and SSDs, each manually annotated for correct rhotic production. We apply our Whisper-Large Rhotic/Derhotic classification model to investigate how the predicted probability of rhotic production varies with chronological age. Specifically, we restrict the prediction only to typical developing children.

Figure 5: Correlation between model-predicted probability of correct rhotic production and chronic age on PERCEPT-R dataset. Each dot represents a single participant, where the model-predicted probability is aggregated from all reading samples of a participant.

7.2.0.2 Observations

Figure 5 shows a moderate positive correlation between chronological age and the model-predicted probability of rhotic production in PERCEPT-R (\(r=0.576\)). Specifically, younger children (\(\sim\)​115-135 months) show larger variability, while predicted probabilities from older children centered more consistently around 0.75. Several lower-probability outliers remain present at older ages, highlighting continued inter-speaker variability. These results indicate the ChildVox provides models that can capture developmental changes in rhotic articulation across late childhood

8 Conclusion and Future Work↩︎

We introduce ChildVox, a benchmark characterizing the diverse acoustic signals through which children communicate from birth through school age, integrating more than 20 sub-tasks across 17 child-centered audio and speech datasets. We would highlight that the scope of the benchmark design is rarely presented in the literature. Evaluated on a representative range of speech and audio foundation models, ChildVox supports downstream applications such as characterizing children’s language levels and tracking speech production with age. In the future, we plan to expand it toward broader phoneme recognition [42], [43], additional recording environments [44][46], and to apply our benchmark models to enrich existing child audio corpora to support early screening [47], longitudinal monitoring [48], and education [49], [50] for children.

9 Limitations↩︎

While ChildVox facilitates the systematic evaluation of audio, speech, and LALMs on child-centered audios, several limitations remain.

9.0.0.1 Language, Demographics, and Task Coverage

The majority of datasets in ChildVox include most English-language recordings in speech categories, with limited representation of other languages or dialects. Specifically, our ASR evaluation is restricted to the English subset. Therefore, conclusions about model performance may not generalize to children speaking other languages, such as Mandarin [51] or Spanish [52]. Similarly, the demographic factors (e.g., educational background or developmental status) across the experimental datasets are, in many cases, not fully documented, which may introduce sampling biases that propagate into model evaluation and result interpretation. Moreover, we would consider integrating more spoken language understanding tasks as described in Dynamic-superb [53], [54].

9.0.0.2 Annotation variability

Many child-centered audio tasks in ChildVox, particularly affective vocalization classification (ReCANVo), cry-cause classification (Donate-a-Cry), and canonical syllables labeling (SpeechMaturity), involve inherently subjective categories with inter-rater disagreement. Therefore, reported classification scores may only reflect the ceiling imposed by annotation reliability.

9.0.0.3 Restricted set of foundation models

Although ChildVox includes several representative self-supervised, ASR-oriented, and LALMs, the development of audio foundation models is expanding rapidly. Specifically, we did not evaluate every recent open-source LALMs, such as GAMA [55], SALMONN [56], Step-Audio [57], and Kimi-Audio [58]. Moreover, comparisons with proprietary baselines are limited to two Gemini Flash models in a zero-shot setting, and our zero-shot prompt approach did not cover different prompt-engineering techniques.

10 Ethical Considerations↩︎

Overall, all models released in this work are derived from publicly available sources, and ChildVox adheres to the scope of each dataset’s license. Moreover, modeling children’s speech can expose sensitive information about children, and to mitigate misuse, we plan to release code and checkpoints under the Responsible AI License (RAIL). Users are expected to respect the privacy of data subjects and comply with applicable laws in their jurisdictions. We encourage the use of these models for research aimed at developing robust speech processing technology for and about children. Specifically, use of the models for clinical or diagnostic applications, surveillance, privacy-invasive applications, or commercial purposes is strictly prohibited.

11 Dataset Details↩︎

11.0.0.1 CirCor

The CirCor DigiScope Phonocardiogram Dataset [8] is one publicly available pediatric heart-sound corpus, which includes 5,272 phonocardiogram recordings from 1,568 subjects aged 0–21 years. Each subject is annotated by an expert cardiac physiologist with a murmur label drawn from, present, absent, or unknown. We use the publicly released training/testing partition to classify these tree labels. We segment each recording into non-overlapping 10-second windows. We further perform between-subject 5-fold cross-validation, partitioning the train data with 80/20 into train and development sets.

11.0.0.2 ICBHI

ICBHI [19] is a publicly available benchmark for respiratory sound analysis, targeting the recognition of crackles, wheezes, and respiratory conditions. It comprises 920 recordings collected from 126 subjects, accompanied by two complementary annotation schemes. The dataset provides cycle-level labels for 6,898 respiratory cycles, each categorized as normal, containing crackles, containing wheezes, or both. Similar to the CirCor dataset, we use the official ICBHI train/test split and apply 5-fold cross-validation over the training subjects to construct train/dev folds. For respiratory condition classification, recordings are segmented into 10-second windows with a 5-second hop, with diagnoses mapped to healthy, COPD, and other. For crackle and wheeze detection, we instead segment each recording according to the provided respiratory-cycle annotations. Segments shorter than 3 seconds (recording-level) or 0.5 seconds (cycle-level) are discarded.

11.0.0.3 SPRSound

SPRSound [20] is a paediatric respiratory sound database, designed for automatic diagnosis. It contains 2683 records and 9089 respiratory sound events collected from 292 participants and annotated by 11 annotators. This dataset categorizes respiratory sound into two major labels: normal and adventitious. Specifically, there are continuous adventitious (CAS, such as Rhonchi, Wheeze, and Stridor) and discrete adventitious (DAS, such as Coarse Crackle and Fine Crackle) based on duration. We use the official SPRSound train/test split and apply 5-fold cross-validation over the training subjects to construct training and development folds. For respiratory sound classification, we use the record-level annotations provided with each recording, which are normal, CAS, DAS, CAS & DAS, and poor quality.

11.0.0.4 Donate-a-cry

Donate-a-cry is a public infant cry audio corpus with labels for crying reasons. The audio samples in the corpus are collected with mobile devices and cleaned in the post-processing. There are 465 samples in the dataset, from children aged from 0 weeks to 2 years. The dataset is distributed with extreme imbalance crying reasons in belly_pain (16), burping (8), discomfort (27), tired (24), and hungry (382). Therefore, we simplify the label as hungry and other.

11.0.0.5 CryBank

CryBank [23] is collected for a longitudinal infant crying study. The samples are recorded from 24 families with 10 girls and 14 boys. In our study, we focus on three primary cry causes, which are hunger, loneliness, and discomfort, discarding segments labeled as pain, other, or unknown due to limited sample availability. We merge consecutive cry segments within each recording when the inter-segment gap is less than 1 second to form continuous cry bouts, and retain only merged segments with a duration of at least 3 seconds. The resulting dataset is partitioned at the subject level using 5-fold cross-validation, where within each fold 80% of the training subjects are used for training and the remaining 20% for validation.

11.0.0.6 AudioSet

AudioSet [21] is a large-scale ontology-driven collection of over one million YouTube video clips. In this benchmark, we focus on child sound event classification and filter data belonging to the following ten child-related audios: speaking, babbling, giggling, shouting, laughing, crying, whimpering, singing, playing, and music for children. We select the AudioSet-balanced dataset as the training set.

11.0.0.7 ReCANVo

ReCANVo [10] is designed to investigate the informative expressions in nonverbal vocalizations from real-world interaction. The participants are 6 males and 2 females, aged from 6 to 23, with different developmental disorders. The speech samples from each speaker have at least 10 pieces with fewer than 10 words. Overall, the data has 7,077 audio samples, recorded in a home environment. In our benchmark, we use six expression categories: delighted, dysregulated, frustrated, request, self-talk, and social, while discarding samples that fall outside this label set. The dataset is partitioned at the utterance level using 5-fold cross-validation, where within each fold 80% of the training and validation samples are used for training and the remaining 20% for testing.

11.0.0.8 BabbleCor

BabbleCor [24] is a crosslingual corpus of infants and children to investigate vocal development in typical speech. 52 participants are exposed to five different languages, including English, Spanish, Yêlí-Dnye, Tseltal Mayan, and bilingual Quechua-Spanish. Audio samples are clipped (400ms) from a day-long real-life recording for each child and annotated into five categories: canonical, non-canonical, crying, laughing, and junk. In our data processing, we use all five vocalization categories, crying, laughing, canonical, non-canonical, and junk, while discarding data marked as NO-LABEL. The dataset is partitioned at the subject level using 5-fold cross-validation, where within each fold 80% of the training subjects are used for training and validation, and the remaining 20% for testing.

11.0.0.9 SpeechMaturity

SpeechMaturity [25], [59] is also designed to investigate children’s vocal development. Following BabbleCor, SpeechMaturity acquires clips from day-long children’s speech recordings and gets clips annotated into the five categories. But SpeechMaturity is collected from a larger population size, containing 242,004 labeled vocalizations from more than 25 languages. Our data processing approach follows the same processing steps as stated in the BabbleCor dataset.

11.0.0.10 C-BESD

C-BESD [11] is a bilingual dataset developed for children’s emotion development research. The dataset contains 4200 utterances from 70 child speakers, aged from 6 to 12, and gender-balanced. Each speaker has 5 utterances for each of the 6 emotions (Happy, Neutral, Angry, Disgust, Fear, and Sad) in both Telugu and English. In our study, we focus on four primary emotion categories: angry, happy, neutral, and sad. This choice is similar to common emotion recognition benchmarks [60], [61]. The dataset is partitioned at the speaker level using 5-fold cross-validation, where within each fold 80% of the training subjects are used for training and validation, and the remaining 20% for testing.

11.0.0.11 PERCEPT-R

PERCEPT-R [26] focuses on the speech of children from fully-rhotic American English with Rhotacism Speech Sound Disorder (RSSD). PERCEPT-R contains 32.47 hours of citation speech recordings, comprising 105,232 word-level utterances. The 281 children (121 girls and 160 boys) are involved in speech recording, ranging from 72 months to 288 months. In data processing, we discard any utterances without a valid perceptual label. We restrict the data to recordings from the pre-treatment (PRE-1) and discharge post-treatment (DP-1) sessions to capture both ends of the intervention spectrum. The dataset is partitioned at the subject level using 5-fold cross-validation, where within each fold 80% of the training subjects are used for training and validation, and the remaining 20% for testing.

11.0.0.12 SpeechOcean762

SpeechOcean762 [28] is an open-source non-native English speech corpus for multi-granularity pronunciation assessment. It contains 5,000 utterances from 250 Mandarin-speaking learners (half children, half adults) with balanced gender and proficiency distributions. Each utterance is annotated by five experts to assess its accuracy, fluency, and prosody. Specifically, for accuracy and prosody, scores of 6 or below are mapped to “Poor”, 7–8 to “Nearly correct”, and 9–10 to “Correct”. For fluency, scores below 5 are mapped to “Intermittent”, 5–6 to “Fluent in general”, and 7 or above to “Fluent”. We adopt the official train and test splits, and further split the training set at the speaker level using 5-fold cross-validation.

11.0.0.13 UltraSuite

UltraSuite [29] is a multimodal repository of synchronized ultrasound and acoustic data from child speech therapy sessions, containing 86 children (58 typically developing, 28 with SSDs) across three sub-datasets totaling 37.28 hours of audio and 14,456 utterances. Data spans six prompt types and includes longitudinal recordings across therapy stages for children with disorders such as phonological delay and childhood apraxia of speech. In our benchmark, we focus on the typically developing children subset (core-uxtd) and restrict the data to Type A prompts, which consist of single-word utterances for word-level articulation analysis. We formulate a multi-label classification task targeting eight places of articulation, alveolar, bilabial, palatal, glottal, labiovelar, velar, dental, and labiodental, by mapping each target word to its constituent consonant phonemes via the CMU Pronouncing Dictionary. We adopt the official train, dev, and test partitions provided with the dataset.

11.0.0.14 Natural Language Sampling (NLS)

The dataset [31] comprises 73 English-speaking child-parent dyadic interaction sessions that occurred in naturalistic or home settings. No child appears more than twice in the dataset. Most children are diagnosed with ASD and are minimally verbal. Recordings were made remotely by research personnel over Zoom. Domain experts assigned each child’s spoken language levels to one of the three levels: pre-verbal (LL-1), first words (LL-2), or word combinations (LL-3). This definition was based on the prior study by  [41].

11.0.0.15 In-house ADOS2-Mod3

The ADOS [33] is a semi-structured assessment protocol lasting 40–60 minutes, where ADOS-Module 3 (Mod3) targets verbally fluent children across 14 pre-defined activities. The in-house ADOS-Mod3 dataset [32] focuses on two interview-style components, Social Difficulties and Annoyance and Emotions. This produces 352 segments from 180 children (mean age 8.53 years, range 3–13). Roughly half received an ASD diagnosis, with most others diagnosed with ADHD or other neurodevelopmental or psychiatric conditions.

11.0.0.16 MyST

MyST [12] is one of the largest children’s conversational speech corpora, containing 393 hours and 228,874 utterances from 1,371 third-to-fifth-grade students engaging in one-on-one spoken dialog with a virtual science tutor. The corpus includes a stratified train/dev/test split for the ASR experiment.

11.0.0.17 TinyVox

TinyVox [13] is a large-scale cross-linguistic corpus of over 500,000 IPA-transcribed child vocalizations (387.7 hours) from 560 children aged 5–96 months, compiled and standardized from 31 PhonBank corpora spanning English, French, Portuguese, German, and Spanish, with a unified 57-phoneme target inventory mapped via phonological feature edit distance. The dataset provides a speaker-independent 80/10/10 train/validation/test split. In this study, we only perform the phoneme experiments with LALM, since the token spaces of the decoders for the evaluated ASR models are not directly compatible with phoneme recognition targets.

12 Training Details↩︎

12.1 Training Parameters↩︎

For non-ASR tasks, during fine-tuning, SSAST models are trained with learning rates from \({1\times10^{-4}, 5\times10^{-4}}\), while other pre-trained speech models use learning rates from \({2\times10^{-4}, 5\times10^{-4}, 1\times10^{-3}}\). We use a learning rate of \(1\times10^{-4}\) in fine-tuning Qwen2Audio and AudioFlamingo3. All experiments are trained for 10 epochs. We set the maximum audio duration of 10 seconds for most audio and speech classification tasks. We apply a batch size of 256 with a mini-batch size of 8 when fine-tuning LALMs. During training, we applied several augmentations to the input waveforms: the Gaussian noise was added with a probability of 1.0, using an SNR range of 3–30 dB; time stretching was used with a probability of 1.0, with stretch rates ranging from 0.9 to 1.1; and polarity inversion was applied with a probability of 0.5.

For ASR fine-tuning, all models are trained with a learning rate of \(1\times10^{-4}\) and a batch size of 32. We train for 5 epochs on the MyST dataset and 20 epochs on the ADOS dataset. Prior to fine-tuning, all transcripts are normalized using Whisper’s BasicTextNormalizer, and the same normalization is applied to the model outputs during inference. Unlike the work in [1], we do not filter samples based on transcription length or the zero-shot Whisper-large-v3 WER, allowing for a more realistic and comprehensive evaluation of model performance.

12.2 Training Architecture↩︎

12.2.0.1 Encoder-based Fine-tuning

Prior works [60], [62] have shown that fine-tuning hidden representations from pre-trained encoder layers of speech foundation models can effectively adapt to a broad range of downstream speech tasks. Based on this, we adopt the architecture in Figure [fig:lora], where a learnable weighted average over all encoder layers (both convolutional and transformer) is passed through a 1D point-wise convolution, followed by temporal averaging, and fed to a fully connected classifier. We further apply LoRA [63] to all fine-tuning experiments for parameter-efficient adaptation of the foundation model backbones, including for the ASR fine-tunings.

12.2.0.2 LALMs

As stated in the main paragraph, we directly apply LoRA modules with rank 64 into the query, key, value, down-projection, and up-projection layers of the language model, and feedforward layers of the audio encoder. We use 4-bit quantization during training and normal 16-bit precision at inference.

12.3 Training Hardware↩︎

All experiments with public datasets are conducted on a High-Performance Computing (HPC) cluster. Most experiments are conducted with a single A40 or V100 GPU. Experiments with private data are conducted on a GPU server with 4 A6000 GPUs. Most experiments on encoder-based models require a fine-tuning time of less than 2 days, and experiments on LALMs require 4-5 days on an A40 GPU.

13 Prompt Design↩︎

While LALMs are trained on generalized speech tasks, they may not have the domain knowledge about vocalizations/speech patterns specific to children. Hence, in our SFT pipeline, we design system prompts for each dataset so as to include descriptions for each of the classes in the datasets. System prompt was first designed by the authors and then filtered using a frontier model (Gemini 3.1-pro). System prompts corresponding to the SFT of each dataset are as follows:

13.1 SpeechMaturity↩︎

Listen to the child audio recording and classify the vocalization type. Output only the label - no explanation or punctuation.

Labels: - Crying: distress vocalizations, whimpers, or wails - Laughing: laughter or giggling sounds - Canonical: well-formed syllables with clear consonant-vowel structure (e.g., “ba”, “da”) - Non-Canonical: vocalizations lacking clear consonant-vowel structure (e.g., squeals, growls, vowel-only sounds) - Junk: non-speech noise, silence, or unintelligible audio

13.2 BabbleCor↩︎

Listen to the child audio recording and classify the vocalization type. Output only the label - no explanation or punctuation.

Labels: - Crying: distress vocalizations, whimpers, or wails - Laughing: laughter or giggling sounds - Canonical: well-formed syllables with clear consonant-vowel structure (e.g., “ba”, “da”) - Non-Canonical: vocalizations lacking clear consonant-vowel structure (e.g., squeals, growls, vowel-only sounds) - Junk: non-speech noise, silence, or unintelligible audio

13.3 PERCEPT-R↩︎

Listen to the child reading a word. Evaluate the pronunciation of the ‘r’ sound specifically. Output only the label - no explanation or punctuation.

Labels: - Rhotic: the ‘r’ sound is produced correctly with appropriate rhoticity - Derhotic: the ‘r’ sound is missing, distorted, or replaced (e.g., substituted with ‘w’)

13.4 SpeechOcean762-Prosody↩︎

Listen to the child reading a sentence. Evaluate the prosody - focus on intonation, stress, and rhythm. Output only the label - no explanation or punctuation.

Labels: - Poor intonation: flat, monotone, or clearly unnatural intonation patterns - Nearly correct intonation: mostly natural prosody with minor errors - Correct intonation: natural, native-like intonation and rhythm throughout

13.5 SpeechOcean762-Fluency↩︎

Listen to the child reading a sentence. Evaluate the fluency - focus on flow, pausing, and hesitations. Output only the label - no explanation or punctuation.

Labels: - Intermittent or little influent: frequent pauses, hesitations, or disruptions in speech flow - Fluent in general: mostly smooth delivery with occasional minor disfluencies - Fluent: smooth, continuous speech with no notable disfluencies

13.6 SpeechOcean762-Accuracy↩︎

Listen to the child reading a sentence. Evaluate the pronunciation accuracy - focus on correct phoneme production. Output only the label - no explanation or punctuation.

Labels: - Poor or Understandable: many mispronunciations; difficult to understand - Good: mostly accurate pronunciation with some errors - Excellent: highly accurate, native-like pronunciation throughout

13.7 CirCor↩︎

Listen to the heart sound recording. Determine whether a cardiac murmur is present. Output only the label - no explanation or punctuation.

Labels: - Absent: no murmur detected; normal heart sounds - Present: a murmur is clearly audible - Unknown: audio quality or ambiguity prevents a confident determination

13.8 ICBHI-Record↩︎

Listen to the lung sound recording. Classify the respiratory condition based on the breath sounds. Output only the label - no explanation or punctuation.

Labels: - Healthy: normal breath sounds with no abnormalities - Obstructive: sounds consistent with airway obstruction (e.g., wheezing) - Infectious: sounds consistent with infection (e.g., crackles, consolidation)

13.9 ICBHI-Crackles↩︎

Listen to the lung sound recording. Determine whether crackles are present. Output only the label - no explanation or punctuation.

Labels: - Crackles: discontinuous, explosive sounds (fine or coarse crackles) are audible - No crackles: no crackle sounds detected

13.10 ICBHI-Wheezes↩︎

Listen to the lung sound recording. Determine whether wheezes are present. Output only the label - no explanation or punctuation.

Labels: - Wheezes: continuous, high-pitched whistling sounds are audible - No wheezes: no wheeze sounds detected

13.11 SPRSound↩︎

Listen to the lung sound recording. Classify the respiratory sound pattern. Output only the label - no explanation or punctuation.

Labels: - Normal: clear breath sounds with no adventitious sounds - CAS: continuous adventitious sounds only (e.g., wheezes, rhonchi) - DAS: discontinuous adventitious sounds only (e.g., crackles) - CAS & DAS: both continuous and discontinuous adventitious sounds present - Poor Quality: recording is too noisy or unclear for reliable classification

13.12 ReCANVo↩︎

Listen to the child audio recording and classify the communicative intent or emotional state of the vocalization. Output only the label - no explanation or punctuation.

Labels: - Delighted: joyful, excited, or positively expressive sounds - Dysregulated: distressed, overwhelmed, or emotionally dysregulated vocalizations - Frustrated: sounds expressing frustration or dissatisfaction - Request: vocalizations used to request an object, action, or attention - Selftalk: vocalizations directed inward, not toward a social partner - Social: vocalizations directed toward engaging or interacting with others

13.13 CryBank↩︎

Listen to the infant cry recording and identify the most likely cause of the cry. Output only the label - no explanation or punctuation.

Labels: - Hunger: cry pattern associated with feeding need - Loneliness: cry pattern associated with desire for social contact or comfort. Loneliness was defined as the infant seeking contact or attention - Discomfort: cry pattern associated with physical discomfort. Examples of discomfort included cold, fever, and full diaper.

13.14 MyST↩︎

Listen to the child audio recording and transcribe exactly what is spoken. Output only the transcript - no labels, explanations, punctuation beyond what is spoken, or formatting. If the audio is unintelligible, output: [unintelligible]

13.15 TinyVox↩︎

Listen to the child speaking a single word. Transcribe the pronunciation using standard phoneme symbols (ARPABET or IPA). Output only the phoneme sequence - no labels, explanations, or formatting. Example format: /b ae t/

14 Proprietary Baselines↩︎

We apply both Gemini 2.5 Flash and Gemini 3.5 Flash APIs on 5 selected public datasets in the main paragraph. We apply a temperature of 1.0 in prompting both models. We run the prompt experiment on these two proprietary baselines for a single run due to the budget limitation.

15 Generative AI Use Disclaim↩︎

Generative AI tools were used solely for grammar checking and polishing the writing of the manuscript. All research concepts, experimental design, implementations, analyses, and reported results were developed and conducted by the authors. The conceptual framing, methodology, and scientific contributions of this work are entirely human-generated.

References↩︎

[1]
R. Fan, N. B. Shankar, and A. Alwan, “Benchmarking children’s ASR with supervised and self-supervised speech foundation models,” Interspeech, 2024.
[2]
A. Ying et al., “Benchmarking training paradigms, dataset composition, and model scaling for child asr in espnet,” Workshop on Child Computer Interaction (WOCCI), 2025.
[3]
V. P. Singh, M. Sahidullah, and T. H. Kinnunen, “Causal analysis of ASR errors for children: Quantifying the impact of physiological, cognitive, and extrinsic factors,” Computer Speech & Language, vol. 95, p. 101859, 2026.
[4]
J. Li et al., “Automated analysis of naturalistic recordings in early childhood: Applications, challenges, and opportunities,” IEEE Signal Processing Magazine, vol. 42, no. 6, pp. 16–34, 2026.
[5]
A. K. Namasivayam, D. Coleman, A. O’Dwyer, and P. Van Lieshout, “Speech sound disorders in children: An articulatory phonology perspective,” Frontiers in psychology, vol. 10, p. 2998, 2020.
[6]
C. Lord, M. Elsabbagh, G. Baird, and J. Veenstra-Vanderweele, “Autism spectrum disorder,” The lancet, vol. 392, no. 10146, pp. 508–520, 2018.
[7]
N. K. Lesaux and L. S. Siegel, “The development of reading in children who speak english as a second language.” Developmental psychology, vol. 39, no. 6, p. 1005, 2003.
[8]
J. Oliveira et al., “The CirCor DigiScope dataset: From murmur detection to murmur classification,” IEEE journal of biomedical and health informatics, vol. 26, no. 6, pp. 2524–2535, 2021.
[9]
E. McGonigle, M. VanDam, C. Wilkinson, and K. T. Johnson, “Benchmarking automatic speech recognition technology for natural language samples of children with and without developmental delays,” in 2024 46th annual international conference of the IEEE engineering in medicine and biology society (EMBC), 2024, pp. 1–5.
[10]
K. T. Johnson, J. Narain, T. Quatieri, P. Maes, and R. W. Picard, “ReCANVo: A database of real-world communicative and affective nonverbal vocalizations,” Scientific Data, vol. 10, no. 1, p. 523, 2023.
[11]
D. V. Rao, K. Radha, S. R. Gudivaka, and P. Polipogu, “Children’s speech emotion recognition using feature fusion-based deep neural networks,” in 2024 IEEE international conference on intelligent signal processing and effective communication technologies (INSPECT), 2024, pp. 1–6.
[12]
S. Pradhan, R. Cole, and W. Ward, “My science tutor (myst)–a large corpus of children’s conversational speech,” in Proceedings of the 2024 joint international conference on computational linguistics, language resources and evaluation (LREC-COLING 2024), 2024, pp. 12040–12045.
[13]
M. Lavechin, E. Bergelson, and R. Levy, “BabAR: From phoneme recognition to developmental measures of young children’s speech production,” arXiv preprint arXiv:2603.05213, 2026.
[14]
J. Li, M. Hasegawa-Johnson, and N. L. McElwain, “Analysis of self-supervised speech models on children’s speech and infant vocalizations,” in 2024 IEEE international conference on acoustics, speech, and signal processing workshops (ICASSPW), 2024, pp. 550–554.
[15]
A. Xu, K. Huang, T. Feng, L. Shen, H. Tager-Flusberg, and S. Narayanan, “Exploring speech foundation models for speaker diarization in child-adult dyadic interactions,” in Proc. Interspeech 2024, 2024, pp. 5193–5197.
[16]
A. Xu et al., “Understanding spoken language development of children with asd using pre-trained speech embeddings,” Interspeech, 2023.
[17]
L. B. Medin, T. Pellegrini, and L. Gelin, “Self-supervised models for phoneme recognition: Applications in children’s speech for reading learning,” Interspeech, 2024.
[18]
A. Sinha, H. K. Kathania, and M. Kurimo, “A study on the layer-wise transferability of self-supervised learning features for children’s speech processing tasks,” Speech Communication, p. 103392, 2026.
[19]
B. M. Rocha et al., “An open access database for the evaluation of respiratory sound classification algorithms,” Physiological measurement, vol. 40, no. 3, p. 035001, 2019.
[20]
Q. Zhang et al., “Sprsound: Open-source SJTU paediatric respiratory sound database,” IEEE Transactions on Biomedical Circuits and Systems, vol. 16, no. 5, pp. 867–881, 2022.
[21]
J. F. Gemmeke et al., “Audio set: An ontology and human-labeled dataset for audio events,” in 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2017, pp. 776–780.
[22]
[23]
M. Lockhart-Bouron et al., “Infant cries convey both stable and dynamic information about age and identity,” Communications Psychology, vol. 1, no. 1, p. 26, 2023.
[24]
M. Cychosz et al., “Vocal development in a large-scale crosslinguistic corpus,” Developmental science, vol. 24, no. 5, p. e13090, 2021.
[25]
K. Hitczenko et al., “Speech maturity dataset: A cross-cultural corpus of naturalistic child and adult vocalizations,” OSF Preprints, 2025.
[26]
N. Benway, J. L. Preston, E. Hitchcock, A. Salekin, H. Sharma, and T. M. Byun, “PERCEPT-r: An open-access american english child/clinical speech corpus specialized for the audio classification of /r/.” in INTERSPEECH, 2022, pp. 3648–3652.
[27]
Y. Rose and B. MacWhinney, “The PhonBank project,” The Oxford Handbook of Corpus Phonology, pp. 380–401, 2014.
[28]
J. Zhang et al., “speechocean762: An open-source non-native english speech corpus for pronunciation assessment,” in Proc. Interspeech 2021, 2021, pp. 3710–3714.
[29]
A. Eshky et al., “UltraSuite: A repository of ultrasound and acoustic data from child speech therapy sessions,” in Proc. Interspeech 2018, 2018, pp. 1888–1892.
[30]
M. S. Ribeiro, J. Cleland, A. Eshky, K. Richmond, and S. Renals, “Exploiting ultrasound tongue imaging for the automatic detection of speech articulation errors,” Speech Communication, vol. 128, pp. 24–34, 2021.
[31]
L. K. Butler et al., “Remote natural language sampling of parents and children with autism spectrum disorder: Role of activity and language level,” Frontiers in Communication, vol. 7, p. 820564, 2022.
[32]
A. Xu, T. Feng, S. Bishop, C. Lord, and S. Narayanan, “End-to-end joint ASR and speaker role diarization with child-adult interactions,” arXiv preprint arXiv:2601.17640, 2026.
[33]
C. Lord et al., Autism diagnostic observation schedule (ADOS): manual. Western Psychological Services Los Angeles, 2008.
[34]
Y. Gong, C.-I. Lai, Y.-A. Chung, and J. Glass, “Ssast: Self-supervised audio spectrogram transformer,” in Proceedings of the AAAI conference on artificial intelligence, 2022, vol. 36, pp. 10699–10709.
[35]
A. Koudounas, M. La Quatra, S. M. Siniscalchi, and E. Baralis, “voc2vec: A foundation model for non-verbal vocalization,” in ICASSP 2025-2025 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2025, pp. 1–5.
[36]
S. Chen et al., “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing, vol. 16, no. 6, pp. 1505–1518, 2022.
[37]
D. Rekesh et al., “Fast conformer with linearly scalable attention for efficient speech recognition,” in 2023 IEEE automatic speech recognition and understanding workshop (ASRU), 2023, pp. 1–8.
[38]
H. Xu, F. Jia, S. Majumdar, H. Huang, S. Watanabe, and B. Ginsburg, “Efficient sequence transduction by jointly predicting tokens and durations,” in International conference on machine learning, 2023, pp. 38462–38484.
[39]
Y. Chu et al., “Qwen2-audio technical report,” arXiv preprint arXiv:2407.10759, 2024.
[40]
S. Ghosh et al., “Audio flamingo 3: Advancing audio intelligence with fully open large audio language models,” Advances in Neural Information Processing Systems, vol. 38, pp. 41819–41886, 2026.
[41]
H. Tager-Flusberg et al., “Defining spoken language benchmarks and selecting measures of expressive language development for young children with autism spectrum disorders,” Journal of Speech, Language, and Hearing Research, vol. 52, no. 3, pp. 643–652, 2009.
[42]
S. Bharadwaj et al., “PRiSM: Benchmarking phone realization in speech models,” arXiv preprint arXiv:2601.14046, 2026.
[43]
C. Guo et al., “HuPER: A human-inspired framework for phonetic perception,” arXiv preprint arXiv:2602.01634, 2026.
[44]
J. Li, M. Hasegawa-Johnson, and N. L. McElwain, “Analysis of acoustic and voice quality features for the classification of infant and mother vocalizations,” Speech communication, vol. 133, pp. 41–61, 2021.
[45]
T. Feng, A. Xu, X. Shi, S. Bishop, and S. Narayanan, “Egocentric speaker classification in child-adult dyadic interactions: From sensing to computational modeling,” in Proc. Interspeech 2025, 2025, pp. 2835–2839.
[46]
B. Long et al., “The BabyView camera: Designing a new head-mounted camera to capture children’s early social and visual environments,” Behavior Research Methods, vol. 56, no. 4, pp. 3523–3534, 2024.
[47]
B. Islam et al., “Preliminary technical validation of LittleBeats™: A multimodal sensing platform to capture cardiac physiology, motion, and vocalizations,” Sensors, vol. 24, no. 3, p. 901, 2024.
[48]
E. Kalenkovich et al., “A year of nouns from english-learning infants’ daily lives: The SEEDLingS-nouns dataset: E. Kalenkovich et al.” Behavior Research Methods, vol. 57, no. 11, p. 298, 2025.
[49]
S. Li et al., “K-function: Joint pronunciation transcription and feedback for evaluating kids language function,” in ICASSP 2026-2026 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2026, pp. 3231–3235.
[50]
L. Wagner et al., “The ohio child speech corpus,” Speech Communication, vol. 170, p. 103206, 2025.
[51]
V. Y. Chua et al., “MERLIon CCS challenge: A english-mandarin code-switching child-directed speech corpus for language identification and diarization,” in Proc. Interspeech 2023, 2023, pp. 4109–4113.
[52]
C. E. Gildersleeve-Neumann, E. S. Kester, B. L. Davis, and E. D. Peña, “English speech sound development in preschool-aged children from bilingual english–spanish environments,” Language, Speech, and Hearing Services in Schools, vol. 39, no. 3, pp. 314–328, 2008.
[53]
C. Huang et al., “Dynamic-superb phase-2: A collaboratively expanding benchmark for measuring the capabilities of spoken language models with 180 tasks,” in International conference on learning representations, 2025, vol. 2025, pp. 82908–82973.
[54]
C. Huang et al., “Dynamic-superb: Towards a dynamic, collaborative, and comprehensive instruction-tuning benchmark for speech,” in ICASSP 2024-2024 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2024, pp. 12136–12140.
[55]
S. Ghosh et al., “Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities,” in Proceedings of the 2024 conference on empirical methods in natural language processing, 2024, pp. 6288–6313.
[56]
C. Tang et al., “Salmonn: Towards generic hearing abilities for large language models,” in International conference on learning representations, 2024, vol. 2024, pp. 16607–16629.
[57]
A. Huang et al., “Step-audio: Unified understanding and generation in intelligent speech interaction,” arXiv preprint arXiv:2502.11946, 2025.
[58]
D. Ding et al., “Kimi-audio technical report,” arXiv preprint arXiv:2504.18425, 2025.
[59]
T. Zhang, M. Suresh, A. Warluamont, K. Hitczenko, A. Cristia, and M. Cychosz, Employing self-supervised learning models for cross-linguistic child speech maturity classification,” in Interspeech 2025, 2025, pp. 2825–2829.
[60]
L. Pepino, P. Riera, and L. Ferrer, “Emotion recognition from speech using wav2vec 2.0 embeddings,” in Proc. Interspeech 2021, 2021, pp. 3400–3404.
[61]
T. Feng and S. Narayanan, “Foundation model assisted automatic speech emotion recognition: Transcribing, annotating, and augmenting,” in ICASSP 2024-2024 IEEE international conference on acoustics, speech and signal processing (ICASSP), 2024, pp. 12116–12120.
[62]
T. Feng et al., “Vox-profile: A speech foundation model benchmark for characterizing diverse speaker and speech traits,” arXiv preprint arXiv:2505.14648, 2025.
[63]
E. J. Hu et al., “Lora: Low-rank adaptation of large language models.” ICLR, vol. 1, no. 2, p. 3, 2022.