October 08, 2025
Current large language models (LLMs) and spoken language models (SLMs) begin thinking and taking actions only after the user has finished their turn. This disables the model from interacting with the user during the user’s turn and can lead to a high response latency for waiting for the model to think. Consequently, thinking after receiving the full input is not suitable for speech-to-speech interaction, where real-time and low-latency interaction is important. We address the above issue by drawing inspiration from the fact that humans can naturally “think while listening”. In this paper, we propose Shanks, a general inference framework that enables SLMs to generate unspoken chain-of-thought reasoning when listening to the user input. Shanks streams the input speech in fixed-duration chunks and, as soon as a chunk is received, generates unspoken reasoning based on all previous speech and reasoning, while the user continues speaking. Shanks uses unspoken reasoning to determine whether to interrupt the user and make tool calls to complete the task. We demonstrate that Shanks enhances the real-time user-SLM interaction in two scenarios: (1) When the user is presenting their step-by-step solution to a math problem, Shanks can listen to and reason over the user’s speech and make an interruption when the user makes a mistake. Shanks interrupts the user 37.1% more accurately compared with a baseline that interrupts the user without thinking. (2) In a tool-augmented dialogue scenario, where the model needs to make tool calls to achieve the user’s request, Shanks can complete 56.9% of the tool calls before the user even ends their turn. Overall, Shanks is a step toward models that keep thinking throughout the conversation, not only after a turn ends. Animated illustrations of Shanks can be found at https://d223302.github.io/SHANKS/.
In recent years, the thinking process has been used to improve Large Language Models (LLMs), where the LLM first generates a hidden chain-of-thought (CoT) reasoning [1], [2] invisible to the users, and then generates the final output response [3], [4]. This thinking process improves LLMs on reasoning-intensive tasks, including mathematics [5], coding [6], and questions that involve significant domain knowledge [7]. However, current reasoning LLMs only start to think after receiving the complete user input, which is reasonable for turn-based interactions, i.e., the model processes the user’s message after it is fully composed and sent.
In contrast, human behavior in spoken communication is different. Humans naturally think while listening, far before the speaker finishes their turn [8]–[10]. Thinking during listening offers two key advantages: (1) It enables timely and well-founded reactions, including interruption, even before the speaker’s turn ends. (2) It reduces response latency by allowing answer preparation to begin before the speaker finishes speaking. Motivated by these observations, we propose a method to enable spoken language models (SLMs) to think while listening to input speech.
In this paper, we introduce Shanks: Simultaneous Hearing and Thinking with Chunked Input Speech. Shanks is a general inference framework for both end-to-end (E2E) and cascade SLMs to achieve thinking while listening. At inference time, Shanks processes the user input in a fixed-size chunk. Once a chunk of speech input is received, Shanks generates a chunk of unspoken thinking tokens based on all previous input speech chunks and previous thinking chunks. Shanks alternates between receiving the input speech chunk and generating an unspoken CoT reasoning when the user is still emitting the next speech chunk, achieving the thinking while listening. During the thinking process, Shanks can decide to interrupt the user or make tool calls to prepare for the final spoken response. The inference workflow of Shanks is depicted in Figure 1.
We use two scenarios to show how Shanks can improve real-time user-SLM interaction. First, we study a scenario where the user first describes a math question and then describes their step-by-step solution. Shanks can listen to the user’s problem-solving process and perform internal thinking in the meantime to interrupt the user when the user makes a mistake in their solution. This scenario has great potential in educational use cases, where the SLM serves as a tutor to guide the student. Compared to a baseline that makes an interruption without thinking, Shanks interrupts 71% more when the user makes a mistake, while the interruption made by Shanks is 37.1% more valid.
Next, we focus on a task-oriented dialogue setting, where the user requests the SLMs to assist with a travel plan, and the model needs to call Booking.com APIs to complete the user’s request and respond to the user. Without Shanks, all the tool calls can only be made after the user’s speech ends, resulting in a higher response latency. We use Shanks to make tool calls when the user is still speaking. Shanks enables the model to successfully complete 56.9% of API calls while the user is still speaking, reducing the latency of the final response.
We summarize our contribution as follows:
We propose Shanks, a general framework for SLMs that enables thinking while listening. To the best of our knowledge, we are the first to explore generating unspoken CoT reasoning when the user is still speaking.
We show that Shanks can interrupt the user more accurately compared to a baseline that interrupts without thinking.
Using Shanks, the SLM can successfully make tool calls before the user even finishes talking.
Current LLMs and SLMs only start to think after the user’s input is completed. In contrast, humans can think while listening [8]–[10], where we reason over what we just heard, guess what the speaker might be up to, and prepare the necessary ingredients to cook up a good response. Thinking while listening allows us to interact with the speaker better when the speaker is still speaking. In this section, we introduce Shanks, a general framework to make SLMs capable of thinking while listening. Here, we only discuss the basic form of Shanks, and we defer the more advanced usages, including interruption or tool call, to later sections.
During inference, Shanks requires that the user’s input speech comes in a streaming manner. Shanks processes the streaming user input speech by a fixed chunk size of \(t_{\rm chunk}\) seconds. We use \(S_i\) to denote the \(i\)-th user input speech chunk, where \(S_i\) is an audio chunk of \(t_{\rm chunk}\) seconds, except for the last chunk \(S_{N}\), which may be shorter. When the user is still speaking, Shanks alternately takes the user speech \(S_{i}\) and generates the thinking chunks \(R_{i}\) conditioning on all previous user speech and all previous thinking chunks. The thinking chunks \(R_i\) are not spoken by the SLM; they only serve as the internal reasoning process of the SLM.
Here, we walk through what happens for Shanks during inference. The following contents are best read with Figure 1. At \(t=0\), the user begins to talk. When
\(t=t_{\rm chunk}\), the user speech from \(0\) to \(t_{\rm chunk}\), i.e., \(S_1\), is sent to the SLM. Here, we append a
special token [EOPA] (end of partial audio) after \(S_1\) to let the SLM know that this is the end of a chunk of partial user speech. Based on \(S_1\), the SLM
generates the first thinking chunk \(R_1\). A thinking chunk is enclosed in two special tokens <think> and </think>. The SLM generates \(R_1\) during
the interval \(t=t_{\rm chunk}\) to \(t=2t_{\rm chunk}\), and the user is still speaking the second chunk \(S_2\) at the same time.
Since \(t_{\rm chunk}\) is the time for the SLM to generate its thinking, the duration of \(t_{\rm chunk}\) cannot be selected too small; otherwise, the SLM may not be able to produce
meaningful thinking chunks. The number of thinking tokens in \(R_i\) is restricted to no more than \(t_{\rm chunk}\times n_{\rm tps}\), where \(n_{\rm tps}\)
is the number of tokens the model can generate per second. Unless specified, we select \(t_{\rm chunk}=4.0s\) in our paper; a 7B model can generate around 320 tokens on a single A100 GPU in this duration. At the end of
\(t_{\rm chunk}\), if the thinking chunk has not finished generating, i.e., the </think> token has not been emitted, we directly stop generating and append the </think> token at the end
of the thinking chunk.
At \(t=2t_{\rm chunk}\), we take the freshly obtained user speech chunk \(S_2\) (from \(t=t_{\rm chunk}\) to \(t=2t_{\rm
chunk}\)) and pass this chunk to the SLM, and again appending the [EOPA] after this chunk to generate the next thinking chunk. (Assume that the user still has not ended their turn at \(t=2t_{\rm
chunk}\).) When generating the next thinking chunk \(R_2\), the SLM conditions on \(S_1\), \(R_1\), and \(S_2\). The
SLM will continue the process of taking user input speech chunks and generating the thinking chunks until the user ends their speech in the \(N\)-th chunk of speech, \(S_N\). After the user
ends their speech, we feed the last speech chunk \(S_{N}\) into the SLM, while this time we append a different special token [EOA] (end of audio), indicating that the user’s speech has ended. Based on \(S_N\) and all the previous interleaved speech/thinking chunks \(\{S_1, R_1,\cdots,S_{N-1},R_{N-1}\}\), the SLM generates the thinking chunk \(R_N\) and then
generates a final response chunk \(O\). Only the final response \(O\) will be spoken by the SLM.
Since Shanks chunks the user input using a fixed-duration chunk \(t_{\rm chunk}\), the model’s thinking will lag behind the user’s speech by at least \(t_{\rm chunk}\) seconds. If the user’s speech is less than \(t_{\rm chunk}\), Shanks cannot think while listening. However, since long speech can easily happen in real-world interaction, this limitation might not be a significant weakness of Shanks.
During inference, Shanks requires the SLM to generate thinking chunks based on all previous user input speech chunks and the model’s own thinking chunks. During training, we prepare datasets to make the SLM learn this behavior. Assume that we have a complete user speech \(S\), we can segment it into \(N\) chunks \(\{S_1,\cdots,S_N\}\) with a fixed duration \(t_{\rm chunk}\). Next, assume that we use some method to obtain the thinking chunks \(\{R_1,\cdots,R_N\}\) and the output response \(O\); we will explain how to obtain them in later sections. After preparing the training data, we use the standard language modeling cross-entropy loss to train the SLM to predict \(R_1\) given \(\{S_1\}\), predict \(R_2\) given \(\{S_1,R_1, S_2\}\), etc., and predict \(R_N\) and \(O\) given \(\{S_1,R_1,\cdots,S_{N-1}\}\). A full training sequence is depicted in Figure 2 (a).
After introducing the basics of Shanks, we use two tasks to demonstrate how Shanks can be applied. In this section, we explain the setup of the two scenarios and how Shanks works at inference time.
In the first scenario, we aim to use Shanks to make SLMs able to interrupt the user when the user is saying something wrong. The significance of this application lies in its potential in educational use cases, where the SLMs can serve as a tutor and listen to the speaker, a student, describing how they solve a problem. The SLM can make a timely interruption to let the student know that they are making a mistake, allowing them to correct it as early as possible.
Interruption is related to the full-duplex ability of spoken language models [11]. While most prior works on full-duplex SLMs focus on user interrupting SLMs, we focus on the reverse scenario: SLM interrupting users. As an important note, we do not advocate that it is good to have a model that interrupts the user. Some users might find it annoying and unpleasant when interrupted by an SLM. Our goal is a modeling mechanism that enables interruption, leaving model deployer policies and users to decide when (or whether) to turn it on.
We explain the precise task we are studying. In this task, the user describes a math problem and then solves the problem. The user’s solution does not simply state the answer; the user describes a step-by-step problem-solving process, which might be correct or wrong. The SLM needs to interrupt the user when the user is making a mistake, and not to interrupt the user when the solution is correct.
As this is a novel task and there is no available data, we built the evaluation data ourselves. First, we construct the user speech. A single user speech includes (1) a math question and (2) a step-by-step solution. We source the math questions from the testing data of GSM8K [12], a grade-school math word problem dataset commonly used for evaluating mathematical reasoning ability [1], [2], [13]. Next, we use two LLMs, Llama-2-7B [14] and Llama3.1-8B [15], to generate step-by-step answers for those questions, and use GPT-4o [16] to determine if the answer generated by the two models matches the ground truth answer in the dataset. We select these two models since they can generate CoT reasoning to solve the math problem, and their performance on GSM8K is very different: Llama-2-7B is a weaker model and prone to generating wrong solutions, while Llama-3.1-8B is a stronger model, which can generate more accurate solutions.
After we have the texts for the step-by-step solution, we convert them into speech. We use GPT-4o to rewrite the answers generated by the two Llama models to make the solution more colloquial. Next, we concatenate the original question, the colloquial step-by-step answer, and prepend a prefix "I want to solve the following question." to form the transcription of a testing instance. We use GPT-4o-mini-TTS [17] to synthesize the speech.
The final testing dataset includes 1280 instances with correct solutions and 1140 with incorrect solutions. We call the former the "correct subset" and the latter the "incorrect subset. The average duration of the user’s speech is around 49.25 seconds.
To teach the model to think while listening and determine whether to interrupt, the training data in this task include two types of instances: (1) The user provides a correct step-by-step solution to the question, and the model does not interrupt the user during the user’s speech. After the user finishes the speech, the output response acknowledges the correctness of the answer. (2) The user’s turn unfolds an erroneous problem-solving process, and the model interrupts the user when the user makes the first mistake and clearly explains what is wrong.
To construct such a training dataset, we use the math questions in Tulu3-Persona-Math-Grade [18] to construct the user speech \(S\) (including the question and the step-by-step solution) following the previously described procedure, and then segment the speech by a fixed duration \(t_{\rm chunk}=4\) seconds to obtain \(\{S_{1},\cdots,S_{N}\}\).
Next, we use GPT-4o to generate the thinking chunk \(R_i\). When generating the \(i\)-th thinking chunk \(R_i\), the input to GPT-4o includes the
transcriptions of all previous user speech chunks \(\{S_1,\cdots,S_{i}\}\) and all previous thinking chunks \(\{R_1,\cdots,R_{i-1}\}\). GPT-4o is required to do the following in the thinking
chunk: (1) Track the information already known and calculate intermediate variables when they are available. (2) Identify if any errors are in the current user’s transcription. If there is an error, GPT-4o should generate a [INTERRUPT] token
at the end of the thinking chunk, indicating that the user should be interrupted. We give GPT-4o four in-context examples to allow GPT-4o to understand the task. The prompt to GPT-4o is given in Table ¿tbl:tab:thinking-while-listening-1?
and ¿tbl:tab:thinking-while-listening-2? in the Appendix.
After generating the thinking chunks, we generate the final output response \(O\). For the user speeches with an error-free solution, the output response simply needs to let the user know that their solution is correct. We prompt GPT-4o to generate the final response based on the transcription of the full user speech and all previous thinking chunks. Now, we can form a training data sequence by interleaving \(S_i\) and \(R_i\) and then appending \(O\) in the end.
For those user speeches with a wrong solution, the output response will be an interruption to the user’s speech. Assume that based on our previous process for generating the \(R_i\)’s, GPT-4o decides to interrupt the
user after the user speech chunk \(S_k\), i.e., the thinking chunk \(R_k\) includes the interruption token [INTERRUPT]. To generate a response for interruption, we give GPT-4o
the user’s speech up to the \(k\)-th user speech chunk and all the previous thinking, and ask GPT-4o to generate a response \(O\) to interrupt the user. The interruption should be precise on
what error is made by the user and how to correct it. After this process, we can interleave \(S_1\) to \(S_k\) with \(R_1\) to \(R_k\) and append \(O\) in the end to form a training sequence. A figurative illustration of this training instance is shown in Figure 2 (b). Note that
in the last thinking chunk \(R_{k}\), there will be a special token [INTERRUPT], indicating that the model is going to interrupt the user.
During inference, we stream the user speech to the SLM by a fixed chunk size \(t_{\rm chunk}\), and follow the inference procedure elaborated in Section 2.1. If the SLM
generates the special token [INTERRUPT] in a thinking chunk \(R_k\) and outputs a response chunk \(O\) (when the user is emitting speech chunk \(S_{k+1}\)), we convert the response token into speech to interrupt the user. Once an interruption happens, the future user speech chunks will not be streamed to the SLM.
We evaluate a model on the testing dataset constructed in Section 3.1.1. We use the following metrics:
Interrupt ratio: The ratio of total interrupted instances among the total instances. A good model should have a low interrupt ratio on the correct subset and a high interrupt ratio on the wrong subset.
Valid interrupt ratio: This is used to evaluate whether an interruption made by the model is valid, and the valid interrupt ratio is defined as the ratio of valid interruptions among the total interruptions. To judge whether an interruption response \(O\) is valid or not, we use LLM-as-a-judge [19], [20]. We give the judge LLM (GPT-4o) the transcription of the user input until the time of interruption3 and output response \(O\) from the model, and ask the LLM judge to determine if the model’s interruption response \(O\) correctly interrupts the user when there are unclear or mistakes in the user’s speech. The prompt given to the judge model is shown in Table ¿tbl:tab:assistant-interruption-judgement? in the Appendix.
Interruption latency: The time of the model interruption compared to the time when the first error happens in the user’s speech, denoted as \(t_{\rm error}\). This metric is only used to evaluate the incorrect subset. For samples in the incorrect subset, we use GPT-4o to determine \(t_{\rm error}\). The details on determining \(t_{\rm error}\) are included in Appendix 10.1. Assume that the model interrupts the user at \(t_{\rm interrupt}\), then the interruption latency is calculated as \(t_{\rm interrupt} - t_{\rm error}\), where \(t_{\rm interrupt}\) is the time when the model starts to emit the first package of the speech for the interruption output \(O\).
In the previous scenario, we have introduced Shanks can be used to think when listening to interrupt the user when necessary. However, the "thinking" of a model can not only include the CoT reasoning generated by the model itself; the model can use external tools in their thinking process to complete the user’s request and improve the answer quality [21]–[25]. In the second scenario, the SLMs will use external tools in their thinking process to achieve the user’s request.
This kind of tool-augmented generation has been widely explored in text-in-text-out LLMs [23]. Given a user request, a tool-augmented LLM selects relevant tools such as calculators [21], search engines [26], and other APIs [22], and integrates the tool execution results into its reasoning process. However, current tool-augmented LLMs begin calling the tools after the full user input is received, which adds delay while the model invokes tools, waits for results, parses them, and composes a response. This delay may be acceptable in text-only settings, but in spoken dialogue, this latency breaks conversational flow.
In the second task, we will use Shanks to make tool calls when the user is still speaking, thus reducing the response latency due to making tool calls. As a proof of concept, we consider a task-oriented dialogue where the SLM serves as a travel-agency agent and is given a set of APIs it may call to complete the task. For example, the user might say, "Help me check the details of the cheapest flight from Hangzhou to Seoul on December 10, 2024, and the car rental information near Seoul airport." The SLM makes API calls to resolve airport names or codes, search for flights, and then query car-rental options, and finally composes the results into a single reply.
Given a user query like the example above, once the destination and date are clear, the flight search API can be called even when the user is only halfway through speaking. This is where Shanks can be useful: processing partial user input and performing early actions. Next, we formally introduce the task we study and how we evaluate it.
To show that SLMs can make tool calls when listening to the user request, we adopt ComplexFuncBench [27], an evaluation dataset that assesses LLMs’ ability to make multi-step API calls. An instance in ComplexFuncBench includes a user query in text specifying some requirements for a travel plan and a list of tool descriptions that are required to complete the task. Some example tools include APIs for searching hotels or flights.
An evaluated model needs to make relevant API calls and provide the resulting information to the user. A single user query typically requires multiple API calls, and these calls may have dependencies in which the output of one call becomes an argument to a subsequent call, so the call order matters and some calls may not be run in parallel. For each instance, the dataset provides the ground truth API calls and their corresponding API responses, which can be used to evaluate whether the model’s API call is correct. To adapt ComplexFuncBench to our spoken-dialogue setting, we use GPT-4o-mini-TTS to synthesize the user’s spoken query from the text instructions from ComplexFuncBench.
To train Shanks to perform Tool calls when listening, we need to teach the model to make API calls in their thinking process \(R_i\) based on user input speech chunk \(S_i\). We split half of the instances in ComplexFuncBench to construct the training data and the other half as the testing data.4 The speech chunks \(S_i\) in the training data can simply be obtained from segmenting the audio of the user query speech.
The next step is to construct the thinking chunks \(R_i\). In this task, the thinking chunk \(R_i\) is the API calls and call responses. For each user query, ComplexFuncBench already provides the ground truth API calls to complete the task, and we only need to determine which API calls, among the ground truth API calls, can be made after a speech chunk \(S_i\). To determine which API calls can be made in \(R_i\), recall that a thinking chunk \(R_i\) is based on the user speech up to \(i\times t_{\rm chunk}\) seconds, so as long as the speech up to \(i\times t_{\rm chunk}\) provides the information to make an ground truth API call, that API call can be included in \(R_i\). We follow the above idea and prompt GPT-o1 [3] with the transcription of the user speech, the time alignment of each word in the user’s speech, and the ground truth APIs, and ask GPT-o1 to determine the earliest time an API call can be made. The prompt is shown in Table ¿tbl:tab:earliest-tool-call? in the Appendix. Based on the above process, we can determine which ground truth APIs should be included in which thinking chunk \(R_i\). A thinking chunk \(R_i\) can have more than one API call and response. If no API calls can be made in \(R_i\), we put a template message that says there are no additional tool calls that can be made currently.
The final response \(O\) is also generated with GPT-4o by prompting it to generate a final response based on the user query, all the ground truth API calls, and the corresponding responses. During training, the descriptions of the API calls necessary to complete the user’s request will be included in the system prompt to let the model know what APIs can be used.
An illustration of a training instance is shown in Figure 2 (c). During training, the loss for the API call response in \(R_{i}\) is masked. Training on this dataset will teach the model to make API calls based on incomplete user queries as long as the information for an API call is sufficient.
During inference, for a testing instance, the model is given the user’s speech in a streaming manner; the descriptions of the APIs that can be used in this testing instance. When the model makes an API call, we use GPT-4o as a judge to determine if the API call matches one of the ground truth API calls, and return the response of the ground truth API call if a match is found. If the API call does not match the API call in the ground truth, we return a generic error message. Using GPT-4o to match the API call made by the model against the ground truth follows one of the evaluation protocols in the original ComplexFuncBench [27].
The testing set includes 500 instances, and each testing instance requires 5.1 API calls to complete on average. The average duration of the audio of the user’s speech is 18.71 seconds. We evaluate the performance using five metrics:
Call accuracy: The number of ground truth API calls that are successfully made by the model, divided by the total number of ground truth API calls. We also calculate the early call accuracy, defined as the ground truth API calls that are successfully made when the user is still speaking, divided by the total number of ground truth API calls. Similarly, we calculate the late call accuracy, where the dividend is the ground truth API calls that are successfully made after the user finishes speaking. This helps us understand how well the model uses the time when the user is still speaking; this can be used as a proxy to measure latency.
Success rate: The percentage of user queries that are successfully completed. If all the ground truth API calls for a user query are successfully made, the task is considered successful.
Correctness: We evaluate if the final response \(O\) is accurate and aligns with the API call responses. Following [27], we use GPT-4o to give a score in \(\{0,1,2\}\), indicating if the transcription of the response \(O\) is completely incorrect, partially correct, or completely correct the user’s request.
Completeness: We evaluate if the final response \(O\) fully satisfies the user’s request. Following [27], we use GPT-4o to give a score in \(\{0,1,2\}\), indicating if the transcription of the response \(O\) does not satisfy, partially satisfies, or fully satisfies the user’s request.
Latency: We evaluate the response latency by the number of tokens the model generates before it emits the final response \(O\) and after the user has finished their turn. We choose to report the token count instead of the exact time since the precise time to emit these tokens depends on hardware and implementation.
In this section, we introduce the models that we will compare in our experiments, including two variants of the Shanks models. Additionally, for each task in Section 3.1 and 3.2, we include a scenario-specific baseline model. The training details and hyperparameters are included in Appendix 9.
We fine-tune an end-to-end (E2E) SLM to make it able to think while listening. We will use Qwen-2.5-Omni (Qwen-omni for short) [28], one of the best open-sourced end-to-end SLMs, in our experiment. Qwen-omni is a thinker-talker SLM. The thinker takes speech representation extracted by a speech encoder [29] as the input and generates text tokens. The talker functions like a text-to-speech (TTS) model, taking the hidden representation from the thinker as the input and generating the output speech.
Originally, Qwen-omni is not capable of performing unspoken thinking – every token (and its corresponding hidden representation) generated by the thinker will be sent to the talker model and synthesized into speech. After fine-tuning the
thinker model on the training dataset mentioned before, the model will learn to enclose the unspoken thinking process within <think> and </think>. Since we only want the Qwen-omni to speak out the response tokens, we
only pass the response tokens and their hidden states to the talker.
As stated in Section 2.1, Shanks uses the time when the user is speaking the next speech chunk \(S_{i+1}\) to generate the thinking chunk \(R_i\), so the number of thinking tokens in \(R_i\) cannot exceed \(t_{\rm chunk}\times n_{\rm tps}\), where \(n_{\rm tps}\) is the number of tokens the model can generate per second.5
We set up a cascade version of Shanks. Precisely, we cascade an ASR (Whisper-large-v3 [30]) with a stronger text-only LLM, Qwen-2.5-7B-Instruct [31], to make the LLM generate thinking chunks while reading the partial transcription. Qwen-2.5-7B-Instruct and Qwen-omni are fine-tuned from the same base model, while Qwen-2.5-7B-Instruct are fine-tuned on a much larger reasoning dataset. This baseline allows us to know what the performance of Shanks can be if we use a model with better reasoning ability as the backbone. The training data of Shanks-E2E and Cascade are almost the same, only differing in the input modality.
We fine-tune a baseline model using Qwen-omni, which we name "No-thinking". We fine-tune it to predict whether it should interrupt the user without any thinking. The model is trained to predict special tokens,
[NO_INTERRUPT] or [INTERRUPT], to indicate whether the model should interrupt the user, given chunked user input speech. This can be thought of as Shanks while the thinking chunks only contain a
[NO_INTERRUPT] or [INTERRUPT] special token.
We do not compare with other models since there are no other models that can interrupt the user. While some full-duplex SLMs should be able to interrupt the user by design [32], we find that these models cannot interrupt the user at all. We also find that closed-source models like GPT-4o cannot interrupt the user when the user is still talking.
For this baseline, we fine-tune a model using Qwen-omni that takes the full user speech and then makes all the API calls. We call this model "Call-after-listen". This serves as a baseline that waits until receiving the full user input and then begins to make API calls; this is how existing tool-augmented models operate. During inference, this model takes the complete user query and then iteratively makes API calls and receives the responses until the model thinks all necessary API calls are made, and then generates the final response.
The results for interrupting the user turn are presented in Table ¿tbl:tab:interruption32results?. We have the following observations.
Shanks is more likely to interrupt on the wrong subset. Comparing the interruption ratio of Shanks on the correct and wrong subsets, the interruption ratio is 54.2% higher on the wrong subset. This shows that Shanks is indeed capable of capturing the errors in the user’s speech and interrupt appropriately. Based on the valid interruption ratio for the wrong subset, about 2 out of 3 interruptions made by Shanks are valid. Interesting, on the correct subset, the valid interruption ratio is non-zero. By looking into the instances in the correct subset, we find that even if their final answers are correct, sometimes their intermediate reasoning may be odd or ambiguous, and the model will interrupt and ask for clarification. Prior works also reported that even if the final answer of the model is correct, the CoT reasoning may be wrong [33]. In this case, the LLM judge treats this kind of interruption as valid.
Shanks’ interruption latency shows that the model mostly interrupts after the error occurs. On the wrong subset, the interruption latency of Shanks-E2E is 5.08 seconds on average. In Figure 5 in the Appendix, we further plot the distribution of the interruption latency. We find that the interruption latency on the wrong dataset is left-skewed, where more samples fall on the right proportion of the distribution and have a positive interruption latency. This indicates that most interruption happens later than the first error.
A qualitative example in Figure 3 shows that Shanks-E2E can interrupt the user when there is a mistake. To allow the readers to have a better idea of how Shanks interrupts the user, we show an example in Figure 3. When the user unfolds the question, Shanks already starts thinking about the math question. For example, when \(t\in[4t_{\rm chunk},5t_{\rm chunk}]\), the model’s thinking already calculates several intermediate variables, including the number of total petunias and sweet potato vines. When the user finishes stating the question after \(t=6t_{\rm chunk}\), the model already completes the calculation and has the correct answer in its mind. The model also correctly identifies the user’s error in falsely calculating that there are 25 petunias and interrupts the user during \(t\in[12t_{\rm chunk},13t_{\rm chunk}]\).
Interruption without thinking leads to much poorer performance. The performance of the no-thinking baseline is much worse than Shanks, which performs reasoning before interrupting. The no-thinking baseline has a much lower interruption ratio on the wrong subset, and the valid interrupt ratio is also much lower than Shanks. This shows that thinking before interruption is important, justifying the design of Shanks.
Cascade version of Shanks with stronger LLM leads to the best performance. When using Qwen-2.5-7B-Instruct as the backbone model for Shanks, the performance can be even better. The interruption ratio on the correct subset is lower, and the valid interruption ratio on the wrong subset also grows higher. This shows that the interruption ability of Shanks is mostly related to the reasoning ability of the backbone model, and using a stronger reasoning LLM can improve the performance.
4pt
max width=
Varying \(t_{\rm chunk}\) at inference time does not significantly affect the performance. When constructing the training data, we fix \(t_{\rm chunk}=4\) seconds. Here, we ask whether we can vary \(t_{\rm chunk}\) at inference time. Since the thinking of Shanks always lags behind the latest user speech by \(t_{\rm chunk}\) seconds, changing \(t_{\rm chunk}\) can affect how soon the SLM can hear the latest user speech and affect the response latency. As an ablation, we change \(t_{\rm chunk}\) to 3 and 5 during inference without retraining Shanks-E2E. The results are shown in the bottom two rows in Table ¿tbl:tab:interruption32results?. On the wrong subset, we do not find the interrupt ratio and valid interrupt ratio to change significantly compared with \(t_{\rm chunk}=4\). Interestingly, we find that the interrupt latency on the wrong subset for \(t_{\rm chunk}=3\) is the smallest, while the \(t_{\rm chunk}=5\) has the largest interrupt latency.
Next, we move on to the second scenario: making tool calls when listening. The experiment results are presented in Table ¿tbl:tab:api-call-results?. We summarize the main findings as follows:
Shanks successfully makes at least 56.9% of API calls when the user is still speaking. Among the successful API calls made by the two Shanks models, about 80% to 90% of the API calls are made during the user speech. In the example in Figure 4, Shanks-E2E makes four out of six API calls when the user is still speaking. Compared to the call-after-listen baseline, which makes all the API calls after the user turn finishes, Shanks can significantly reduce the response latency by using the time the user is speaking to make tool calls.
Call-after-listen has a higher success rate and response quality. While Shanks elegantly uses the user speaking time to make API calls, the success rate and response quality (correctness and completeness) lag behind call-after listen. We find that this is because during inference, if the API call fails during \(R_i\), Shanks seldom retries the failed API call in future \(R_j\), where \(j>i\). On the other hand, the call-after-listen baseline is more likely to retry failed API calls.
Combining Shanks and call-after-listen yields the best performance. To solve the previously observed issue, a simple method is to use Shanks when the user is still speaking and back off to call-after-listening when the user’s speech ends. Precisely, when the user is still speaking, we use Shanks to call APIs while listening, and only keep the successful API calls and their responses. When the user finishes their speech, we switch to the call-after-listening mode, where the input to the SLM is the complete user speech. We also feed the success API calls and responses previously made by Shanks to the model, as if the call-after-listen model had made those API calls by itself. Based on the full input speech and previous success API calls and responses, the model continues to make the remaining API calls. Since some API calls have already been made by Shanks when the user speaks, this combined method enjoys the thinking-while-listening advantage of Shanks.
max width=
In the last row in Table ¿tbl:tab:api-call-results?, we show the result of combining Shanks-E2E with call-after-listening. This combined method has a high number of early call accuracy while also having a higher task success rate and response quality. Among the API calls that are successful, over 60% of API calls are made during the user speech, while only 40% of the API calls are made after the user finishes. This is in stark contrast to call-after-listen, where 100% of the successful API calls are made after the user finishes, resulting in a much higher latency. As a result, the combined method only needs to generate on average 117 tokens after the user has finished, while the call-after-listen baseline needs to generate on average 313 tokens after the user’s turn ends. So thinking while listening reduce the response by 62.3%.
Thinking before responding has been widely explored in text-only LLMs [3], [4]. Recently, this "think then respond" paradigm has been applied to audio-aware language models, which takes speech (or audio) as the input and output texts [34]. Note that these audio-aware language models, which only output text, are different from the speech-in-speech-out SLMs focused on in our paper. While thinking before responding can improve the response quality and yield significant performance on many challenging benchmarks [5], [35], thinking before responding creates great response latency. As a result, it is impractical to directly apply thinking-before-responding to SLMs, speech-in-speech-out models, which require real-time and low-latency interaction [28], [36], [37]. Developing reasoning methods that preserve real-time interaction in SLMs remains an open problem.
Concurrent to us, [38] introduce thinking to SLMs by a thinking-while-speaking method called Stitch. Stitch uses the fact that a chunk of audio in the speech response takes less time to generate than it does to play to the user, and the model can use the remaining time to generate thinking tokens when the SLM is still speaking. While both Shanks and Stitch explore unspoken thinking processes for SLMs, the main distinction is when the thinking happens. In Shanks, the thinking process happens when the user is still speaking, while Stitch thinks when the SLM is speaking. In fact, the two methods can be combined: an SLM can think when listening and speaking. We believe this will be the future of SLM, and we leave this as a promising future direction.
Another concurrent work released on arXiv less than one week ago (10/02/2025), Stream RAG [39], also studies calling tools (web search and knowledge graph APIs) during the user’s speech. This is similar to our second scenario introduced in Section 3.2. However, Stream RAG focuses on when to issue retrieval/tool queries while listening and does not introduce an explicit silent chain-of-thought (‘thinking’) process like we do. In contrast, our paper studies a broader thinking-while-listening paradigm, with tool-calling as one application, and show benefits such as improved user-interruption decisions.
In this paper, we introduce Shanks, a framework that enables SLMs to think while listening. Shanks achieves thinking-while-listening by chunking the user input speech and progressively reasoning over the available user inputs. When the user is speaking, Shanks is generating thinking chunks for all previous input speech, achieving thinking while listening. We demonstrate the potential of Shanks on two scenarios: First, Shanks can listen to the user solving a math problem step-by-step, and interrupt the user when the user is making a mistake. Second, we focus on a tool-augmented task-oriented dialogue setting and show that Shanks can listen to the user speech and evoke necessary API calls when the user is still speaking. On ComplexFunxBench, Shanks successfully calls more than half of the APIs that are required to complete the user’s request when the user is still speaking. This reduces the response latency, as the model only needs to call the remaining half APIs after the user has finished.
While Shanks shows great potential in improving the user-SLM interaction, we see the following limitations of the method. First, Shanks requires the user’s speech to have certain structures: the user’s speech needs to be long enough to allow the model to perform meaningful reasoning when listening, and the information in the user’s speech needs to be able to be processed in a sequential order. This kind of speech can naturally occur, as shown in the two scenarios we studied.
Next, Shanks uses a fixed chunk size to segment the user input speech. The chunking nature of Shanks means that Shanks always lags behind the user’s speech by \(t_{\rm chunk}\) seconds, incurring latency in the thinking process. We encourage future work to reduce the latency between thinking and listening by using more sophisticated chunking methods.
Last, the goal of the user might be unclear when the user’s speech is not completed, and the thinking tokens generated during listening may be redundant and not always be useful to address the goal of the user. While thinking during listening does not incur additional latency after the user’s speech ends, Shanks still significantly increases the compute cost during inference. Although Shanks has some limitations, we believe that our effort in proposing a novel modeling method, thinking while listening, together with the reasonable scenarios and convincing results, already contributes greatly to the research community by shedding light on a potentially fruitful research direction.
The authors would like to thank Ke-Han Lu, Chen-An Lee, Yi-Cheng Lin, and Wei-Chih Chen for their valuable feedback on the draft of this paper.
In the main content of the paper, we say that we chunk the user input audio into fixed-size chunks of \(t_{\rm chunk}\) seconds. In fact, what we do is chunking at the level of feature representation instead of the level of the audio waveform. Precisely, when encoding the \(i\geq2\) speech chunks \(S_i\), we feed the full speech through the audio encoder, and only take the speech representation for the corresponding speech chunk. If we directly chunk the audio waveform and encode each audio chunk independently, the representation of later audio chunks will not be able to depend on the earlier audio chunks, which can potentially lead to performance degradation.
We fine-tune the models using the Llamafactory [40] toolkit. When generating the training data using GPT-4o, we do not feed the audio of the user’s speech into GPT-4o. Instead, we feed the transcription of the speech chunks. This is because using the speech chunk will increase the cost and time to call the API. To obtain the transcription of each chunk, we use Whisper-large-v3 [30] to obtain the transcription and timestamp for each word in the user speech, and then segment the transcriptions into chunks based on the timestamp. While the timestamp obtained from Whisper may not be very precise, this is already sufficient for preparing the training data.
To prepare the training data, we randomly sample 5K samples from Tulu-3-SFT-Math-Grade [18], which can be loaded from Huggingface datasets [41]: https://huggingface.co/datasets/allenai/tulu-3-sft-personas-math-grade. We follow the procedure detailed in Section 3.1.2 to construct the training data. We additionally filter out audios that are longer than 80 seconds, so the final training dataset is slightly less than 5K.
We fine-tune the thinker on the training data for two epochs on 8 A100 GPUs. The effective batch size is 64. We set the learning rate to \(1.0e-4\) with cosine learning rate scheduling and a 0.1 warm-up ratio [42]. The same training hyperparameters are used across all three models, including the Shanks-E2E, Shanks-Cascade, and no-thinking model.
The training data is mostly generated by GPT-4o. We include the prompts to generate the reasoning chunks in Table ¿tbl:tab:thinking-while-listening-1? and ¿tbl:tab:thinking-while-listening-2?, the prompt to generate the interruption in Table ¿tbl:tab:interrupt-correction-template?, and the prompt to generate the response without interruption in Table ¿tbl:tab:assistant-full-turn32interruption?.
We use the procedure detailed in Section 3.2.2 to construct the training data. The training data consists of 500 samples. The prompt used to determine when an API call can be made is shown in Table ¿tbl:tab:earliest-tool-call?. The prompt used to generate the final response is shown in Table ¿tbl:tab:tool-results-final-response?.
For the think-after-listen and Shanks-E2E model, we fine-tune them using LoRA [43], as the sequence length for this dataset is very large and full fine-tuning will result in out-of-memory. We also fine-tune the LM head and the token embedding of the talker model; otherwise, the model will not be able to recognize and generate special tokens. For the Shanks-Cascade model, we fine-tune all the parameters. As the training dataset is smaller, we fine-tune the model for 10 epochs, while other training hyperparameters follow those in Appendix 9.1.
To determine the time of interruption \(t_{\rm interrupt}\) when evaluating the interrupt latency, we apply the following procedure. We use Whisper-large [44] to obtain the timestamp of each word in the user speech from the testing set, and we use GPT-4o to determine when the first error in the user speech occurs by giving GPT-4o the question, the ground truth answer, the transcription of the user speech, and the word-timestamp alignment. The prompt used to determine the first error time \(t_{\rm error}\) is shown in Table ¿tbl:tab:error-timestamp-template?.
To evaluate the valid interrupt ratio, we use GPT-4o as the judge. The prompt used to determine whether an interruption is valid is shown in Table ¿tbl:tab:assistant-interruption-judgement?.
When evaluating the Shanks models, when the user is still speaking, each thinking chunk can only include at most 320 tokens generated by the SLM itself. In some cases, the API call augments may include very long tokens, which can easily exceed the token limit for the thinking chunk, and the API call will be incomplete and unsuccessful. While we do not specifically handle this kind of case, there is a simple workaround that can resolve the above issue: The long arguments for the API call generated by the model must be the returned value of previous tool calls, so we can include previous API call responses in a reservoir of the draft for speculative decoding [45], [46], and use speculative decoding to speed up the inference speed.
ComplexFuncBench is originally designed to evaluate a model’s ability to parse long-context information, and the API call response can be very long. However, since Qwen-omni only has a context length of 32,768, once the token exceeds this limit, we directly terminate the inference for a tested instance. As a result, some of the testing samples we evaluate may fail because the number of tokens exceeds the max sequence length of the model.
max width=
max width=
max width=
max width=
max width=
max width=
max width=
max width=
Work done during an internship at Microsoft GenAI. dcml0714@gmail.com↩︎
Correspondence: xiaofei.wang@microsoft.com↩︎
If the interruption happens in \(R_i\), the user is currently speaking the \((i+1)\)-th speech chunk \(S_{i+1}\), as the thinking \(R_i\) happens simultaneously when the user says \(S_{i+1}\). Consequently, we also feed the transcription of \(S_{i+1}\) into the judge model when determining whether the interruption is valid. Nevertheless, we do not find the results to differ significantly if we only give the judge model up to the speech chunk \(S_i\).↩︎
ComplexFuncBench is originally designed as an evaluation dataset. Here, we train the model directly on this dataset since the models we use were not trained to perform tool call. Since our training data has the exactly same distribution as the testing data, our results should not be compared with other models that are not trained on this dataset.↩︎
The tokens in the API call responses are not included in the \(t_{\rm chunk}\times n_{\rm tps}\) limit since these are not the tokens generated by the SLM itself. For simplicity, we do not consider the API call latency, as our environment already prepares all the ground truth API responses.↩︎