October 11, 202611 min read

Introducing Indic Smart Turn: Turn Detection for Indian Languages


Most voice agents run the same loop. They turn the caller's speech into text, send the text to a language model, and speak the model's reply. Before any of that, the agent has to decide that the caller has finished speaking, because that moment starts the agent's reply.

Most agents make that decision with a voice activity detector (VAD) and a timer. The VAD is a small model that answers one question about each slice of audio: is this speech or silence? When it reports silence, a timer starts. If the silence outlasts a fixed limit, for example 0.8 seconds, the agent treats the turn as finished and replies.

People also pause in the middle of a sentence:

text
I was going to the market [pause] to buy milk.

The VAD reports the pause after "market" and the pause after "milk" the same way: silence. The only thing the agent can adjust is the length of the limit, and each choice fails in its own way.

  • A short limit interrupts. At 0.3 seconds, the agent starts replying after "market", while the caller is still mid-sentence.
  • A long limit delays every reply. At 1.5 seconds, the agent waits through the pause after "market". It also waits 1.5 seconds after "milk", and after every other finished sentence in the call, before it starts to reply.

What Smart Turn adds

Pipecat's Smart Turn gives the agent a second check. When the VAD reports silence, Smart Turn takes the audio of the caller's current turn, up to its last eight seconds, and returns a probability that the caller has finished. A turn shorter than eight seconds gets silence added at the start. Above 0.5, the agent replies. Below it, the agent keeps listening.

Smart Turn listens to the audio itself, so it hears how something was said as well as what was said. Written down, "I was going to the market" could be a whole sentence. Said aloud with the voice trailing off at "market", it sounds unfinished, and anyone listening can tell more is coming. A transcript keeps the words and drops that tone.

Why I built it

Smart Turn v3.2 lists 23 supported languages. Hindi, Marathi and Bengali are the only Indian ones, and Pipecat trained them mostly on synthetic speech generated by text-to-speech. Tamil, Telugu, Kannada, Malayalam, Gujarati, Punjabi, Odia and Assamese are absent.

On the real Indian phone calls in my test set, Smart Turn v3.2 gets 64 to 75% of pauses right in those eight absent languages. On English it gets 94% right. The gap matters to anyone building a voice agent for Indic languages, because a turn detector that is wrong on a quarter to a third of pauses interrupts people several times in a single call.

So I built Indic Smart Turn, and I am releasing it as open source. It covers English and eleven Indian languages: Hindi, Marathi, Bengali, Tamil, Telugu, Kannada, Malayalam, Gujarati, Punjabi, Odia and Assamese. On a public test set of 20,431 clips, the recommended model gets 89.8% of pauses right, against 81.9% for Smart Turn v3.2, and it is ahead on every Indian language. It takes the same input as Smart Turn v3, so it works in Pipecat with a one-line change. The models, the 50,421-clip dataset and the code that trained and scored them are all public.

GitHubadimyth/indic-smart-turnTraining and evaluation code, reports, charts and a stage-by-stage record of how the model was built. GitHubadimyth/indic-smart-turn on Hugging FaceThe models, with the model card and usage code. GitHubadimyth/indic-smart-turn-data on Hugging Face50,421 labelled clips from real phone conversations in eleven Indian languages.

Finding the speech

Three sources supplied the audio, all under CC BY 4.0.

  • IndicVoices, from AI4Bharat, holds real two-person phone calls with human transcripts. Each recording has one speaker, cut into segments.
  • TamilEOT holds 18,485 Tamil turn boundaries from 116 calls, checked by people.
  • Pipecat's own training data covers English, Hindi, Marathi and Bengali, so the new model keeps what Smart Turn v3.2 already does well.

IndicVoices is 30 to 50 GB per language, and only about 17% of it is conversation. I downloaded only the files holding conversation rows, about 2 to 3 GB per language.

Labelling every pause

IndicVoices cuts its segments at pauses, not at the ends of turns. A segment ending tells you the speaker stopped for a moment. It does not tell you they had finished. Every clip needed a label: complete or incomplete.

I used two labellers and kept both answers.

The audio label. gemini-3.7-flash heard each clip, along with its transcript when one existed. It answered complete or incomplete, with a confidence and a one-line reason. For the market example, it returned this:

json
{"verdict": "incomplete", "confidence": 0.9, "reason": "sentence still open"}

This is the label the model trains on.

The text label. gpt-6-luna read one speaker's whole session at once, one numbered line per segment, and marked whether the speaker had finished at the end of each line. Because it sees the whole session, it can use the next line to judge the current one.

text
[1] (2.1s) I was going to the market
[2] (1.4s) to buy milk
[3] (0.8s) okay thanks

It marked line 1 incomplete because line 2 finishes the sentence, line 2 complete, and line 3 complete because a short acknowledgement is a full turn.

The two labels agreed on 89 to 93% of clips in each language. Most disagreements came down to tone of voice, which the text model cannot hear. One Hindi clip reads "सर आप लच्छा पराठा कर दीजिए", which means "Sir, please make the lachha paratha." From the transcript alone, the text model thought the speaker would go on. The audio model heard a polite request ending on a falling tone, and marked it complete. In the test set, I moved the 405 clips where the two labellers disagreed into a separate group and scored it on its own. When the labellers disagree, neither label is a trustworthy answer to grade a model against.

Pauses inside a segment. Segment ends are mostly finished turns, so I added unfinished examples by cutting segments at their internal pauses, the moment a VAD would report silence in the middle of a turn. I kept a cut only when Gemini heard it as unfinished.

Short replies. Segments under 1.5 seconds are almost all one-word replies, such as "haan" or "okay". In some languages they made up a large share of the clips, leaving fewer of the longer cases where the decision is harder. I kept them to at most 20% of each language's clips.

The result is 50,421 labelled clips across eleven languages, divided by speaker so that nobody in the test set appears in training. Between 67% and 75% of each language's clips are complete turns.

Checking the measuring stick

Before I trusted any number from my evaluation code, I used it to reproduce two published results on the TamilEOT test set, which people had checked by hand.

ModelMy evaluatorPublished
Smart Turn v3.2, CPU model70.2%70.3%
The TamilEOT paper's Tamil model86.1%86.13%

Both landed within 0.1 point, so I trusted the evaluator with the new models.

Training the model

The recipe is Pipecat's, applied to the new data. It starts from Whisper, OpenAI's speech recognition model. Whisper has two halves: an encoder that turns audio into a summary of what it heard, and a decoder that writes the transcript. The recipe keeps the encoder, drops the decoder, and adds a small layer that turns the encoder's summary into one number, the probability that the turn is complete.

I trained two sizes: whisper-tiny, with 8M parameters in its encoder, and whisper-base, with 20M. Both take up to eight seconds of audio before a pause as a spectrogram, a picture of which frequencies are loud at each moment.

Both sizes come in two formats, fp32 and int8.

Size and speed

Each model scores one clip at a time on a single CPU thread:

ModelSizeCloud serverMacBook
Smart Turn v3.2, int88.7 MB80 ms34 ms
Smart Turn v3.2, fp3232 MB90 ms22 ms
Indic Smart Turn tiny, int88.8 MB83 ms34 ms
Indic Smart Turn tiny, fp3232 MB84 ms22 ms
Indic Smart Turn base, int824 MB122 ms44 ms
Indic Smart Turn base, fp3281 MB190 to 208 ms37 ms

Tiny int8 matches Smart Turn v3.2's CPU model on size and speed. The model runs only when the VAD reports silence, not on every frame of audio.

Results on the public test set

The test set holds 20,431 clips that no model trained on. It combines the IndicVoices test split for all eleven languages, TamilEOT's test set, and Pipecat's own test set for English, Hindi, Marathi and Bengali.

The recommended model

Across the whole test set, the recommended model, base int8 after the fine-tuning described below, gets 89.8% of pauses right. Smart Turn v3.2's CPU model gets 81.9%. Indic Smart Turn is ahead on every Indian language, and its AUC is higher in all twelve languages, English included.

LanguageSmart Turn v3.2, int8Indic Smart Turn base, int8Gain
Malayalam63.8% / 0.66485.6% / 0.873+21.8
Assamese64.9% / 0.69183.3% / 0.892+18.4
Punjabi69.4% / 0.74986.5% / 0.892+17.1
Tamil70.0% / 0.74986.7% / 0.922+16.7
Kannada74.0% / 0.78389.1% / 0.920+15.1
Gujarati72.5% / 0.76587.1% / 0.896+14.6
Telugu70.9% / 0.76684.8% / 0.923+13.9
Odia75.2% / 0.77887.3% / 0.933+12.1
Marathi77.7% / 0.86588.1% / 0.955+10.4
Hindi85.8% / 0.94191.5% / 0.971+5.7
Bengali81.4% / 0.88486.8% / 0.936+5.4
English94.3% / 0.98393.9% / 0.988−0.4
All81.9% / 0.90289.8% / 0.958+7.9

How to read it: Each cell shows accuracy, then AUC. Gain is in accuracy points. Both models were scored in one run on one machine.

The gains are largest where Smart Turn v3.2 has no training data. On the eight Indian languages it does not support, Indic Smart Turn is ahead by 12 to 22 points. On Hindi, Marathi and Bengali, which it does support, Indic Smart Turn is still ahead by 5 to 10 points. On English the two are level: 0.4 points apart on accuracy, with Indic Smart Turn's AUC slightly higher.

The other models

The other Indic Smart Turn models also beat Smart Turn v3.2's CPU model on every Indian language.

Indic Smart Turn modelGain on the Indian languagesGain on English
tiny, fp32+4.0 to +17.6−0.6
base, int8, before fine-tuning+5.0 to +19.5+0.5
base, fp32, after fine-tuning+5.1 to +21.6−0.3

How to read it: Gains are in accuracy points over Smart Turn v3.2's CPU model, scored on the same machine as the table above. Tiny int8 was scored on a different machine, so it is not in this table. The README has per-language charts for every model and an interactive version.

A side finding: int8 results depend on the CPU, and fp32 results do not. On the same clips with the same code, Smart Turn v3.2's int8 model scored 3.7 points lower on Kannada on an Apple Silicon MacBook than on an x86 cloud server, while every fp32 model gave identical numbers on both. ONNX Runtime uses different int8 kernels on x86 and Arm, so a quantised model behaves slightly differently on each. Statically quantised models moved by up to 3.7 points per language, in either direction, and Indic Smart Turn base int8 by about 1 point. Benchmark an int8 model on the hardware you deploy it on, and use fp32 where it is fast enough, as it is on Apple Silicon and GPUs. The measurements are in the repository.

Two external test sets

TamilEOT and Pipecat's test set are other projects' published benchmarks, so I also compared the models on each of them on its own. Indic Smart Turn holds up on both.

  • TamilEOT, 4,168 Tamil clips labelled by people. The recommended model scores 86.2%, against about 70% for Smart Turn v3.2 and 86.1% for the TamilEOT paper's own model, a whisper-base trained only on Tamil. One model covering twelve languages matches a model built for Tamil alone.
  • Pipecat's own test set for English, Hindi, Marathi and Bengali, 10,878 clips. The recommended model scores 92.7%, against 92.1% for Smart Turn v3.2's CPU model. Adding eleven Indian languages cost nothing on the languages Smart Turn v3.2 already handled.

Testing on sales-training calls

I also tested the models on recordings from a sales-training product, where trainees rehearse a sales pitch by talking to an AI voice agent. Each session is one recording with both voices mixed together.

I separated the trainee's voice from the agent's, cut a clip at every trainee pause of 200 ms or more, and had Gemini label each clip the same way as the public data. That left 1,207 sessions in ten languages. The 12,710 test clips come from sessions held out of all training. The recordings are private, so I publish only aggregate numbers.

Smart Turn v3.2, int8Indic Smart Turn tiny, int8Indic Smart Turn base, int8, before fine-tuning
Accuracy66.6%72.7%77.1%
AUC0.7710.8210.859
Replied when the trainee had finished84.5%84.5%85.7%
Interrupted when the trainee had not finished47.2%36.4%29.6%

Base int8 is ahead in every language, by 8 to 17 points. Tiny int8 is ahead in every language except Odia, where 95 clips from two sessions are too few to say.

All three reply to a finished turn about equally often. The difference is in unfinished sentences: Smart Turn v3.2 interrupts nearly half of them, and Indic Smart Turn base interrupts fewer than a third.

Fine-tuning it further

I then fine-tuned the base model on sales-training clips from sessions outside the test set, mixed with public clips so it would keep what it learned there.

Test setBefore fine-tuningAfter fine-tuning, recommended
Sales-training calls, 12,710 clips77.1% / AUC 0.85979.4% / AUC 0.875
Public test, 20,431 clips89.9% / AUC 0.95589.8% / AUC 0.958

The fine-tuned model gains 2.3 points on the sales-training calls and holds level on the public test, where it is ahead of the earlier version in eight of twelve languages and within a point behind in the other four. The gain on those calls comes from the same kind of audio, with the same labeller, as the tuning data. It shows the model fitting this product's calls better and says little about other products.

Try it in your voice agent

The models are on Hugging Face and take the same input as Smart Turn v3.

Pipecat

python
from pipecat.audio.turn.smart_turn.local_smart_turn_v3 import LocalSmartTurnAnalyzerV3

analyzer = LocalSmartTurnAnalyzerV3(smart_turn_model_path="indic-smart-turn-base-int8.onnx")

Use it wherever you would pass the built-in analyzer, for example in the user-turn stop strategy. The default threshold is 0.5.

Directly

The model takes an 80 x 800 log-mel spectrogram of the last 8 seconds of 16 kHz mono audio and returns the probability that the turn is complete. config.json records the same contract.

python
import numpy as np, onnxruntime as ort, soundfile as sf
from transformers import WhisperFeatureExtractor

MODEL = "indic-smart-turn-base-int8.onnx"
fe = WhisperFeatureExtractor(chunk_length=8)           # 8 s window, 80 mel bins, 800 frames
so = ort.SessionOptions(); so.intra_op_num_threads = 1  # one core is enough at batch 1
session = ort.InferenceSession(MODEL, so, providers=["CPUExecutionProvider"])

def last_8s(audio: np.ndarray, sr: int = 16000) -> np.ndarray:
    n = 8 * sr
    return audio[-n:] if len(audio) > n else np.pad(audio, (n - len(audio), 0))  # keep the end, zero-pad the front

def turn_complete_prob(audio: np.ndarray) -> float:
    feats = fe(last_8s(audio), sampling_rate=16000, return_tensors="np", padding="max_length",
               max_length=8 * 16000, truncation=True, do_normalize=True).input_features.astype(np.float32)
    return float(session.run(None, {"input_features": feats})[0][0, 0])

audio, sr = sf.read("clip.wav", dtype="float32")        # 16 kHz mono; resample first if not
p = turn_complete_prob(audio)
print("complete" if p > 0.5 else "incomplete", round(p, 3))

Indic Smart Turn is open source. If you build voice agents in any of these languages, try it on your own calls and open an issue where it gets the turn wrong. A test set labelled by people, in any language other than Tamil, would help the most.