Cloud, Services and Security

What Is Call Transcription, and How Does It Work?

Call transcription turns a call recording into written text. A computer engine listens to the recording, recognizes the words and writes them down, and sometimes also marks who spoke when. This way you can read a call instead of listening to it, search it for a word and produce a summary from it.

TranscriptionSpeech to TextSTTSpeech recognitionConverting speech to textAutomatic transcription

Reading time: about 8 minutes

What Is Transcription, Exactly

Transcription is the conversion of speech into written text. It used to be done only by people: a typist sat with headphones, paused the recording every few seconds, and typed. Today most of the work is done by a computer engine called speech recognition (Speech to Text, or STT for short).

In the phone system, transcription is usually done on the call recording. That is, the call is recorded first, and only then does the recording go to the transcription engine. The result is a text file attached to the call, next to the details that appear in the call history: who called, when, and how long they talked.

The closest comparison is the minutes of a yeshiva meeting. Even someone who wasn't in the room can read what was said, find the moment when something was decided, and pass the summary along without listening to an hour of recording.

Why Transcribe Calls at All

Listening to a recording takes exactly as long as the call itself. A ten-minute call takes ten minutes to hear, and sometimes more, because you go back and replay parts. Reading the same call as text takes a minute or two, and you can skip straight to the important part.

A second advantage is search. You can't search for a word inside an audio file, but you can in text. A manager who wants to know which call a customer mentioned a "refund" in, or a gabbai looking for who asked for a "seat in the synagogue", finds it in seconds.

A third advantage is documentation. In a busy office, details get forgotten: what was promised to the customer, what date was set, what number they left. A transcript keeps everything in writing, and you can return to it even months later.

How the Transcription Engine Works

The engine doesn't "hear" the way a person does. It takes the sound waves, divides them into very short segments, and calculates for each segment which sounds were likely spoken. Then it joins the sounds into words, and the words into sentences, based on what it learned from huge amounts of speech and text.

In the last stage, the engine guesses the context. If it heard something that could be one of two similar-sounding words, it chooses based on the rest of the sentence. That is why a clear, complete sentence is transcribed better than single words thrown out amid noise.

Newer engines are based on artificial intelligence and learn from large collections of language. This greatly improved accuracy, but it also created a new problem: sometimes the engine "fills in" words that were never said, so that the sentence looks logical. We will say more about this below.

Accuracy: What It Depends On

No engine transcribes at one hundred percent. Accuracy varies from call to call, and a few things affect it more than anything else:

  • Audio quality — a call on a good network, with a quality codec and no dropouts, is transcribed far better than a call from a mobile phone inside an elevator.
  • Background noise — a street, children, a noisy air conditioner, or an open speaker all hurt recognition.
  • Overlapping speech — when two people talk at once, the engine has trouble separating them.
  • Accent and pace — very fast speech, a heavy accent, or swallowed syllables lower accuracy.
  • Special terms — personal names, street names, professional terms, and Hebrew religious language (lashon hakodesh).

Rule of thumb: a transcript is good enough to understand what happened in a call and to find things in it. It is not always good enough to quote a phone number or an amount of money without checking against the recording.

Speaker Identification: Who Said What

A continuous transcript of the whole call as one block of text is very hard to read. It isn't clear where the agent is speaking and where the customer is. That is why a good transcript adds speaker labels (Speaker Diarization): "Speaker 1", "Speaker 2", or "Agent" and "Customer".

There are two ways to get there. The first is guessing from the voice itself: the engine recognizes that two different voices are speaking and separates them. The second, which is more accurate, is recording each side on a separate channel, so it is clear in advance which side is which.

In practice, it is worth checking the speaker labels on several calls before relying on them. A transcript in which the speakers get mixed up can mislead more than a transcript with no labels at all, because it attributes words to the wrong person.

Hebrew and Yiddish

Hebrew is a relatively hard language for transcription engines. It is written without vowel marks, so a single word can be read in several ways. It has many short words that attach to the next word (such as the prefixes for "and", "the", "in" and "that"), and there are big differences between everyday speech and written language. Even so, today's engines transcribe spoken Hebrew at a reasonable to good level.

Harder is speech that mixes Hebrew with lashon hakodesh, Aramaic and terms from the yeshiva world. "A question in the laws of Shabbat", "a stipend for an avrech", "the Daf Yomi" — the engine knows some of this, and sometimes writes something that sounds similar but means something else.

Yiddish is harder still, and it is important to say so honestly. There is far less recorded and transcribed Yiddish material for the engines to learn from, and there are several ways of spelling and pronouncing it. What often happens is that the engine "corrects" the Yiddish into Hebrew and returns text that looks readable but does not faithfully reflect what was said. That is why, in an organization that speaks Yiddish, automatic transcription is only a helping tool, and it must be checked on real calls before you rely on it.

The Danger of Text That Looks "Too Good"

In older transcription, a mistake looked like a mistake: garbled words, a sentence that didn't hang together. In transcription based on artificial intelligence, mistakes sometimes look like perfectly fine text. The engine completes a sentence in a logical way, but not necessarily the correct one.

Example: a customer says, "I'll come on Tuesday, maybe Wednesday." The engine might write "I'll come on Tuesday" and leave out the "maybe". No one will notice, because the sentence looks natural. When the detail matters — a date, an amount, a promise — it is worth clicking on the recording and listening to the passage itself.

So the general recommendation: the text is the key to finding the information, and the recording is the source. When in doubt, the recording decides.

From Summary to Action: What to Do With the Transcript

Transcription alone is a first step. The next step, which is becoming more common, is analysis of the text: a short summary of the call, a list of the topics that came up, and the tasks that follow from it ("send a price quote", "call him back on Thursday").

Common uses in the field:

  • Customer service — a manager goes over the day's call summaries in ten minutes, instead of listening for hours.
  • Sales — the summary goes into the customer's record, and the next agent who talks to them knows what has already been said.
  • Training — look for calls where a particular question came up, and learn from them how to answer well.
  • Institutions and nonprofits — documentation of inquiries from donors, parents or people receiving support, without writing them down by hand.

All of these work better when the transcript is connected to a CRM system, so the information is kept in one place next to the customer's name, and not in a folder of files.

Live Transcription vs. Transcription After the Call

There are two main types. Transcription after the call (Batch) receives the full recording once the call has ended. It sees the whole call from beginning to end, so it is more accurate and labels speakers better.

Live transcription (Real-time) writes during the call, word after word. It is useful, for example, for captions for people who are hard of hearing or for real-time suggestions to an agent, but it is less accurate, because it has to decide quickly without hearing the rest of the sentence.

After the callLive
When you get the textAfter the call endsDuring the call
AccuracyHigherLower
Speaker labelsBetterPartial
Typical useDocumentation, summaries, searchCaptions, help for the agent

For most offices, transcription after the call is entirely enough. What matters is getting reliable text, not getting it within a second.

Privacy and Security in Transcription

Transcription turns a call into text that is easy to copy, send and search. That is an advantage, and it is also a responsibility. The text of a call with a customer can include names, addresses, health conditions or financial details.

So it is worth deciding in advance who may read transcripts, how long they are kept, and when they are deleted. Everything that is true of recording privacy is true of transcription too, and even more so, because text spreads easily. It is also worth making sure that the permissions in the system limit viewing to only those who need it.

How to Start the Right Way

An organization that wants to start working with transcription can save itself a lot of disappointment by starting small:

  1. Choose sample calls — ten to twenty real calls of different kinds: short and long, clear and noisy.
  2. Compare to the recording — someone from the office reads the transcript while listening at the same time. That shows where the engine makes mistakes.
  3. Decide what you will use it for — if the accuracy is enough for a summary but not for a quote, that is fine. You just need to know it in advance.
  4. Improve the audio — good headsets for call center staff and reasonable quiet in the office improve transcription more than any setting.

In an organization that has calls in Yiddish or other languages, it is worth including them in the very first test, and not discovering the gap after everyone has gotten used to trusting the text.

How it works with us at Kesher

At Kesher, transcription is part of the CRM connection. When a call is recorded, the recording is transcribed, and then analyzed by artificial intelligence, which produces a summary, topics and tasks. All of this is saved in the customer's record.

The customer is identified by the number they called from, and a customer record is opened automatically. On the next call, the customer's name appears on the display, and the summaries of the previous calls are already waiting in the record.

So that there is something to transcribe, calls need to be recorded. This is set up in the control panel, in recording groups with a retention period, and in the recording setting of each extension.

About Yiddish — we say this up front: it is the hardest challenge in transcription. If your organization speaks Yiddish, we will ask to check together on real calls before you rely on the transcript. And if anything is unclear, every question is answered by a person at our company, not an automated system.

FAQ

Can a call that wasn't recorded be transcribed?

No. Transcription after the call works on the recording. If the call wasn't recorded, there is nothing to transcribe.

How long does it take to get a transcript?

It depends on the system and the length of the call. Usually it is not instant, because the recording has to be ready first, and only then is it transcribed and analyzed.

Does the transcription understand people's names?

Partly. Common names are usually recognized, but rare family names, street names and special terms are sometimes written wrong. When the name matters, check against the recording.

Why is the transcript wrong in certain calls in particular?

Usually because of the audio: background noise, an open speaker, weak mobile reception, or two people talking at once. Improving the audio quality improves the transcription.

Can I rely on a Yiddish transcript?

Only with caution. Yiddish is much harder for the engines, and they tend to return text that looks like proper Hebrew but does not reflect what was said. You must check on real calls before relying on it.

What is the difference between a transcript and a summary?

A transcript is everything that was said, word for word. A summary is a few lines describing the main point of the call. The summary is built from the transcript, so a mistake in the transcript can carry over to the summary too.

Who can read the transcripts?

Whoever has been given permission. It is worth limiting viewing to only the people who need the information for their work.

Does the transcript replace the recording?

No. The recording is the source, and the transcript is a fast way to find things in it. When there is a dispute about what was said, the recording decides.

Back to the Knowledge Center — all terms

Want to hear how it would work for you?

Tell us how your phones work today — how many calls, who answers, what gets in the way — and we'll get back to you with an organized proposal.

Leave your details and we'll get back to you
077-921-9000