What Are Resilience and Redundancy?
Resilience is the ability of a system to keep working even when something fails. Redundancy is the way to achieve it: more than one component for each role — servers, connections, power — so that if one fails, another carries on.
Two words, one idea
Resilience (in English, High Availability) is the result: the phone keeps working even on a bad day. Redundancy is the method: keeping "one more" of every important component, even if it isn't needed on an ordinary day. The thinking behind it is simple: there is no component that never breaks, so you build a system that doesn't depend on any single component.
In professional terms we talk about a Single Point of Failure: any place where one component failing brings down everything. The goal of the design is to find these points and eliminate them, one after another.
Another term worth knowing is Failover: the automatic switch from a failed component to the backup component. Redundancy without automatic switching is half a solution: the backup exists, but someone has to notice and activate it.
Availability is commonly measured in percentages, and it's worth understanding what the numbers mean. 99% sounds excellent, but it is more than three days a year without service. 99.9% is a little under nine hours a year. Each additional "nine" cuts downtime by a factor of ten, and requires another layer of redundancy.
The four layers
In business telephony there are four layers, and in each of them something can go wrong:
- The servers — how many servers, in how many data centers, in how many countries.
- The provider's network — how many carriers and how many connections to the telephone network.
- The office — a UPS for the router and switch, backup internet, and sometimes two internet providers.
- The plan — what happens to calls when the whole office is unavailable: forwarding to a mobile phone, to another branch, or to voicemail.
The first two layers are the provider's responsibility. The third is mostly yours. The fourth is shared: the provider makes it possible, and you decide what is right for your organization.
What happens at the moment of failure
Suppose one data center loses power. In a correctly built system, new calls start flowing through the other servers. Customers who call you don't know anything happened. A call that was in progress at the moment of failure may be dropped, which is the small price, but redialing already goes through.
Now suppose the failure is at your end: the electric company cut the supply on your street. Without a plan, callers hear endless ringing or a busy signal. With a plan, the PBX detects that the phones in the office are not registered and sends the calls to the alternate destination. The business owner in Bnei Brak keeps receiving the calls on his mobile phone, and the customer doesn't know the office is dark.
Cloud PBX versus an office PBX
| Scenario | Physical PBX in the office | Cloud PBX |
|---|---|---|
| Power outage at the office | Everything is down, unless there is a large UPS | The PBX keeps working; calls go to an alternate destination |
| Internet outage at the office | Internal calls may work; external calls depend on the lines | The PBX keeps working; calls go to an alternate destination |
| PBX equipment failure | Until a technician arrives and replaces it | The provider's responsibility, with backup servers |
| Fire or flood in the office | Equipment and settings may be lost | Settings and recordings are kept in the cloud |
The table shows why many organizations move to the cloud specifically for resilience: in the cloud, most of the expensive layers — duplicate servers, separate sites — already exist and are shared among customers. Building the same thing in a single office would cost a fortune. For more: On-Premises PBX vs. Cloud.
Common mistakes
- "We have a backup" — but nobody ever tested it. A backup internet line that was never tried, or a UPS whose battery died two years ago.
- A backup that depends on the same thing. Two internet providers running through the same cable in the street, or a UPS plugged into the same outlet as the air conditioner.
- No fallback destination. The PBX works, but nobody set where to send calls when the office doesn't answer — so calls just ring into empty space.
- A fallback destination where nobody answers. Forwarding to the mobile of someone on vacation, or to a voicemail box nobody listens to.
How to build a small, simple plan
- Ask: what does an hour without a phone cost? For a small business, maybe a few lost calls. For a service center, or a nonprofit on a fundraising evening — much more.
- Set a fallback destination for every important number: the mobile of the person in charge, a branch, or a queue whose agents work from home.
- Put a small UPS on the router and the switch — it's cheap and prevents disconnections during short power outages.
- If the phone is the heart of the business, add a backup internet line from a different provider.
- Every few months, disconnect the internet for a few minutes and call the office from outside. See what happens.
The last test is the most important one on the list, and the one done least often. Five minutes of testing saves you a surprise on the day it matters.
A real-life example: a yeshiva with a dormitory sets it up so that when the office doesn't answer, calls go to the mobile of the mashgiach on duty, and at night to a voicemail box whose messages arrive by email. That way a parent who calls in the middle of the night about something urgent reaches someone, and regular messages wait until morning.
The layers table: what happens if
The best way to understand redundancy is to go layer by layer and ask, "What happens if this one fails?" The table summarizes the four layers, what can break in each one, who steps in, and what you feel:
| Layer | What goes down | Who keeps going | What you feel |
|---|---|---|---|
| Server | One server or an entire data center | Servers at other sites; phones re-register | At most, one call dropped mid-conversation and a redial |
| Provider network | One connection to the telephone network or to an internet provider | Parallel connections through other providers | Usually nothing |
| The office | Power, router, internet line | A UPS for the first few minutes, then a backup internet line | A short break if there is no backup; nothing if there is |
| The plan | The entire office is unavailable for a long time | The fallback destination: mobile, branch, voicemail box | Calls are answered on a mobile; the caller doesn't know |
Note the last column. As you go down the table, responsibility shifts to you, and so does the feeling of the failure. In the top two layers, the provider absorbs the failure for you. In the bottom two, you decide in advance how much of it will be felt, with relatively cheap decisions: a small UPS, a backup line, and a fallback destination set in the control panel.
A full example: a bad day at a service center
The service center of a small insurance company in Bnei Brak: eight agents, one queue, about 400 calls a day. This is what a day looks like when everything goes wrong, at a center that prepared in advance.
9:15 — A power blip in the street. The power drops out for three seconds and comes back. The agents' computers shut off and restart. The phones, the switch and the router are connected to a UPS and didn't notice anything. A call that was in progress continued. Without a UPS, all the phones would have rebooted for about a minute, and eight calls would have been cut off.
10:40 — The main internet line goes down. The fiber provider announces a regional outage. Within about half a minute, the router switches to the cellular backup line. The phones re-register with the PBX through the new line. Two calls were dropped; the callers dialed again and reached the queue. The cellular line is slower, so the center manager asks the agents not to watch videos until the fiber is back — so that bandwidth stays available for calls.
1:00 PM — A real power outage. This time for two hours. The UPS keeps the router and the switch running for about twenty minutes, and then shuts down. Now the office is completely cut off. But the cloud PBX is not: it sees that the office extensions are not registered, and activates the plan — the queue keeps receiving calls, and four agents who have the mobile app join the queue from their mobiles. Customers wait a little longer, but they get answered. Anyone who waits too long is offered a callback.
3:05 PM — The power returns. The phones come up, register, and the agents go back to their desks. In that day's call log you see: 391 calls, six that were cut off mid-call, and not a single call that went nowhere. Without the preparation, the same day would have ended with several hours of "busy" and dozens of customers looking for another company.
Quarterly checklist
Redundancy you don't test is an assumption, not a fact. Once a quarter, fifteen minutes:
- UPS — unplug it from the wall for two minutes. The router and the switch should stay on. A weakening battery gives less and less time, until one day it gives zero.
- Backup internet — pull the main line's cable out of the router and make calls out and in. Check how long the switchover took.
- The fallback destination — turn off the phones (or disconnect the switch) and call the main number from a mobile. Make sure the call reaches whoever is supposed to answer, and that they actually answer.
- The list itself — who is the mobile on duty? Do they still work for you? Is the number up to date in the control panel?
- The messages — if calls go to a voicemail box, who receives the messages by email, and do they read them?
Write down the results: the date, what was tested, how long the UPS lasted. After a year you have a real picture of what works and what is weakening — and you can replace a battery before it fails, not after.
The fifth layer: your data
The four layers deal with calls that keep flowing. There is one more layer that is easy to forget: settings and recordings. The cloud PBX keeps them for you, but some things are worth having on your side too:
- Settings documentation — a short document: which numbers you have, where each one is routed, who is in each queue, and what the fallback destination is. When someone new takes over managing the phones, this saves days. You can take the diagram from the call flow screen.
- Important recordings — recordings are deleted after the retention period. For a call that involves a dispute or a commitment, download it and save it separately.
- The call log — a periodic export to a file, for reports and bookkeeping.
- The audio files — the recordings of your menus and greetings. If you recorded a voice professional, keep the original files in an organized folder.
The principle is the same as for the other layers: don't rely on a single copy of anything important to you, even when that copy sits in a protected place.
How it works with us at Kesher
Kesher's PBX is built with backup and redundancy at every layer: servers in data centers in London, Paris, Rome, Athens and Moscow, alongside additional sites. A failure in one place does not take the service down.
On your side, we define in advance what happens when the office is unavailable — with no power or no internet. The PBX keeps receiving calls and forwards them to a mobile, another branch or a voicemail box, according to what is right for you.
In the control panel you can see, for each extension, whether it is registered and from which device, and in the call flow screen you see in a diagram where each call goes and which rule takes precedence. This is how you check that the plan really does what you thought.
When we set up a communications cabinet, the UPS is part of the design, alongside the router and the PoE switch. And if you'd like to review your backup plan together, talk to us and we'll go over it.
FAQ
What is the difference between resilience and redundancy?
Resilience is the outcome — the system keeps working. Redundancy is the method — an extra component for every role, which takes over when the first one fails.
If the office internet goes down, do customers hear a busy signal?
Not if a fallback destination has been set. The cloud PBX keeps receiving calls and forwards them to a mobile, a branch or a voicemail box.
Do you need a UPS with a cloud PBX too?
It's a good idea. A small UPS on the router and the switch keeps the phones running through short power outages and prevents reboots.
How do you check that the backup plan works?
Disconnect the office internet for a few minutes and call the main number from an outside phone. If the call reaches the fallback destination, the plan works.
How long does the UPS need to last?
For the router and the switch, 15 to 30 minutes is usually enough — that covers most short outages. For a longer outage, the plan is forwarding to a mobile, not a bigger battery.
Is a cellular backup internet line enough for calls?
Usually yes. A call uses little bandwidth, and a decent cellular line can hold several calls at once. Just make sure that during the backup, the bandwidth is kept for calls and not for heavy downloads.