Skip to main content

Unit 4 · Topic 4.2

4.2 Fault Tolerance

Parts of a network fail all the time, but the internet keeps working. This topic covers redundancy and fault tolerance: how extra paths let data route around failures, what that costs, and how to find the weak spots in a network.

Key terms

  • fault tolerance
  • redundancy
  • path
  • rerouting
  • reliability

Redundancy and fault tolerance

Redundancy means including extra components that can take over if others fail. In a network, the most important kind is having more than one path between devices.

A system is fault tolerant if it can keep working even when some of its parts fail. This matters because parts of complex systems fail at unexpected times, often several at once, as when a storm knocks out a whole area. Fault tolerance lets people keep using the network anyway.

The internet was engineered to be fault tolerant. If a router or connection fails, later packets are sent along a different route, if one exists. Because routing is dynamic, this happens automatically.

Benefits and costs

Redundant routes make the internet more reliable, and they help it scale to more devices and more people, because traffic can spread across many paths.

Redundancy isn't free. Extra cables, routers and servers cost money to build and maintain. Designers weigh how much reliability they need against what it costs. A hospital or bank may pay for several backup connections; a small home network usually has just one.

The same idea shows up beyond network paths. Cloud services keep copies of your files on several servers in different places, and data centers have backup power, so one failed machine or a local power cut doesn't make your data disappear. Each copy or backup is a redundant component that costs money but keeps the system running.

Finding weak spots

When a question shows a network, look for single points of failure: one device or connection whose failure would cut some devices off. If every path between two devices passes through the same device, that device is a single point of failure for them.

To find how many connections must fail before two devices can't communicate, look for the smallest set of connections that, once removed, leaves no path between them. To make a network more fault tolerant, add connections that give a second route around its weak spots.

Picture this network, used in the worked example: devices A, B, C, D, E and F, with connections A–B, A–C, B–D, C–D, D–E, D–F and E–F. Writing out the paths from A to F is the first step to answering any question about it.

Worked examples

Try each one yourself first, then open the solution.

  1. Example 1

    Analyzing a network

    In the network with connections A–B, A–C, B–D, C–D, D–E, D–F and E–F: (a) list every path from A to F that doesn't repeat a device; (b) what is the smallest number of connections that could fail and leave A unable to reach F; (c) is there a single device whose failure would cut A off from F?

    Show the solution
    1. Step 1: (a) From A you can go to B or C; both lead only to D. From D you can go straight to F or through E. Paths: A–B–D–F, A–B–D–E–F, A–C–D–F, A–C–D–E–F.
    2. Step 2: (b) No single connection breaks every path: each one has a way around it. But two can. If A–B and A–C both fail, A is cut off. If D–F and E–F both fail, F is cut off. So the answer is 2.
    3. Step 3: (c) Every path goes through D. If D fails, A can't reach F. D is a single point of failure.
    4. Step 4: To fix that weak spot, add a connection that bypasses D, such as B–E or C–F.

    Answer: (a) Four paths: A–B–D–F, A–B–D–E–F, A–C–D–F, A–C–D–E–F. (b) 2 connections. (c) Yes, device D.

  2. Example 2

    Explaining fault tolerance

    A city's traffic cameras send video to a control center. Currently each camera has one connection to one router, and that router has one connection to the control center. Explain one vulnerability and one way to make the system more fault tolerant.

    Show the solution
    1. Step 1: Vulnerability: the single router and its single link are single points of failure. If either fails, no video reaches the center.
    2. Step 2: Improvement: add redundancy, such as a second router or a second connection to the control center, so data can be rerouted if one fails.
    3. Step 3: Mention the trade-off: the extra equipment costs more, but the cameras keep working during a failure.

    Answer: The single router is a single point of failure; adding a second router or a backup connection lets data reroute, at extra cost.

Common mistakes

  • Counting a path that repeats a device. Paths in these questions go from sender to receiver without looping back.
  • Thinking redundancy has no downside. It costs more to build and maintain.
  • Assuming that more connections always remove every single point of failure. Check whether every path still passes through one device.

On the exam

  • Network diagram questions are common: they ask how many connections can fail before two devices are cut off, or which added connection improves fault tolerance most. Label the diagram and check each choice.
  • Expect questions on why the internet is fault tolerant: redundant paths plus dynamic routing.

Connected topics

Videos

  • AP CSP Topic 4.2 - Routers and Redundancy - Code.org Unit 2.4 walkthrough - PLUS 4 practice MCQs!!

    Dr_WuWatch on YouTube (opens in a new tab)

  • The Internet: Packets, Routing & Reliability

    CodeAIWatch on YouTube (opens in a new tab)

  • Quick Bit: Redundancy and Fault Tolerance

    UTeach Computer ScienceWatch on YouTube (opens in a new tab)

  • Redundancy and Fault Tolerance (AP Computer Science Principles, Unit 2)

    Professor CunninghamWatch on YouTube (opens in a new tab)

  • AP CSP U2 L4 Routers and Redundancy

    Janelle WhalenWatch on YouTube (opens in a new tab)

  • AP CSP Topic 4.2 - Fault Tolerance - Speedrun! 43 MCQs!

    Dr_WuWatch on YouTube (opens in a new tab)

Check yourself

4 questions on 4.2 Fault Tolerance. Pick an answer to see if you got it, and why.

Network X connects five devices in a loop: P is connected to Q, Q to R, R to S, S to T, and T back to P.

Network Y connects the same five devices differently: P is connected to Q, R, S and T, and there are no other connections.

In both networks, data can travel either way along a connection and can pass through any number of devices.

Invented scenario

Question 1 of 4

Which statement best compares how fault tolerant the two networks are?

Question 2 of 4

In Network X, what is the minimum number of connections that must fail to stop device Q from communicating with device S?

Question 3 of 4

In Network Y, the failure of which single device would leave none of the remaining devices able to communicate with each other?

Question 4 of 4

Which of the following best describes a fault-tolerant system?

0 of 4 answered