English· Español· Deutsch· Nederlands· Français· 日本語· ქართული· 繁體中文· 简体中文· Português· Русский· العربية· हिन्दी· Italiano· 한국어· Polski· Svenska· Türkçe· Українська· Tiếng Việt· Bahasa Indonesia

nu

invité
1 / ?
retour aux leçons

Welcome

Reading Data Without Being Fooled

Numbers feel honest. A chart feels like proof. But a graph can be technically true and still trick your eyes.

Dashboards at work, charts in the news, and automated reports all show real numbers. The trouble hides in how those numbers get drawn: where an axis starts, which dates get shown, and who got counted.

Today you will learn four detective moves: read the spread behind an average, catch a misleading graph, question the sample, and demand a fair comparison.

Warm-up: think of a time a number or a chart surprised you (a score, a price, a headline, a game stat). What was it, and did you trust it right away?

An Average Is Not the Whole Story

The Average Hides the Spread

An average (the mean) adds up all the values and divides by how many there are. It gives one tidy number. That tidiness is the danger: one number cannot show how spread out the data is.

The range is the distance from the smallest value to the largest. The spread is how far apart the values sit in general.


Two swimming pools, average depth 1 meter each.

- Pool A: every part is 1 meter deep. Safe to wade.

- Pool B: the shallow end is 0.1 meter and the deep end is 3 meters. The average is still about 1 meter, but a small child could drown in the deep end.


Same average. Completely different reality. The average hid the spread. Whenever you see an average, ask: what is the range, and how spread out are the values?

The Pool Problem

Your Turn

A report says: The average depth of this pool is 1 meter, so it is safe for toddlers.

Why might this claim be misleading, even though the average is a real, correctly calculated number? Use the idea of range or spread in your answer.

Three Ways a Graph Lies

Two bar charts of the same data: a truncated y-axis starting at 100 exaggerates a tiny gap, while a full y-axis starting at 0 shows the bars are nearly equal

A Chart Can Be True and Still Lie

A graph can show correct numbers and still fool you. Three of the most common tricks:


1. Truncated (zoomed) y-axis. An honest bar chart starts its y-axis at zero. A misleading one starts partway up, say at 100. Look at the picture: Brand A sold 102 and Brand B sold 105, a difference of just 3 units. When the axis starts at 100, Brand B's bar looks about 2.5 times taller. Start the axis at 0 and the bars are nearly equal. The data never changed. Only the baseline did.


2. Cherry-picked time range. Show only the slice of dates that supports your story. A stock that fell all year looks great if you only chart the one good month. Ask: why does the chart start and end on those exact dates?


3. Changed or inconsistent scale. Uneven spacing between numbers, or a scale that switches partway, can stretch or squash changes. A line can be made to look like a cliff or a gentle hill using the same numbers.


The detective move is always the same: read the axis and the range before you trust the shape.

Why Is This Graph Misleading?

Spot the Trick

A news chart shows two bars: last month's crime count was 500, this month's is 515. But the y-axis starts at 495 instead of 0. On screen, this month's bar looks roughly four times taller than last month's. The headline reads: Crime Explodes!

Name the specific trick this graph uses, and explain its effect: what does the trick make the viewer believe that the real numbers do not support?

Asking a Few vs Asking Everyone

The Sample: A Few Standing In for the Whole

Often you cannot ask everyone, so you ask a sample: a smaller group meant to stand in for the whole group (the population).

A sample only tells the truth about the whole if it represents the whole. When it does not, it is a biased sample.


Examples of biased samples:

- A website asks its own visitors whether they like the website. The people who hate it already left.

- A survey about school lunch is handed out only to students leaving the pizza line.

- A poll runs only during a weekday afternoon, missing everyone at work.


Two questions expose most bias: Who got left out? and Would the left-out people answer differently? A big sample does not fix bias; a poll of a million wrong people is still wrong.

The Survey Problem

Your Turn

A gym posts: We surveyed 200 people at our gym on a Saturday morning. 95 percent say exercise is easy to fit into a busy week. So exercise is easy for almost everyone.

Is this sample biased? Explain who is likely missing from the sample, and why that makes the 95 percent claim untrustworthy for people in general.

Same Units, Same Baseline

What Makes a Comparison Fair

The last detective move is the simplest: when two things get compared, check that the comparison is fair.

A fair comparison uses the same units and the same baseline for both sides.


Unfair comparisons hide in plain sight:

- Different units. Our phone has 5000 of battery, theirs has only 12! One number is milliamp-hours, the other is hours. They cannot be compared.

- Different baseline. City A had 50 crimes, City B had 200, so City B is far more dangerous. But City B has 20 times the people. Per person, City A is worse. Compare rates (per 1000 people), not raw totals.

- Different time spans. A whole year of one thing against one month of another is not a fair race.


Before you accept any comparison, ask: same units? same baseline? same time span?

The City Problem

Your Turn

A post claims: City A reported 50 thefts last year. City B reported 200 thefts last year. City B is four times as dangerous. City A has 10,000 people. City B has 500,000 people.

Explain why this is not a fair comparison, and describe what you would compare instead to make it fair.

What Will You Remember?

One Last Thought

Numbers and charts are not lies by nature. They are tools, and like any tool they can be aimed well or badly. A dashboard, a news graphic, and an automated report can each be technically true and still leave you with a false picture.

You now have four detective moves: read the spread behind an average, check where a graph's axis starts and which range it shows, question who is in the sample, and demand the same units and baseline in any comparison.

Read the axis and the sample, not just the shape.

In one or two sentences, which of the four moves do you think you will use most, and where might you use it?