Learning Data Science, Day 4 – Introduction to Statistics
Oh, W3Schools, we're in for it now. We're doing Statistics, a course I should've taken at university but didn't because everyone said Symbolic Logic was easier, and I just needed a math credit since I didn't get a good enough grade on my AP Calculus test. And, after all, you don't really need Statistics when you decide that instead of finishing up that Anthropology degree, you're going to switch to Religious Traditions of the West, because then you don't have to take that mandatory Linguistics course at 9am or skip out on Magic and Religion in the Ancient World for some random Anthro seminar.
Okay, yeah, I was pretty much a slacker while I was at university. I'm not gonna lie.
But now my middle-aged ass is paying the price. And now... Statistics.
Descriptive Statistics
So if Statistics is the science of analysing data (as W3Schools says), then Descriptive Statistics summarises the details of the data set, including:
- Count
- Sum
- Standard Deviation
- Percentile
- Average
It also says "Etc.", which, c'mon, W3Schools, that is no help.
To see descriptive statistics in Python, we use .describe(), as we saw two days ago with our action figure data.
I feel like I should have more detailed data this time around. I mean, the two sets of action figure details worked really well for the initial study, and the bookshelf example worked really well for linear graphs, but since the example W3Schools is providing is a six-column many-row beast of health data, I think I need something a little...denser.
Figuring out my data
So after staring blankly at things around my house, and going "god, it'd take me forever to log all of this into a database", I remembered that Letterboxd lets you export out all your data as a CSV.
Yesss. There we go. Movie data!
Of course, this isn't exactly like the data they have for the example, which is all numbers all the time. It's pretty much just name of film, when I saw it, and if I rewatched it. Especially since I don't give ratings on films because then I end up overthinking things.
So it's going to be very basic data, but okay. Let's take my You, me, and Talking Pictures TV list, since that has over 400 films that I've watched on Talking Pictures TV (my favourite television channel forever and ever). And we'll just do films by decade. That goes from 1912 to 2017, so we're talking a good range.
The table is this:
| 1910s | 1920s | 1930s | 1940s | 1950s | 1960s | 1970s | 1980s | 1990s | 2000s | 2010s |
|---|---|---|---|---|---|---|---|---|---|---|
| 3 | 3 | 11 | 44 | 100 | 111 | 88 | 44 | 3 | 1 | 1 |
Letterboxd does give me a list of all the films I've seen, but I have seen a lot of movies, so I don't want to do that. So, instead, I'll just count the ones I've seen since I started my Letterboxd account, which will be over 1,000 (did I tell you I've seen a lot of movies?). It does include the 400-odd in the Talking Pictures TV list, but since it also has 600-odd more, I figure it gives me a good guide.
Organise that by year, and my data table looks like this:
| 1910s | 1920s | 1930s | 1940s | 1950s | 1960s | 1970s | 1980s | 1990s | 2000s | 2010s | 2020s |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 3 | 3 | 11 | 44 | 100 | 111 | 88 | 44 | 3 | 1 | 1 | 0 |
| 3 | 4 | 16 | 56 | 123 | 142 | 132 | 116 | 34 | 26 | 189 | 174 |
(By the way, I then organised it by name and looked to see what movie I've rewatched the most in the 9 years I've had a Letterboxd account?
It was Dune. Of course it was Dune. Sure, The Thing and The Wicker Man came close, but it's Dune or nothing.)
So there we go. I now have a table. With some data. And it's good data.
Using .describe()
And now that I've turned that data into a CSV, it's time to actually do .describe().
So first I import in Pandas
import pandas as pd
Then I tell Pandas to read the CSV and set the maximum columns and rows to none.
film_data = pd.read_csv("data.csv", header=0, sep=",")
pd.set_option('display.max_columns',None)
pd.set_option('display.max_rows',None)
Finally, it's time to print!
print (film_data.describe())
And there we go, the descriptive statistics of films.
1910s 1920s 1930s 1940s 1950s 1960s \
count 2.0 2.000000 2.000000 2.000000 2.000000 2.00000
mean 3.0 3.500000 13.500000 50.000000 111.500000 126.50000
std 0.0 0.707107 3.535534 8.485281 16.263456 21.92031
min 3.0 3.000000 11.000000 44.000000 100.000000 111.00000
25% 3.0 3.250000 12.250000 47.000000 105.750000 118.75000
50% 3.0 3.500000 13.500000 50.000000 111.500000 126.50000
75% 3.0 3.750000 14.750000 53.000000 117.250000 134.25000
max 3.0 4.000000 16.000000 56.000000 123.000000 142.00000
1970s 1980s 1990s 2000s 2010s 2020s
count 2.000000 2.000000 2.00000 2.00000 2.000000 2.00000
mean 110.000000 80.000000 18.50000 13.50000 95.000000 87.00000
std 31.112698 50.911688 21.92031 17.67767 132.936075 123.03658
min 88.000000 44.000000 3.00000 1.00000 1.000000 0.00000
25% 99.000000 62.000000 10.75000 7.25000 48.000000 43.50000
50% 110.000000 80.000000 18.50000 13.50000 95.000000 87.00000
75% 121.000000 98.000000 26.25000 19.75000 142.000000 130.50000
max 132.000000 116.000000 34.00000 26.00000 189.000000 174.00000
Now the course says
Do you see anything interesting here?
And...not really? I guess? I don't know? Whatever. I now know that for films in the 80s, I watched a minimum of 44 and a maximum of 116, but the average was 80. I watched on average 80 films set in the 80s. Haaaa.
Percentiles
We then get into percentiles, which is roughly the number that you would be at for a given percentage. So if I were to watch 25% of movies from the 2020s that I've watched, I'd be watching 43.5 movies. Maybe I fell asleep during The Odyssey or something. Who knows.
If I want to find the 10% percentile, I can do it using NumPy.
import pandas as pd
import numpy as np
film_data = pd.read_csv("data.csv", header=0, sep=",")
double_20s = film_data["2020s"]
percentile10 = np.percentile(double_20s, 10)
print(percentile10)
And I get 17.400000000000002. So I didn't fall asleep during halfway through the Odyssey, I fell asleep about an hour in, and I watched 17 other films from the 2020s.
Hmm, and based on the list I have, I'm going to say those 17 were...
- Annette
- Dune
- Coming Out
- Everything Everywhere All At Once
- Godzilla Minus One
- Hundreds of Beavers
- Last Night in Soho
- Late Night with the Devil
- The People's Joker
- Pig
- Pillion
- Prey
- RoboDoc: The Creation of Robocop
- Summer of Soul (...Or When the Revolution Could Not Be Televised)
- The Unbearable Weight of Massive Talent
- The Viewing
- X
Goddamn, I love movies.
Day 4 — Results
- Statistics is just analysing data scientifically.
- Descriptive statistics summarise the details of the data set.
- This includes Count, Sum, Standard Deviation, Percentile, and Average.
- To get descriptive statistics in Python, you use
.describe(). - You need to import in Pandas first in order to read the CSV.
.describe()gives you Count, Mean, Standard Deviation, Minimum, Maximum, and the 25%, 50%, and 75% Percentiles.- When you want to know a different percentile, you can use NumPy and
.percentile(). - Just import NumPy, and then use
NUMPYNAME.percentile(COLUMNNAME,PERCENTAGE). - I sure have seen Dune a lot.
Tomorrow, I'll tackle Standard Deviation and Variance. I might do Correlation, but that's broken up into Correlation, Correlation Matrix, and Correlation vs Causality, so I might hold off and tackle those on Thursday, because otherwise, my braaaaains.
Today's Sticker

Godzilla said TRANS RIGHTS!
I wish I could remember the name of the artist I got this from. There was an entire little parade of Godzillas holding flags. I have them all and I love them all very dearly.
If you haven't seen Coming Out, you really really need to. Oh my god, it is short, it is adorable, it is perfect, and it's excellent stop-motion animation. WATCH IT NOW.