Katemonkey (In Most Places)

Learning Data Science, Day 4 – Introduction to Statistics

Oh, W3Schools, we're in for it now. We're doing Statistics, a course I should've taken at university but didn't because everyone said Symbolic Logic was easier, and I just needed a math credit since I didn't get a good enough grade on my AP Calculus test. And, after all, you don't really need Statistics when you decide that instead of finishing up that Anthropology degree, you're going to switch to Religious Traditions of the West, because then you don't have to take that mandatory Linguistics course at 9am or skip out on Magic and Religion in the Ancient World for some random Anthro seminar.

Okay, yeah, I was pretty much a slacker while I was at university. I'm not gonna lie.

But now my middle-aged ass is paying the price. And now... Statistics.

Descriptive Statistics

So if Statistics is the science of analysing data (as W3Schools says), then Descriptive Statistics summarises the details of the data set, including:

It also says "Etc.", which, c'mon, W3Schools, that is no help.

To see descriptive statistics in Python, we use .describe(), as we saw two days ago with our action figure data.

I feel like I should have more detailed data this time around. I mean, the two sets of action figure details worked really well for the initial study, and the bookshelf example worked really well for linear graphs, but since the example W3Schools is providing is a six-column many-row beast of health data, I think I need something a little...denser.

Figuring out my data

So after staring blankly at things around my house, and going "god, it'd take me forever to log all of this into a database", I remembered that Letterboxd lets you export out all your data as a CSV.

Yesss. There we go. Movie data!

Of course, this isn't exactly like the data they have for the example, which is all numbers all the time. It's pretty much just name of film, when I saw it, and if I rewatched it. Especially since I don't give ratings on films because then I end up overthinking things.

So it's going to be very basic data, but okay. Let's take my You, me, and Talking Pictures TV list, since that has over 400 films that I've watched on Talking Pictures TV (my favourite television channel forever and ever). And we'll just do films by decade. That goes from 1912 to 2017, so we're talking a good range.

The table is this:

1910s 1920s 1930s 1940s 1950s 1960s 1970s 1980s 1990s 2000s 2010s
3 3 11 44 100 111 88 44 3 1 1

Letterboxd does give me a list of all the films I've seen, but I have seen a lot of movies, so I don't want to do that. So, instead, I'll just count the ones I've seen since I started my Letterboxd account, which will be over 1,000 (did I tell you I've seen a lot of movies?). It does include the 400-odd in the Talking Pictures TV list, but since it also has 600-odd more, I figure it gives me a good guide.

Organise that by year, and my data table looks like this:

1910s 1920s 1930s 1940s 1950s 1960s 1970s 1980s 1990s 2000s 2010s 2020s
3 3 11 44 100 111 88 44 3 1 1 0
3 4 16 56 123 142 132 116 34 26 189 174

(By the way, I then organised it by name and looked to see what movie I've rewatched the most in the 9 years I've had a Letterboxd account?

It was Dune. Of course it was Dune. Sure, The Thing and The Wicker Man came close, but it's Dune or nothing.)

So there we go. I now have a table. With some data. And it's good data.

Using .describe()

And now that I've turned that data into a CSV, it's time to actually do .describe().

So first I import in Pandas

import pandas as pd

Then I tell Pandas to read the CSV and set the maximum columns and rows to none.

film_data = pd.read_csv("data.csv", header=0, sep=",")
pd.set_option('display.max_columns',None)
pd.set_option('display.max_rows',None)

Finally, it's time to print!

print (film_data.describe())

And there we go, the descriptive statistics of films.

       1910s     1920s      1930s      1940s       1950s      1960s  \
count    2.0  2.000000   2.000000   2.000000    2.000000    2.00000   
mean     3.0  3.500000  13.500000  50.000000  111.500000  126.50000   
std      0.0  0.707107   3.535534   8.485281   16.263456   21.92031   
min      3.0  3.000000  11.000000  44.000000  100.000000  111.00000   
25%      3.0  3.250000  12.250000  47.000000  105.750000  118.75000   
50%      3.0  3.500000  13.500000  50.000000  111.500000  126.50000   
75%      3.0  3.750000  14.750000  53.000000  117.250000  134.25000   
max      3.0  4.000000  16.000000  56.000000  123.000000  142.00000   

            1970s       1980s     1990s     2000s       2010s      2020s  
count    2.000000    2.000000   2.00000   2.00000    2.000000    2.00000  
mean   110.000000   80.000000  18.50000  13.50000   95.000000   87.00000  
std     31.112698   50.911688  21.92031  17.67767  132.936075  123.03658  
min     88.000000   44.000000   3.00000   1.00000    1.000000    0.00000  
25%     99.000000   62.000000  10.75000   7.25000   48.000000   43.50000  
50%    110.000000   80.000000  18.50000  13.50000   95.000000   87.00000  
75%    121.000000   98.000000  26.25000  19.75000  142.000000  130.50000  
max    132.000000  116.000000  34.00000  26.00000  189.000000  174.00000

Now the course says

Do you see anything interesting here?

And...not really? I guess? I don't know? Whatever. I now know that for films in the 80s, I watched a minimum of 44 and a maximum of 116, but the average was 80. I watched on average 80 films set in the 80s. Haaaa.

Percentiles

We then get into percentiles, which is roughly the number that you would be at for a given percentage. So if I were to watch 25% of movies from the 2020s that I've watched, I'd be watching 43.5 movies. Maybe I fell asleep during The Odyssey or something. Who knows.

If I want to find the 10% percentile, I can do it using NumPy.

import pandas as pd
import numpy as np

film_data = pd.read_csv("data.csv", header=0, sep=",")

double_20s = film_data["2020s"]
percentile10 = np.percentile(double_20s, 10)

print(percentile10)

And I get 17.400000000000002. So I didn't fall asleep during halfway through the Odyssey, I fell asleep about an hour in, and I watched 17 other films from the 2020s.

Hmm, and based on the list I have, I'm going to say those 17 were...

Goddamn, I love movies.

Day 4 — Results

Tomorrow, I'll tackle Standard Deviation and Variance. I might do Correlation, but that's broken up into Correlation, Correlation Matrix, and Correlation vs Causality, so I might hold off and tackle those on Thursday, because otherwise, my braaaaains.

Today's Sticker

A drawing of Godzilla smiling widely and holding up a trans flag

Godzilla said TRANS RIGHTS!

I wish I could remember the name of the artist I got this from. There was an entire little parade of Godzillas holding flags. I have them all and I love them all very dearly.

If you haven't seen Coming Out, you really really need to. Oh my god, it is short, it is adorable, it is perfect, and it's excellent stop-motion animation. WATCH IT NOW.

#data science #kate learns data science #kate learns python #programming