Katemonkey (In Most Places)

Learning Data Science, Day 5 – Standard Deviation, Coefficient of Variation, and Variance

Yesterday was a total write-off, and today wasn't looking so hot either. I just got a Fairphone, and I'm doing all the transferring over and trying to figure out how to make things like they were on my Samsung and going "Why the hell did you duplicate everything on my SD card to the actual phone" and "Where the hell are my podcasts" and "Why aren't the widgets loading" and...well...I'm going to take a break before I end up breaking something.

So, instead, more data science! Because that won't annoy me at all!

Standard Deviation

Standard deviation is apparently "a number that describes how spread out the observations are".

That means nothing to me.

But, actually, what they mean is that maths has difficulties predicting precise values. Like, sure, you know that you watched three movies from the 1910s on Talking Pictures TV, but even if you predict that, in the future, you'll watch three more movies, that doesn't mean you actually will. You might watch one. Or none. Or they'll have an entire Buster Keaton marathon and you'll end up watching 11 films from the 1910s that he did.

But we can roughly guess what we'd get, and the standard deviation is the amount that it can vary.

So, in this case, since we have our handy dandy Letterboxd data from last time, we can find out the standard deviation for all our decades by using NumPy.

First we import in Pandas to read the CSV and NumPy to do the maths.

import pandas as pd
import numpy as np

Then we tell Pandas to read the CSV.

film_data = pd.read_csv("data.csv", header=0, sep=",")

Then we tell NumPy to give us the standard deviation of film_data and print the variable that stores that information.

std = np.std(film_data)
print(std)

And there we go, we get our standard deviation for each decade.

1910s     0.0
1920s     0.5
1930s     2.5
1940s     6.0
1950s    11.5
1960s    15.5
1970s    22.0
1980s    36.0
1990s    15.5
2000s    12.5
2010s    94.0
2020s    87.0

I like the 2010s and 2020s standard deviation, because, yeah, sure, it's unlikely I'll see films from those decades on Talking Pictures TV, but, as Borley Rectory shows, it's not impossible.

Coefficient of Variation

Now we need to know how large the standard deviation is compared to the data. This is called the coefficient of variation, where we divide the standard deviation by the mean.

So if we're like "well, I watched 111.5 movies from the 1950s on Talking Pictures TV and other sources, and the standard deviation is 11.5, that means the standard deviation would be around 10%. Which isn't bad but it's not that amazing."

And to do that in Python, it's NumPy again.

import pandas as pd
import numpy as np

film_data = pd.read_csv("data.csv", header=0, sep=",")

cv = np.std(film_data) / np.mean(film_data)

print(cv)

Which gives us

1910s    0.000000
1920s    0.008427
1930s    0.042135
1940s    0.101124
1950s    0.193820
1960s    0.261236
1970s    0.370787
1980s    0.606742
1990s    0.261236
2000s    0.210674
2010s    1.584270
2020s    1.466292

So the 2010s and the 2020s have a very high standard deviation rate, but when we go back 100 years, we have a very very low standard deviation rate.

This isn't perfect, because, like I said last time, the Talking Pictures TV data is included in the general film watching data.

So let's pull slightly different data from my Letterboxd. This time, we're taking my 2018-2026 viewing data, and matching it up with films that I officially "loved" on Letterboxd.

(If I rated films, this would provide me with some amazing data to play with, but I don't, because I'd overthink it, so...eh. Let's play with what we have.)

(This would also be immensely easier if they included "loves" on the diary spreadsheet, but, eh, we can't have everything.)

After a lot of spreadsheet finagling, because I still don't actually know how to combine spreadsheets and make them do weird and wonderful things with just formulas (I know it'll happen sometime, but, right now, it's still just copy/paste/move), I now have my new data table.

Decade 1910s 1920s 1930s 1940s 1950s 1960s 1970s 1980s 1990s 2000s 2010s 2020s
Watched 3 4 16 56 123 142 132 116 34 26 189 174
Loved 0 0 0 3 3 6 5 17 11 3 12 10

So now when we run the standard deviation, we get:

1910s     1.5
1920s     2.0
1930s     8.0
1940s    26.5
1950s    60.0
1960s    68.0
1970s    63.5
1980s    49.5
1990s    11.5
2000s    11.5
2010s    88.5
2020s    82.0

Which, as we can see, is different from our previous data set, and a bit more obvious as to what would happen.

And when we do the coefficient of variation:

1910s    0.033180
1920s    0.044240
1930s    0.176959
1940s    0.586175
1950s    1.327189
1960s    1.504147
1970s    1.404608
1980s    1.094931
1990s    0.254378
2000s    0.254378
2010s    1.957604
2020s    1.813825

So now, I know that my standard deviation for the 1990s is lower than my standard deviation for the 2010s. That could just be because I watched more films from the 2010s during this period, but it could also be that the 1990s did have a lot of films I love.

(And considering the 1990s included Jurassic Park, Ghost in the Shell, Bound, Jackie Brown, and Showgirls? I think we know the truth.)

I did wonder what my standard deviation and coefficient of variation would be if I did "Movies I love" against "Movies I've seen on Talking Pictures TV", and...

Standard deviation:

1910s     1.5
1920s     1.5
1930s     5.5
1940s    20.5
1950s    48.5
1960s    52.5
1970s    41.5
1980s    13.5
1990s     4.0
2000s     1.0
2010s     5.5
2020s     5.0

Coefficient of variation:

1910s    0.075157
1920s    0.075157
1930s    0.275574
1940s    1.027140
1950s    2.430063
1960s    2.630480
1970s    2.079332
1980s    0.676409
1990s    0.200418
2000s    0.050104
2010s    0.275574
2020s    0.250522

I think this is probably volume rather than quality. Like, Talking Pictures TV shows a lot of short films, and while I do love them, I don't love Love them. So I'm more likely to watch something from the 1950s and not love it than I am from the 1930s.

And considering my first loved film in the list was 1942's The Black Swan? Yeah, the data's not amazing, but it's still fun.

Variance

Variance is how spread out the values are.

There's a lot of maths involved here. They go step by step, which I'll do, but I'm only going to have two numbers in my data set, because otherwise I will lose the plot.

Our data is 3 and 8. There we go.

  1. Find the mean for the data - (3+8) / 2 = 5.5
  2. Find the difference from the mean for each value - 3 - 5.5 = -2.5, 8 - 5.5 = 2.5
  3. Find the square value for each difference - -2.5 ^ 2 = -6.25, 2.5 ^ 2 = 6.25
  4. Sum the squared values and find the average - (-6.25 + 6.25) / 2 = 0

The variance, in my example, is 0, and I cannot even begin to think about doing that with more detailed data, my head is already swimming, but, thank god, NumPy has one ridiculously simple command to do all that - .var().

import pandas as pd
import numpy as np
film_data = pd.read_csv("data.csv", header=0, sep=",")
var = np.var(film_data)
print(var)

We pop our Watched/Loved data into the CSV, and get...

1910s       2.25
1920s       4.00
1930s      64.00
1940s     702.25
1950s    3600.00
1960s    4624.00
1970s    4032.25
1980s    2450.25
1990s     132.25
2000s     132.25
2010s    7832.25
2020s    6724.00

And, yeah. The 2020s are very spread out. I really need to figure out what variance means, or if I should add more data. Or something. Because those are some weird-ass numbers there.

I try it with my Talking Pictures TV/Loved data, and get

1910s       2.25
1920s       2.25
1930s      30.25
1940s     420.25
1950s    2352.25
1960s    2756.25
1970s    1722.25
1980s     182.25
1990s      16.00
2000s       1.00
2010s      30.25
2020s      25.00

Huh. I need to really dig into what variance is. But maybe when I'm not hungry and tired and watching my old phone slowwwwwwwly move things over.

Day 4 — Results

I am definitely too fried for the correlation bits today. Tomorrow. Maybe. If my phone behaves.

Today's Sticker

An illustration of a cat winking, coloured like the rainbow pride flag

New photo from new phone! And because new phone means new stickers on the back, I get to take a photo of the one I used to have. My megomobile This 90s Cat Is Gay sticker!

I do rather love this sticker. I am tempted to put it back on my new phone, but, at the same time, I did put Every Day I'm Grunklin, and that is definitely more of how I am right now.

#data science #kate learns data science #kate learns python #programming