Learning Data Science, Day 5 – Standard Deviation, Coefficient of Variation, and Variance
Yesterday was a total write-off, and today wasn't looking so hot either. I just got a Fairphone, and I'm doing all the transferring over and trying to figure out how to make things like they were on my Samsung and going "Why the hell did you duplicate everything on my SD card to the actual phone" and "Where the hell are my podcasts" and "Why aren't the widgets loading" and...well...I'm going to take a break before I end up breaking something.
So, instead, more data science! Because that won't annoy me at all!
Standard Deviation
Standard deviation is apparently "a number that describes how spread out the observations are".
That means nothing to me.
But, actually, what they mean is that maths has difficulties predicting precise values. Like, sure, you know that you watched three movies from the 1910s on Talking Pictures TV, but even if you predict that, in the future, you'll watch three more movies, that doesn't mean you actually will. You might watch one. Or none. Or they'll have an entire Buster Keaton marathon and you'll end up watching 11 films from the 1910s that he did.
But we can roughly guess what we'd get, and the standard deviation is the amount that it can vary.
So, in this case, since we have our handy dandy Letterboxd data from last time, we can find out the standard deviation for all our decades by using NumPy.
First we import in Pandas to read the CSV and NumPy to do the maths.
import pandas as pd
import numpy as np
Then we tell Pandas to read the CSV.
film_data = pd.read_csv("data.csv", header=0, sep=",")
Then we tell NumPy to give us the standard deviation of film_data and print the variable that stores that information.
std = np.std(film_data)
print(std)
And there we go, we get our standard deviation for each decade.
1910s 0.0
1920s 0.5
1930s 2.5
1940s 6.0
1950s 11.5
1960s 15.5
1970s 22.0
1980s 36.0
1990s 15.5
2000s 12.5
2010s 94.0
2020s 87.0
I like the 2010s and 2020s standard deviation, because, yeah, sure, it's unlikely I'll see films from those decades on Talking Pictures TV, but, as Borley Rectory shows, it's not impossible.
Coefficient of Variation
Now we need to know how large the standard deviation is compared to the data. This is called the coefficient of variation, where we divide the standard deviation by the mean.
So if we're like "well, I watched 111.5 movies from the 1950s on Talking Pictures TV and other sources, and the standard deviation is 11.5, that means the standard deviation would be around 10%. Which isn't bad but it's not that amazing."
And to do that in Python, it's NumPy again.
import pandas as pd
import numpy as np
film_data = pd.read_csv("data.csv", header=0, sep=",")
cv = np.std(film_data) / np.mean(film_data)
print(cv)
Which gives us
1910s 0.000000
1920s 0.008427
1930s 0.042135
1940s 0.101124
1950s 0.193820
1960s 0.261236
1970s 0.370787
1980s 0.606742
1990s 0.261236
2000s 0.210674
2010s 1.584270
2020s 1.466292
So the 2010s and the 2020s have a very high standard deviation rate, but when we go back 100 years, we have a very very low standard deviation rate.
This isn't perfect, because, like I said last time, the Talking Pictures TV data is included in the general film watching data.
So let's pull slightly different data from my Letterboxd. This time, we're taking my 2018-2026 viewing data, and matching it up with films that I officially "loved" on Letterboxd.
(If I rated films, this would provide me with some amazing data to play with, but I don't, because I'd overthink it, so...eh. Let's play with what we have.)
(This would also be immensely easier if they included "loves" on the diary spreadsheet, but, eh, we can't have everything.)
After a lot of spreadsheet finagling, because I still don't actually know how to combine spreadsheets and make them do weird and wonderful things with just formulas (I know it'll happen sometime, but, right now, it's still just copy/paste/move), I now have my new data table.
| Decade | 1910s | 1920s | 1930s | 1940s | 1950s | 1960s | 1970s | 1980s | 1990s | 2000s | 2010s | 2020s |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Watched | 3 | 4 | 16 | 56 | 123 | 142 | 132 | 116 | 34 | 26 | 189 | 174 |
| Loved | 0 | 0 | 0 | 3 | 3 | 6 | 5 | 17 | 11 | 3 | 12 | 10 |
So now when we run the standard deviation, we get:
1910s 1.5
1920s 2.0
1930s 8.0
1940s 26.5
1950s 60.0
1960s 68.0
1970s 63.5
1980s 49.5
1990s 11.5
2000s 11.5
2010s 88.5
2020s 82.0
Which, as we can see, is different from our previous data set, and a bit more obvious as to what would happen.
And when we do the coefficient of variation:
1910s 0.033180
1920s 0.044240
1930s 0.176959
1940s 0.586175
1950s 1.327189
1960s 1.504147
1970s 1.404608
1980s 1.094931
1990s 0.254378
2000s 0.254378
2010s 1.957604
2020s 1.813825
So now, I know that my standard deviation for the 1990s is lower than my standard deviation for the 2010s. That could just be because I watched more films from the 2010s during this period, but it could also be that the 1990s did have a lot of films I love.
(And considering the 1990s included Jurassic Park, Ghost in the Shell, Bound, Jackie Brown, and Showgirls? I think we know the truth.)
I did wonder what my standard deviation and coefficient of variation would be if I did "Movies I love" against "Movies I've seen on Talking Pictures TV", and...
Standard deviation:
1910s 1.5
1920s 1.5
1930s 5.5
1940s 20.5
1950s 48.5
1960s 52.5
1970s 41.5
1980s 13.5
1990s 4.0
2000s 1.0
2010s 5.5
2020s 5.0
Coefficient of variation:
1910s 0.075157
1920s 0.075157
1930s 0.275574
1940s 1.027140
1950s 2.430063
1960s 2.630480
1970s 2.079332
1980s 0.676409
1990s 0.200418
2000s 0.050104
2010s 0.275574
2020s 0.250522
I think this is probably volume rather than quality. Like, Talking Pictures TV shows a lot of short films, and while I do love them, I don't love Love them. So I'm more likely to watch something from the 1950s and not love it than I am from the 1930s.
And considering my first loved film in the list was 1942's The Black Swan? Yeah, the data's not amazing, but it's still fun.
Variance
Variance is how spread out the values are.
There's a lot of maths involved here. They go step by step, which I'll do, but I'm only going to have two numbers in my data set, because otherwise I will lose the plot.
Our data is 3 and 8. There we go.
- Find the mean for the data -
(3+8) / 2 = 5.5 - Find the difference from the mean for each value -
3 - 5.5 = -2.5, 8 - 5.5 = 2.5 - Find the square value for each difference -
-2.5 ^ 2 = -6.25, 2.5 ^ 2 = 6.25 - Sum the squared values and find the average -
(-6.25 + 6.25) / 2 = 0
The variance, in my example, is 0, and I cannot even begin to think about doing that with more detailed data, my head is already swimming, but, thank god, NumPy has one ridiculously simple command to do all that - .var().
import pandas as pd
import numpy as np
film_data = pd.read_csv("data.csv", header=0, sep=",")
var = np.var(film_data)
print(var)
We pop our Watched/Loved data into the CSV, and get...
1910s 2.25
1920s 4.00
1930s 64.00
1940s 702.25
1950s 3600.00
1960s 4624.00
1970s 4032.25
1980s 2450.25
1990s 132.25
2000s 132.25
2010s 7832.25
2020s 6724.00
And, yeah. The 2020s are very spread out. I really need to figure out what variance means, or if I should add more data. Or something. Because those are some weird-ass numbers there.
I try it with my Talking Pictures TV/Loved data, and get
1910s 2.25
1920s 2.25
1930s 30.25
1940s 420.25
1950s 2352.25
1960s 2756.25
1970s 1722.25
1980s 182.25
1990s 16.00
2000s 1.00
2010s 30.25
2020s 25.00
Huh. I need to really dig into what variance is. But maybe when I'm not hungry and tired and watching my old phone slowwwwwwwly move things over.
Day 4 — Results
- Standard deviation is the range that your prediction can vary by.
- You use NumPy to get the standard deviation by using
.std(). - The coefficient of variation is how large the standard deviation is compared to the data. That's the standard deviation divided by the mean.
- You use NumPy again, and this time it's
.std() / .mean(). - Variance is how spread out the data values are.
- It's a lot of maths, but in NumPy, it's just
.var().
I am definitely too fried for the correlation bits today. Tomorrow. Maybe. If my phone behaves.
Today's Sticker

New photo from new phone! And because new phone means new stickers on the back, I get to take a photo of the one I used to have. My megomobile This 90s Cat Is Gay sticker!
I do rather love this sticker. I am tempted to put it back on my new phone, but, at the same time, I did put Every Day I'm Grunklin, and that is definitely more of how I am right now.