Katemonkey (In Most Places)

Learning Statistics, Day 1 - Introduction to Statistics

Okay, I had a day of leisure and then I got bored. And I discovered that I could just continue with my data science excursions by actually studying statistics on W3Schools.

Actual statistics!

There are 36 separate lessons in this course, so I don't know exactly how long I'll stick with it, what with the next module in my Masters starting up at the end of the month (come on...Information Management and Visualisation...you'll be so good for me...), but I figure it'll help in general.

I mean, hell, if anything, it gives me a chance to play around with more Letterboxd data, huh?

Introduction to Statistics

Like I said before, I really should've taken a stats course while I was at university. But it was my first year, I was kinda frazzled, I knew I had to get my math credit out of the way, I went for symbolic logic.

I also ended up with mono that semester, so, really, it might have been a very good thing I did symbolic logic, because I spent pretty much all of February dead to the world.

(It was also my first Mardi Gras, dammit. Too sick for parades, booooooo.)

But that's in the past, and now I learn statistics. And, to be honest, it might be better to do it now, when it's a lot easier to play with data that I actually care about, instead of...hell, whatever they'd have to give me in 1996. Probably really boring data on really boring things. Like, I don't know. Crop yields. Census data. Whatever.

The three steps of statistics are:

They say "Knowing which questions you want to answer can help guide what sort of data you need." but I'm hoping they cover confirmation bias and the like later on.

They also have a long list of important concepts in statistics, which they say they'll cover. Some of these I've already covered in my Data Science, course, but it's good to have even more details.

Here's what I'll get to learn:

Apparently, also, along with using Python, the course uses R? I've never even heard of R, so that's going to be exciting.

Gathering Data

If I want to know something about an entire population, I need to pull together a sample that has the same characteristics of the population, but is small enough to gather information on.

I kinda like their example. If I want to know about all the people in France, I don't go and survey all the people of France. I gather a sample, making sure that I pull together the right people to make up a representative sample of French people, and then survey them.

If you only include people named Jacques living in Paris who are 48 years old, the sample will not be similar to the whole population.

Maybe I want to study the population of 48-year-old Parisians named Jacques, huh? Maybe that is my population! Maybe I want to know how many of them wear berets and stripy shirts! How many of them have little moustaches! How many of them are holding baguettes at any given time! Other French stereotypes!

Describing Data

Once I have my representative sample of Jacques, I can visualise their data with graphs or summarise it with numbers.

Graphs are a good visual way to show the data distribution. And lord knows I love a good graph. I have a hard time reading some of them, but that's probably more because I get distracted by colours and shapes, like a baby looking at a mobile.

They bring up some types of graphs, like histograms, pie charts, bar graphs, and box plots, but they also have a giant section on Descriptive Statistics, so I'll explain those later.

Summary statistics sum up large quantities of information into a few key values, like mean, median, mode, range, quartiles, percentiles, and standard deviation. Again, they'll get into these later, so I'll get into them later.

Making Conclusions

And once you've summarised and visualised the data, you can infer conclusions about the entire population. My tiny collection of Jacques stands in for all the 48-year-old Jacques in all of Paris.

Probability theory is used to calculate my certainty about these statistics, and my uncertainty is expressed as a confidence interval.

Ooh, do you think I can start talking like C-3PO and start telling Han Solo the odds?

There's also hypothesis testing, which I did in the Data Science lessons, when I was trying to figure out whether or not I could determine how many movies I'd like based on how many I'd seen in a given year.

That was also causal inference, when I was trying to figure out if those numbers correlated. They didn't, but that's not too surprising, because I watch a lot of movies, but that doesn't mean I like them.

Day 1 — Results

I think that'll do for today. Next time will be the rest of the basics of statistics before I start in on the different types of descriptive statistics. None of this is particularly difficult yet, but we'll see.

Today's Sticker

Hello Kitty is standing, holding a heart behind her back, because she has so much love to give.

Hey, maybe I should do a data analysis of all my stickers. How many Hello Kitty stickers do I have? How many of them have I used as preview images? How many of them are from the Panini sticker collection and how many are from a Miniso sticker pack like this one?

...

I think statistics could be a very dangerous thing for me to learn...

#kate learns statistics #statistics