A population is the entire group to which we want to generalize our results. A sample is a subset of the population that we actually study.
A numerical summary of a population is called a parameter, while the same numerical summary of a sample is called a statistic.
Sampling error is the sample-to-sample variability in a statistic. (Note that the term “error” here does not mean that anyone has made a mistake. “Error” here just refers to variability.) In general we do not know the population parameters that we are interested in. Instead, we have to draw conclusions about them based on sample statistics … usually from a single sample. In principle, any relationship that we observe in a sample could reflect nothing more than sampling error. In other words, this sampling error hypothesis is a possible explanation for any relationship that we observe in a sample. It could just be a random thing that happened for this particular sample, but it is not true in the population in general.
Central limit theorem states that
the random sampling distribution of means will always tend to be normal, irrespective of the shape of the population distribution from which the samples were drawn.
the mean of the random sampling distribution of means is equal to the mean of the original population.
Statistical Relationships
There are two basic kinds of statistical relationships- relationship between two group means - null hypothesis testing
- relationships between two quantitative variables - correlation and regression
- Correlation
- Regression
Types of Data: