Stats-240 notes
Lecture #1: 9/10
Chapter 2.1
➢ Visualizing the qualitative variable
○ Bar chart -One axis, each bar is a category
○ Other axis has either frequency of the category, or relative frequency of the
category
○ Relative frequency is frequency/total
○ In a bar chart, the points can be organized in any order
○ Pareta chart is a bar chart, whose bars are drawn in descending order of
frequency or relative frequency
○ Pie chart -Is a circular chart divided into sectors. Each sector represents one
category. The area of each sector is proportional to the frequency of each
category
○ A bar chart can tell us the total number of each frequency in given categories.
From a bar chart, as long as you are given the frequency of each, you are able
to figure out the total of each of them
○ Angle of unknown: take the sector, which will be the total distribution of that
category, out of the total, and then divide the sector over the total. For
example: there are 7 people in a class that are unknown in terms of the grades
they are in, 5 are seniors, 3 are freshman, 5 are juniors, you would take 7/20,
which is 0.35. From the 0.35, you divide by the total angle of unknown, (360,
entire circle), and then are able to figure out the angle of unknown, which then
will be able to find the number of people in the grade that you do not know
○ Bar chart vs. Pie chart:
■ Bar charts are easier to compare the frequency or relative frequency of
each category
■ Pie charts are more visual. However, some similar angles are difficult
to judge. They are not commonly used for comparing two specific
values of the categories
■ Organizing quantitative data
■ Construct a histogram. A histogram is constructed by drawing
rectangles for each class of data. The height of each rectangle is the
frequency or relative frequency. The width of each rectangle is the
same, and the rectangles touch each other
○ Classes: Classes are intervals into which data grouped when a data set
consists of a large number of quantitative data values, we create classes
by using intervals of numbers
○ Example: Suppose we have a variable: “age.” The data range of this
variable= 30 to 80. Solution: We group the data into classes.
○ The lower limit is the smallest value within the class
○ The upper value is the largest value within the class
, ○ Classes cannot be overlapped
○ Dot plot- is a second way to visualize quantitative data, count the
number of data points for each number, and put a dot out for each
value over that value
○ Different distributions have different shapes: Uniform (flat), bell
shaped, (symmetric), skewed right, and skewed left. Distribution
shapes will be dependent on how the data values are set up
○ Skewed distributions, can either be right skewed or left skewed,
depending on the direction that the data is set up
○ Left skewed means there are some extremely small values in the data
set
○ Right skewed means there are some extremely large values
3.1 Measures of a center tendency
Population vs. Sample:
Population is a group of people we wish to study
ALL individuals Example: All iphone users at Suffolk
Sample is a collection, (subset) of objects or people taken from population
All iphone users in STATS 240A
The arithmetic Mean: (average)
Population Mean: M= X1+X2+X3+,+,+Xn all divided by N
Sample mean = X bar
Sample mean x bar = X1+X2+,+,+,+,Xn all divided by N again
Suppose there is a small population data
23, 11, 25, 17, 19
Find the population mean (M)
M = 23+11+25+7+19 /5 =95/5 =19
Suppose there is a sample of test scores
82,72,90,62,74,68
Find the mean
, 82+72+90+62=74+68 Total of these= 448/6
74.667 =mean
Textbook Notes (Refer to these every single day when I do MY Stats Lab
Chapter1 Part 1
● Statistics is a lot to do with numbers and proportions of, how many of a certain group
out of the total, fit into a specific category
● Statistics defined: the science of collecting, organizing, summarizing, and analyzing
information to draw conclusions or answer questions.
● Stats has a lot to do with data, and data collection
● An entire group is called a population, one person is called an individual
● A statistic is a numerical summary of a sample. Descriptive statistics consist of
organizing and summarizing data
● Inferential statistics use methods that take a result from a sample, and extend it to a
population. They will then measure the reliability of the result
● Quantitative statistics are strictly numerical. Qualitative are information based
● A discrete variable is a quantitative variable that has either a finite number of values,
or a countable number of possible values
● Variables are the characteristics of the individuals within the population
● Obtaining a simple random sample: has to be done by selecting a random group of
people. No person can be pre-selected, because then it is not random
● Alternatively, in a sample with replacement, the same person can be sampled twice
● A cluster sample is obtained by selecting all individuals within a randomly selected
group
● A convenience sample is one where individuals are chosen because they are easily
accessible to the people sampling
● With collecting data, often times there are multiple different ways someone can go
about collecting it. It should always be a simple random samples, but the source one
uses to obtain the numbers has options
Lecture #2: 9/12
● Continue to read textbook later tonight and take notes. Try to get a significant portion
more done
𝑧 𝑥𝑖
● Measures of center: Sample mean𝑥=
𝑛
M= Z bar, X i/ N
● Median: A value that lies in the middle of the data when arranged in ascending order
● Steps to finding the median: 1) arrange the data in ascending order, 2) If we have an
odd number of observations, the middle value is the median. If we have an even
number of observations, the median is the average of the middle two values.
Lecture #1: 9/10
Chapter 2.1
➢ Visualizing the qualitative variable
○ Bar chart -One axis, each bar is a category
○ Other axis has either frequency of the category, or relative frequency of the
category
○ Relative frequency is frequency/total
○ In a bar chart, the points can be organized in any order
○ Pareta chart is a bar chart, whose bars are drawn in descending order of
frequency or relative frequency
○ Pie chart -Is a circular chart divided into sectors. Each sector represents one
category. The area of each sector is proportional to the frequency of each
category
○ A bar chart can tell us the total number of each frequency in given categories.
From a bar chart, as long as you are given the frequency of each, you are able
to figure out the total of each of them
○ Angle of unknown: take the sector, which will be the total distribution of that
category, out of the total, and then divide the sector over the total. For
example: there are 7 people in a class that are unknown in terms of the grades
they are in, 5 are seniors, 3 are freshman, 5 are juniors, you would take 7/20,
which is 0.35. From the 0.35, you divide by the total angle of unknown, (360,
entire circle), and then are able to figure out the angle of unknown, which then
will be able to find the number of people in the grade that you do not know
○ Bar chart vs. Pie chart:
■ Bar charts are easier to compare the frequency or relative frequency of
each category
■ Pie charts are more visual. However, some similar angles are difficult
to judge. They are not commonly used for comparing two specific
values of the categories
■ Organizing quantitative data
■ Construct a histogram. A histogram is constructed by drawing
rectangles for each class of data. The height of each rectangle is the
frequency or relative frequency. The width of each rectangle is the
same, and the rectangles touch each other
○ Classes: Classes are intervals into which data grouped when a data set
consists of a large number of quantitative data values, we create classes
by using intervals of numbers
○ Example: Suppose we have a variable: “age.” The data range of this
variable= 30 to 80. Solution: We group the data into classes.
○ The lower limit is the smallest value within the class
○ The upper value is the largest value within the class
, ○ Classes cannot be overlapped
○ Dot plot- is a second way to visualize quantitative data, count the
number of data points for each number, and put a dot out for each
value over that value
○ Different distributions have different shapes: Uniform (flat), bell
shaped, (symmetric), skewed right, and skewed left. Distribution
shapes will be dependent on how the data values are set up
○ Skewed distributions, can either be right skewed or left skewed,
depending on the direction that the data is set up
○ Left skewed means there are some extremely small values in the data
set
○ Right skewed means there are some extremely large values
3.1 Measures of a center tendency
Population vs. Sample:
Population is a group of people we wish to study
ALL individuals Example: All iphone users at Suffolk
Sample is a collection, (subset) of objects or people taken from population
All iphone users in STATS 240A
The arithmetic Mean: (average)
Population Mean: M= X1+X2+X3+,+,+Xn all divided by N
Sample mean = X bar
Sample mean x bar = X1+X2+,+,+,+,Xn all divided by N again
Suppose there is a small population data
23, 11, 25, 17, 19
Find the population mean (M)
M = 23+11+25+7+19 /5 =95/5 =19
Suppose there is a sample of test scores
82,72,90,62,74,68
Find the mean
, 82+72+90+62=74+68 Total of these= 448/6
74.667 =mean
Textbook Notes (Refer to these every single day when I do MY Stats Lab
Chapter1 Part 1
● Statistics is a lot to do with numbers and proportions of, how many of a certain group
out of the total, fit into a specific category
● Statistics defined: the science of collecting, organizing, summarizing, and analyzing
information to draw conclusions or answer questions.
● Stats has a lot to do with data, and data collection
● An entire group is called a population, one person is called an individual
● A statistic is a numerical summary of a sample. Descriptive statistics consist of
organizing and summarizing data
● Inferential statistics use methods that take a result from a sample, and extend it to a
population. They will then measure the reliability of the result
● Quantitative statistics are strictly numerical. Qualitative are information based
● A discrete variable is a quantitative variable that has either a finite number of values,
or a countable number of possible values
● Variables are the characteristics of the individuals within the population
● Obtaining a simple random sample: has to be done by selecting a random group of
people. No person can be pre-selected, because then it is not random
● Alternatively, in a sample with replacement, the same person can be sampled twice
● A cluster sample is obtained by selecting all individuals within a randomly selected
group
● A convenience sample is one where individuals are chosen because they are easily
accessible to the people sampling
● With collecting data, often times there are multiple different ways someone can go
about collecting it. It should always be a simple random samples, but the source one
uses to obtain the numbers has options
Lecture #2: 9/12
● Continue to read textbook later tonight and take notes. Try to get a significant portion
more done
𝑧 𝑥𝑖
● Measures of center: Sample mean𝑥=
𝑛
M= Z bar, X i/ N
● Median: A value that lies in the middle of the data when arranged in ascending order
● Steps to finding the median: 1) arrange the data in ascending order, 2) If we have an
odd number of observations, the middle value is the median. If we have an even
number of observations, the median is the average of the middle two values.