POS 3713 EXAM 3 #Questions with correct
answers
How do we calculate predicted values of Y if we know the values of the IVs, alpha,
and betas? - ✔✔Use regression to predict values of Y; regression gives you
coefficients, which estimate how much of an effect X has on Y.
To predict values of Y using IV's, alphas, and betas, we use the systematic (not
stochastic) components and equation. YHat= alpha + beta (Xi); we leave out the
random component (ui), because it defines the variation in Y that is NOT related
to the variation in x.
Alpha - ✔✔called the constant/intercept. Measures the value where the
regression line crosses the y-axis.
Beta - ✔✔called the coefficient/slope. Measures the steepness of the regression
line. A one unit change in X results in a beta change in (the expected value) of Y.
R-squared - ✔✔ranges from 0-1, measures the 'quality of the fit.' Measures the
proportion of the variance (Y) that the model explains. Ex; an r-squared of 1.0
would mean the model fits the data perfectly, with a line going right through
every data point. If you get an r-squared of around .85, you'd conclude that 85%
of the data is explained by pure variation, the other 15% is due to 'luck'.
RMSE - ✔✔Root mean squared error. *The difference between each observed
value of the DV and the predicted value of the DV. Indicates the absolute fit of the
model to the data points—how close the observed data points are to the model's
predicted values. Good measure of how accurately the model physically predicts
the results; measure of the typical deviations from the regression line. Used to
compare models. The lower the values of RMSE, the better the fit.
, Standard errors - ✔✔The overall measure of uncertainty of our data/model.
Smaller standard errors of alpha and beta mean more confidence in the estimated
values.
What is the difference between a sample and a population and why (and to what
extent) do we care about each? - ✔✔Sample—a subset of cases drawn from an
underlying population.
Population—data that is considered to be relevant for every possible case.
These two are important because in order to have a successful, non-
skewed/altered OLS, the population we draw from must be general enough to
represent the case, and the sample must be random.
What is a p-value? - ✔✔The probability that we would see the relationship that
we are finding because of purely random chance. Ranges from 0-1. It's based on
the assumption that we are drawing a perfectly random sample from the
population. *The lower the p-value, the greater confidence we have that there IS
a systematic relationship between the two variables. *does NOT tell us whether
the relationship is causal. **if our p-value is <.05, it is statistically significant, and
we'd reject the null.
What is a null hypothesis and how does it relate to the p-value? What is an
alternative hypothesis? - ✔✔Null hypothesis (Ho:) is implied by every hypothesis;
the expectation of no relationship between two concepts. Statistical significance
relies on the rejection of the null—if we can reject the null, we accept the
alternate hypothesis—if our p-value is less than .05, we can reject the null.
Alternative hypothesis (Ha:) is the claim for which we are trying to find
evidence—this is the one we really care about. It is used to represent the idea
that something has changed; a change in X leads to an effect on Y.
Ex; yellow pad
answers
How do we calculate predicted values of Y if we know the values of the IVs, alpha,
and betas? - ✔✔Use regression to predict values of Y; regression gives you
coefficients, which estimate how much of an effect X has on Y.
To predict values of Y using IV's, alphas, and betas, we use the systematic (not
stochastic) components and equation. YHat= alpha + beta (Xi); we leave out the
random component (ui), because it defines the variation in Y that is NOT related
to the variation in x.
Alpha - ✔✔called the constant/intercept. Measures the value where the
regression line crosses the y-axis.
Beta - ✔✔called the coefficient/slope. Measures the steepness of the regression
line. A one unit change in X results in a beta change in (the expected value) of Y.
R-squared - ✔✔ranges from 0-1, measures the 'quality of the fit.' Measures the
proportion of the variance (Y) that the model explains. Ex; an r-squared of 1.0
would mean the model fits the data perfectly, with a line going right through
every data point. If you get an r-squared of around .85, you'd conclude that 85%
of the data is explained by pure variation, the other 15% is due to 'luck'.
RMSE - ✔✔Root mean squared error. *The difference between each observed
value of the DV and the predicted value of the DV. Indicates the absolute fit of the
model to the data points—how close the observed data points are to the model's
predicted values. Good measure of how accurately the model physically predicts
the results; measure of the typical deviations from the regression line. Used to
compare models. The lower the values of RMSE, the better the fit.
, Standard errors - ✔✔The overall measure of uncertainty of our data/model.
Smaller standard errors of alpha and beta mean more confidence in the estimated
values.
What is the difference between a sample and a population and why (and to what
extent) do we care about each? - ✔✔Sample—a subset of cases drawn from an
underlying population.
Population—data that is considered to be relevant for every possible case.
These two are important because in order to have a successful, non-
skewed/altered OLS, the population we draw from must be general enough to
represent the case, and the sample must be random.
What is a p-value? - ✔✔The probability that we would see the relationship that
we are finding because of purely random chance. Ranges from 0-1. It's based on
the assumption that we are drawing a perfectly random sample from the
population. *The lower the p-value, the greater confidence we have that there IS
a systematic relationship between the two variables. *does NOT tell us whether
the relationship is causal. **if our p-value is <.05, it is statistically significant, and
we'd reject the null.
What is a null hypothesis and how does it relate to the p-value? What is an
alternative hypothesis? - ✔✔Null hypothesis (Ho:) is implied by every hypothesis;
the expectation of no relationship between two concepts. Statistical significance
relies on the rejection of the null—if we can reject the null, we accept the
alternate hypothesis—if our p-value is less than .05, we can reject the null.
Alternative hypothesis (Ha:) is the claim for which we are trying to find
evidence—this is the one we really care about. It is used to represent the idea
that something has changed; a change in X leads to an effect on Y.
Ex; yellow pad