BUSINESS S TATISTICS
4TH EDITION
NOREAN R. SHARPE
RICHARD D. DE VEAUX
PAUL F. VELLEMAN
, TABLE OF CONTENTS
Part I Exploring and Collecting Data Case Study: Health Care Costs 20-33
Chapter 1 Data and Decisions 1-1 Part VI Analytics
Chapter 2 Displaying and Describing Categorical Data 2-1 Chapter 21 Introduction to Big Data and Data Mining 21-1
Chapter 3 Displaying and Describing Quantitative Data 3-1 Part VII Online Topics
Chapter 4 Correlation and Linear Regression 4-1 Chapter 22 Quality Control 22-1
Case Study: Paralyzed Veterans of America 4-49 Chapter 23 Nonparametric Methods 23-1
Part II Modeling with Probability Chapter 24 Decision Making and Risk 24-1
Chapter 5 Randomness and Probability 5-1 Chapter 25 Analysis of Experiments and Observational Studies
25-1
Chapter 6 Random Variables and Probability Models 6-1
Chapter 7 The Normal and Other Continuous Distributions 7-1
Part III Gathering Data
Chapter 8 Data Sources: Observational Studies and Surveys 8-1
Chapter 9 Data Sources: Experiments 9-1
Part IV Inference for Decision Making
Chapter 10 Sampling Distributions and Confidence Intervals for
Proportions 10-1
Case Study: Real Estate Simulation
Chapter 11 Confidence Intervals for Means 11-1
Chapter 12 Testing Hypotheses 12-1
Chapter 13 More about Tests and Intervals 13-1
Chapter 14 Comparing Two Means 14-1
Chapter 15 Inference for Counts: Chi-Square tests 15-1
Brief Case: Loyalty Program 15-27
Part V Models for Decision Making
Chapter 16 Inference for Regression 16-1
Chapter 17 Understanding Residuals 17-1
Chapter 18 Multiple Regression 18-1
Copyright © 2019 Pearson Education, Inc.
Chapter 19 Building Multiple Regression Models 19-1
Chapter 20 Time Series Analysis 20-1
, Chapter 1 – Data and Decisions
SECTION EXERCISES
SECTION 1.1
1. a) Each row represents a different house that was recently sold. It can be described as a case.
b) There are six quantitative variables in each row plus a house identifier for a total of seven variables.
2. a) Each row represents a different transaction (not customer or book). It can be described as a case.
b) There are six quantitative variables plus two identifiers in each row for a total of eight variables.
SECTION 1.2
3. a) House_ID is an identifier (categorical, not ordinal); Neighborhood is categorical (nominal); Mail_ZIP is
categorical (nominal – ordinal in a sense, but only on a national level); Acres is quantitative (units – acres);
Yr_Built is quantitative (units – year); Full_Market_Value is quantitative (units – dollars); Size is
quantitative (units – square feet).
b) These data are cross-sectional. Each row corresponds to a house that recently sold so at approximately
the same fixed point in time.
4. a) Transaction ID is an identifier (categorical, nominal, not ordinal); Customer ID is an identifier
(categorical, nominal); Date can be treated as quantitative (how many days since the transaction took place,
days since Jan. 1 2009, for example) or categorical (as month, for example); ISBN is an identifier
(categorical, nominal); Price is quantitative (units – dollars); Coupon is categorical (nominal); Gift is
categorical (nominal); Quantity is quantitative (unit – counts).
b) These data are cross-sectional. Each row corresponds to a transaction at a fixed point in time. However,
the date of the transaction has been recorded so the data could be reconfigured as a time series. It is likely
that the store had more sales in that time period so a time series is not appropriate.
SECTION 1.3
5. It is not specified whether or not the real estate data of Exercise 1 are obtained from a survey. The data
would not be from an experiment, a data gathering method with specific requirements. Rather, the real
estate major’s data set was derived from transactional data (on local home sales). The major concern with
drawing conclusions from this data set is that we cannot be sure that the sample is representative of the
population of interest (e.g., all recent local home sales or even all recent national home sales). Therefore,
we should be cautious about drawing conclusions from these data about the housing market in general.
6. The student is using a secondary data source (from the Internet). No information is given about how, when,
where and why these data were collected or if it was the result of a designed experiment. It is also not
stated that the sample is representative of companies. There are concerns about using these data for
generalizing and drawing conclusions because the data could have been collected for a different purpose
(not necessarily for developing a stock investment strategy). Therefore, the student should be cautious
about using this type of data to predict performance in the future.
CHAPTER EXERCISES
7. The news. Answers will vary.
8. The Internet. Answers will vary.
9. Survey. The description of the study has to be broken down into its components in order to understand the
study. Who– who or what was actually sampled–college students; What–what is being measured–opinion of
electric vehicles: whether there will more electric or gasoline powered vehicles in 2025 and the likelihood
of whether they would purchase an electric vehicle in the next 10 years; When–current; Where–your
location; Why–automobile manufacturer wants college student opinions; How–how was the study
1-1
, 1-2 Chapter 1 Data and Decisions
conducted–survey; Variables–there are two categorical variables–what students think about whether or not
there will be more electric or gasoline powered vehicles in 2025 and the second categorical variable is also
ordinal–how likely, using a scale, would the student be to buy an electric vehicle in the next 10 years;
Source –the data are not from a designed survey or experiment; Type–the data are cross-sectional;
Concerns–none.
10. Your survey. Answers will vary.
11. World databank. Answers will vary but chosen from the following possible indicators:
GDP growth (annual %)
GDP (current US$)
GDP per capita (current US$)
GNI per capita, Atlas method (current US$)
Exports of goods and services (% of GDP)
Foreign direct investment, net inflows (BoP, current US$)
GNI per capita, PPP (current international $)
GINI index
Inflation, consumer prices (annual %)
Population, total
Life expectancy at birth, total (years)
Internet users (per 100 people)
Imports of goods and services (% of GDP)
Unemployment, total (% of total labor force)
Agriculture, value added (% of GDP)
CO2 emissions (metric tons per capita)
Literacy rate, adult total (% of people ages 15 and above)
Central government debt, total (% of GDP)
Inflation, GDP deflator (annual %)
Poverty headcount ratio at national poverty line (% of population)
12. Arby’s menu. Who–Arby’s sandwiches; What–type of meat, number of calories (in calories), and serving
size (in ounces); When–not specified; Where–Arby’s restaurants; Why–assess the nutritional value of the
different sandwiches; How–information was gathered from each of the sandwiches on the menu at Arby’s,
resulting in a census; Variables–there are 3 variables: the number of calories and serving size are
quantitative, and the type of meat is categorical; Source–data are not from a designed survey or experiment;
Type–data are cross-sectional; Concerns–none.
13. MBA admissions. Who–MBA applicants (in northeastern U.S.); What–sex, age, whether or not accepted,
whether or not they attended, and the reasons for not attending (if they did not accept); When–not specified;
Where–a school in the northeastern United States; Why–the researchers wanted to investigate any patterns
in female student acceptance and attendance in the MBA program; How–data obtained from the admissions
office; Variables–there are 5 variables: sex, whether or not the students accepted, whether or not they
attended, and the reasons for not attending if they did not accept (all categorical) and age which is
quantitative; Source–data are not from a designed survey or experiment; Type–data are cross-sectional;
Concerns–none.
14. MBA admissions II. Who–MBA students (in program outside of Paris); What–each student’s standardized
test scores and GPA in the MBA program; When–2009 to 2014; Where–outside of Paris; Why–to
investigate the association between standardized test scores and performance in the MBA program over
five years (2009–2014); How–not specified; Variables–there are 2 quantitative variables: standardized test
scores and GPA; Source–data are not from a designed survey or experiment, data are available from student
records; Type–although the data are collected over 5 years, the purpose is to examine them as cross-
sectional rather than as time-series; Concerns–none.