WGU C207 DATA-DRIVEN DECISION STUDY GUIDE |
LATEST UPDATE
Module 1: The Case for Quantitative Analysis
Analytics
Analytics – extensive use of data, statistical and quantitative analysis, explanatory and predictive
models, and fact-based management to drive decisions and add value.
• Descriptive statistics are used to inform/explain
• Inferential statistics are used to predict/trend
Big Data
Refers to both structured and unstructured data in such large volumes that it's difficult to process using
traditional database and software techniques.
Data Mining
Process of discovering patterns in large data sets. Data mining is performed on big data to decipher
patterns from these large databases.
Davenport-Kim 3-stage model
• Framing the Problem
o Problem Recognition
1. Identifying Stakeholders
2. Focusing on decisions
3. Identifying the kind of story
4. Determining the scope of the problem
5. Getting specific about what data to analyze
o Review of Previous Findings
, • Solving the Problem
o Modeling Step
o Data Collection Step
o Data Analysis Step
• Communicating Results
Data Management
Refers to cleaning and organizing a data set that has been collected
• Available
• Accurate
• Complete
• Relevant
• Timely
4 Levels of Measurement
• Continuous Data – a data point can lay along any point in a range of data (age)
o Interval Data – all objects are an equal interval apart, cannot have a natural zero (time)
o Ratio Data – has a unique zero point (age, Kevin scale, income, stock price, inventory)
• Discrete Data – can only take on whole values and has clear boundaries (number of cars)
o Nominal Data – called categorical data, used to label subjects in a study (males/females)
o Ordinal Data – places data objects into an order according to some quality (degrees)
Reliability and Validity of Data
• Random Error –will not repeat over time, minimized by larger sample size
• Systematic Error – it repeats itself, constant measurement error, measurement instruments
• Omission Error – when relevant data is not included in study or action has not been taken
• Outlier – observation points (numbers) that are distant from other observations
• Measurement Bias
o Sample is not representative of the population
o Sample tested is not sufficiently random
• Information Bias
o Response Bias – Respondent says what they believe the questioner wants to hear
o Conscious Bias – Surveyor is actively seeking a certain response
Skewness (Bias) – is a measure of the degree to which data leans toward one side.
2
, Research Design
• Observational Studies – when it’s impractical or impossible to control the conditions of the study
o Cohort Study
o Case Control Study
• Experimental Studies – variable measurements and subjects are under the researcher’s control
o Experimental units – subjects of objects under observation
o Treatments – the procedures applied to each subject
o Responses – the effects of the experimental treatments
Experimental Studies: Explanatory Variable
Also known as the independent or predictor variable, it explains variations in the response variable; in
an experimental study, it is manipulated by the researcher
Experimental Studies: Response Variable
Also known as the dependent or outcome variable, its value is predicted, or its variation is explained by
the explanatory variable; in an experimental study, this is the outcome of study.
• Blind Study – participants are not told
• Double Blind Study – data gatherers and participants are not told
• Triple Blind Study – data analyzer, data gatherers, and participants are not told
Experimental Design
• Qualitative Research – exploratory research, data not characterized by numbers
• Quantitative Research – uses numerical data and measurements
Module 2: Statistics as a Managerial Tool
The Misuse of Statistics
• Not a truly representative sample
• Response bias
• Conscious bias
• Missing data and refusals
• Small sample sizes
• Association and causality
• Training and test data
• Unfounded assumptions
• Faulty operationalization
• Lack of blinding
3
LATEST UPDATE
Module 1: The Case for Quantitative Analysis
Analytics
Analytics – extensive use of data, statistical and quantitative analysis, explanatory and predictive
models, and fact-based management to drive decisions and add value.
• Descriptive statistics are used to inform/explain
• Inferential statistics are used to predict/trend
Big Data
Refers to both structured and unstructured data in such large volumes that it's difficult to process using
traditional database and software techniques.
Data Mining
Process of discovering patterns in large data sets. Data mining is performed on big data to decipher
patterns from these large databases.
Davenport-Kim 3-stage model
• Framing the Problem
o Problem Recognition
1. Identifying Stakeholders
2. Focusing on decisions
3. Identifying the kind of story
4. Determining the scope of the problem
5. Getting specific about what data to analyze
o Review of Previous Findings
, • Solving the Problem
o Modeling Step
o Data Collection Step
o Data Analysis Step
• Communicating Results
Data Management
Refers to cleaning and organizing a data set that has been collected
• Available
• Accurate
• Complete
• Relevant
• Timely
4 Levels of Measurement
• Continuous Data – a data point can lay along any point in a range of data (age)
o Interval Data – all objects are an equal interval apart, cannot have a natural zero (time)
o Ratio Data – has a unique zero point (age, Kevin scale, income, stock price, inventory)
• Discrete Data – can only take on whole values and has clear boundaries (number of cars)
o Nominal Data – called categorical data, used to label subjects in a study (males/females)
o Ordinal Data – places data objects into an order according to some quality (degrees)
Reliability and Validity of Data
• Random Error –will not repeat over time, minimized by larger sample size
• Systematic Error – it repeats itself, constant measurement error, measurement instruments
• Omission Error – when relevant data is not included in study or action has not been taken
• Outlier – observation points (numbers) that are distant from other observations
• Measurement Bias
o Sample is not representative of the population
o Sample tested is not sufficiently random
• Information Bias
o Response Bias – Respondent says what they believe the questioner wants to hear
o Conscious Bias – Surveyor is actively seeking a certain response
Skewness (Bias) – is a measure of the degree to which data leans toward one side.
2
, Research Design
• Observational Studies – when it’s impractical or impossible to control the conditions of the study
o Cohort Study
o Case Control Study
• Experimental Studies – variable measurements and subjects are under the researcher’s control
o Experimental units – subjects of objects under observation
o Treatments – the procedures applied to each subject
o Responses – the effects of the experimental treatments
Experimental Studies: Explanatory Variable
Also known as the independent or predictor variable, it explains variations in the response variable; in
an experimental study, it is manipulated by the researcher
Experimental Studies: Response Variable
Also known as the dependent or outcome variable, its value is predicted, or its variation is explained by
the explanatory variable; in an experimental study, this is the outcome of study.
• Blind Study – participants are not told
• Double Blind Study – data gatherers and participants are not told
• Triple Blind Study – data analyzer, data gatherers, and participants are not told
Experimental Design
• Qualitative Research – exploratory research, data not characterized by numbers
• Quantitative Research – uses numerical data and measurements
Module 2: Statistics as a Managerial Tool
The Misuse of Statistics
• Not a truly representative sample
• Response bias
• Conscious bias
• Missing data and refusals
• Small sample sizes
• Association and causality
• Training and test data
• Unfounded assumptions
• Faulty operationalization
• Lack of blinding
3