Written by students who passed Immediately available after payment Read online or as PDF Wrong document? Swap it for free 4.6 TrustPilot
logo-home
Document preview thumbnail
Preview 4 out of 34 pages
Exam (elaborations)

QMB3302 Final Exam, QMB3302 Final UF UPDATED Study Guide QUESTIONS AND CORRECT ANSWERS

Document preview thumbnail
Preview 4 out of 34 pages

QMB3302 Final Exam, QMB3302 Final UF UPDATED Study Guide QUESTIONS AND CORRECT ANSWERS

Content preview

QMB3302 Final Exam, QMB3302 Final UF
UPDATED Study Guide QUESTIONS
AND CORRECT ANSWERS
- Pipelines are useful (in analytics with Python sense) for the
following reasons? (choose all that apply)
- Pipelines make it easy to repeat/replicate steps and run multiple
models
- Pipelines are good for moving data into your programming
environment
- Pipelines automatically update to new versions of Python
- Pipelines help organize code you used to clean and treat your data
- Pipelines make it very easy to change small things in your model,
like which variable to include - CORRECT ANSWERS -
Pipelines make it easy to repeat/replicate steps and run multiple
models
- Pipelines help organize code you used to clean and treat your data
- Pipelines make it very easy to change small things in your model,
like which variable to include


The basic idea of a regression is very simple. We have some X value
(which we call ______) and some Y value that we are trying to _____.
We could have multiple Y value, but that is not something we have
covered. - CORRECT ANSWERS features; predict


Y and y-hat are a little different. Y is our target vector, and y-hat is an
output in our model that is a.... (choose one of the following)
- estimate or prediction of y

,- the actual value of y
- an axis on our 2 way graph
- a combination of XY intercept coordinates - CORRECT ANSWERS
estimate or prediction of y


When looking at the code in the videos, we sometimes used a variable
to hold out model. What is the significance of the word "model" in the
below code?


model = LinearRegression(fit_intercept=True) - CORRECT
ANSWERS 'model' is a named variable and is just holding our
linear regression model. It could be renamed anything. The word itself
is not important. It is just a container.


What is a good model fit value? - CORRECT ANSWERS
unknowable without knowing/understanding the context of the
domain


Imagine X in the below is a missing value. If I were to run a median
imputer on this set of data, what would the return value be?


50, 60, 70, 80, 100, 60, 5000, X - CORRECT ANSWERS 70


Which of the below were discussed as being problems with the
holdout method for validation?
- Data is not available for test and control differences
- Outliers can skew the result

,- The model is not trained on all of the data
- K=3 is not sufficiently large enough
- Validation is sometimes too challenging - CORRECT ANSWERS
- Outliers can skew the result
- The model is not trained on all of the data


The features in a model...
- are used as proxies for y-hat divided by y
- are always functions of each other
- keep the model validation process stable
- none of these answers are correct - CORRECT ANSWERS
none of these answers are correct


What is the first variable in a decision tree called (before any of the
branches)? - CORRECT ANSWERS root


One problem with decision trees is that they are prone to _____ if you
are not careful or do not set the _____ appropriately. - CORRECT
ANSWERS overfitting; max depth


True or False: The random forest algorithm prevents, or at least
avoids to some extent, the problems with overfitting found in decision
trees. - CORRECT ANSWERS True


True or False: Random Forests can only be used on classification
problems - CORRECT ANSWERS False

, True or False: In order to interpret Decision Tree's, it is necessary to
first run a linear regression - CORRECT ANSWERS False


True or False: Decision Tree's are nice because they are fairly simple
and straightforward to interpret - CORRECT ANSWERS True


When running our first decision tree, we took out "maxdepth=". This
had the unfortunate result of... - CORRECT ANSWERS
Building a very large hard to understand tree


What is the terminal node as discussed in the lecture? - CORRECT
ANSWERS The last node (sometimes called a leaf is you
google the term); the tree doesn't split after this


Models, such as the random forest model we ran, often have a number
of parameters that the analyst can choose or set.
What is a the best source of up to date information about the different
parameters that can be set? - CORRECT ANSWERS The scikit
learn documentation


Random forests are _____ interpretable than decision trees. -
CORRECT ANSWERS less


True or False: The correct number of clusters in Hierarchical
clustering can be determined precisely using approaches such as
silhouette scores. - CORRECT ANSWERS False

Document information

Uploaded on
January 15, 2026
Number of pages
34
Written in
2025/2026
Type
Exam (elaborations)
Contains
Questions & answers
$18.39

Wrong document? Swap it for free Within 14 days of purchase and before downloading, you can choose a different document. You can simply spend the amount again.
Written by students who passed
Immediately available after payment
Read online or as PDF

Sold
1
Followers
0
Items
1928
Last sold
4 months ago



Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

Student with book image

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Working on your references?

Create accurate citations in APA, MLA and Harvard with our free citation generator.

Working on your references?

Frequently asked questions