100% satisfaction guarantee Immediately available after payment Both online and in PDF No strings attached 4.2 TrustPilot
logo-home
Summary

JADS Master - Data Mining Summary

Rating
-
Sold
1
Pages
39
Uploaded on
02-01-2022
Written in
2020/2021

Summary for the Data Mining course of the Master Data Science and Entrepreneurship.

Institution
Course









Whoops! We can’t load your doc right now. Try again or contact support.

Written for

Institution
Study
Course

Document information

Uploaded on
January 2, 2022
Number of pages
39
Written in
2020/2021
Type
Summary

Subjects

Content preview

1. Introduction
Machine Learning
Learn to perform a task based on experience 𝑋 and minimizing error ϵ
● If the data is biased the model is biased
𝑓θ(𝑋) = 𝑦
Inductive Bias
The assumptions put into the model (β):
● What should the model look like
● User-defined settings (hyperparameters)
● Assumptions about the distribution of the data (i.e. 𝑋 ∼ 𝑁)
● Knowledge transferred from previous tasks (𝑓 , 𝑓 , 𝑓 , ... ⇒ 𝑓 )
1 2 3 𝑛𝑒𝑤
𝑎𝑟𝑔 min ϵ(𝑓θ,β(𝑋))
θ,β

Statistics Machine Learning
● Help humans understand the world ● Automated task entry
● Assume data generated according ● Assume data generation process is
understandable model unknown

Supervised Learning
Learn model 𝑓 from labeled data (𝑋, 𝑦) (ground truth).
● Classification: predict a class label (category), discrete and unordered
○ Result can be binary (0, 1) or multi-class (a, b, c, d)
○ Can return confidence per class
○ Predictions yield a decision boundary separating classes
● Regression: predict a continuous value (i.e. temperature)
○ Target variable is numeric
○ Some algorithms can return confidence interval
○ Find the relationship between predictors and the target variable




Unsupervised Learning
Explore structure of unlabeled data (𝑋) to extract meaningful information.
● Clustering: organize information into meaningful subgroups (clusters)
○ Objects in the cluster share a certain degree of similarity (and dissimilarity to
other clusters)


1

, ● Dimensionality reduction: can compress data into fewer dimensions while retaining most
of the information
○ New features lose original meaning
○ New representation can be easier to model/visualize

Semi-supervised Learning
Learn a model from a few labeled and many unlabeled data points.

Reinforcement Learning
Develop an agent that improves performance based on interactions with the environment.
● Search a (large) space of actions and states
● Learn a series of actions (policy) that maximizes reward through exploration
● Reward function: defines how well a (series of) actions works

Learning = Representation + Evaluation + Optimization
● Representation: defines concepts it can learn (hypothesis space)
● Evaluation: an way to choose one hypothesis over the other using a object function,
scoring function or loss function (ℓ) (diff. between correct output and predictions)
● Optimization: efficient way to search hypothesis space

Overfitting
A model that is too complex for the amount of data you have (high train score & low test score)
→ Solution: make the model simpler (regularization), collect more data, remove features or
scale data.

Underfitting
A model that is too simple given the complexity of the data (low train score & low test score) →
Solution: use more complex model

Model Selection
By using an (external) evaluation function we can check:
● If we’re learning the right thing (feedback signal) → underfitting/overfitting
● Choose to fit the application
● Choose different hyperparameter settings

Data Split
Data needs to be split to avoid data leakage (optimizing hyperparameters or preprocessing
based on the test data).
● Train model → train set
● Optimize hyperparameters → validation set
● Evaluate → test set




2

Get to know the seller

Seller avatar
Reputation scores are based on the amount of documents a seller has sold for a fee and the reviews they have received for those documents. There are three levels: Bronze, Silver and Gold. The better the reputation, the more your can rely on the quality of the sellers work.
tomdewildt Jheronimus Academy of Data Science
Follow You need to be logged in order to follow users or courses
Sold
29
Member since
4 year
Number of followers
13
Documents
22
Last sold
6 months ago

5.0

1 reviews

5
1
4
0
3
0
2
0
1
0

Recently viewed by you

Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

Student with book image

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Frequently asked questions