Summary

JADS Master - Data Mining Summary

Rating

Sold

Pages

Uploaded on

02-01-2022

Written in

2020/2021

Summary for the Data Mining course of the Master Data Science and Entrepreneurship.

Institution

Course

Whoops! We can’t load your doc right now. Try again or contact support.

Report Copyright Violation

Written for

Institution: Tilburg University (UVT)
Study: Data Science And Entrepreneurship
Course: Data Mining (JM0150M6)

All documents for this subject (1)

Document information

Uploaded on: January 2, 2022
Number of pages: 39
Written in: 2020/2021
Type: Summary

Subjects

jads
master
data mining
linear models
kernelization
data preprocessing
neural networks
neural networks for text
summary
evaluation amp model specification
convolutional neural networks

Content preview

1. Introduction
Machine Learning
Learn to perform a task based on experience 𝑋 and minimizing error ϵ
● If the data is biased the model is biased
𝑓θ(𝑋) = 𝑦
Inductive Bias
The assumptions put into the model (β):
● What should the model look like
● User-defined settings (hyperparameters)
● Assumptions about the distribution of the data (i.e. 𝑋 ∼ 𝑁)
● Knowledge transferred from previous tasks (𝑓 , 𝑓 , 𝑓 , ... ⇒ 𝑓 )
1 2 3 𝑛𝑒𝑤
𝑎𝑟𝑔 min ϵ(𝑓θ,β(𝑋))
θ,β

Statistics Machine Learning
● Help humans understand the world ● Automated task entry
● Assume data generated according ● Assume data generation process is
understandable model unknown

Supervised Learning
Learn model 𝑓 from labeled data (𝑋, 𝑦) (ground truth).
● Classification: predict a class label (category), discrete and unordered
○ Result can be binary (0, 1) or multi-class (a, b, c, d)
○ Can return confidence per class
○ Predictions yield a decision boundary separating classes
● Regression: predict a continuous value (i.e. temperature)
○ Target variable is numeric
○ Some algorithms can return confidence interval
○ Find the relationship between predictors and the target variable

Unsupervised Learning
Explore structure of unlabeled data (𝑋) to extract meaningful information.
● Clustering: organize information into meaningful subgroups (clusters)
○ Objects in the cluster share a certain degree of similarity (and dissimilarity to
other clusters)

1

, ● Dimensionality reduction: can compress data into fewer dimensions while retaining most
of the information
○ New features lose original meaning
○ New representation can be easier to model/visualize

Semi-supervised Learning
Learn a model from a few labeled and many unlabeled data points.

Reinforcement Learning
Develop an agent that improves performance based on interactions with the environment.
● Search a (large) space of actions and states
● Learn a series of actions (policy) that maximizes reward through exploration
● Reward function: defines how well a (series of) actions works

Learning = Representation + Evaluation + Optimization
● Representation: defines concepts it can learn (hypothesis space)
● Evaluation: an way to choose one hypothesis over the other using a object function,
scoring function or loss function (ℓ) (diff. between correct output and predictions)
● Optimization: efficient way to search hypothesis space

Overfitting
A model that is too complex for the amount of data you have (high train score & low test score)
→ Solution: make the model simpler (regularization), collect more data, remove features or
scale data.

Underfitting
A model that is too simple given the complexity of the data (low train score & low test score) →
Solution: use more complex model

Model Selection
By using an (external) evaluation function we can check:
● If we’re learning the right thing (feedback signal) → underfitting/overfitting
● Choose to fit the application
● Choose different hyperparameter settings

Data Split
Data needs to be split to avoid data leakage (optimizing hyperparameters or preprocessing
based on the test data).
● Train model → train set
● Optimize hyperparameters → validation set
● Evaluate → test set

2

$6.66

Get access to the full document:

100% satisfaction guarantee

Immediately available after payment

Both online and in PDF

No strings attached

Get to know the seller

tomdewildt

5.0

(1)

Get to know the seller

tomdewildt Jheronimus Academy of Data Science

View profile

Sold

Member since

4 year

Number of followers

Documents

Last sold

6 months ago

5.0

1 reviews

Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their tests and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can instantly pick a different document that better fits what you're looking for.

Pay as you like, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card and download your PDF document instantly.

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Frequently asked questions

What do I get when I buy this document?

You get a PDF, available immediately after your purchase. The purchased document is accessible anytime, anywhere and indefinitely through your profile.

Satisfaction guarantee: how does it work?

Our satisfaction guarantee ensures that you always find a study document that suits you well. You fill out a form, and our customer service team takes care of the rest.

Who am I buying these notes from?

Stuvia is a marketplace, so you are not buying this document from us, but from seller tomdewildt. Stuvia facilitates payment to the seller.

Will I be stuck with a subscription?

No, you only buy these notes for $6.66. You're not tied to anything after your purchase.

Can Stuvia be trusted?

4.6 stars on Google & Trustpilot (+1000 reviews) 44104 documents were sold in the last 30 days Founded in 2010, the go-to place to buy study notes for 15 years now

JADS Master - Data Mining Summary

Written for

Document information

Subjects

Content preview

Get to know the seller

Recently viewed by you

Why students choose Stuvia

Created by fellow students, verified by reviews

Didn't get what you expected? Choose another document

Pay as you like, start learning right away

Frequently asked questions

What do I get when I buy this document?

Satisfaction guarantee: how does it work?

Who am I buying these notes from?

Will I be stuck with a subscription?

Can Stuvia be trusted?