Summary

Data Mining 2017/2018 - Short Summary

Rating

Sold

Pages

Uploaded on

10-01-2018

Written in

2017/2018

Short summary (samenvatting) Data Mining Data Science Regression Classification Clustering Dimensionality Reduction

Institution

Course

Whoops! We can’t load your doc right now. Try again or contact support.

Report Copyright Violation

Written for

Institution: Tilburg University (UVT)
Study: Data Science
Course: Data Mining

All documents for this subject (2)

Document information

Uploaded on: January 10, 2018
Number of pages: 4
Written in: 2017/2018
Type: Summary

Subjects

data
mining
summary

Content preview

Data Mining Essentials
Supervised vs Unsupervised Learning
- Supervised learning
o Classification (cat | dog | mouse)
o Regression (24 | 3 | 32 | 10)
- Unsupervised ‘learning’
o Clustering ( a b c | k l m | x y z)
o Dimensionality reduction (X1, X2, X3, X4, X5  –X3, –X5)

Overall goal of both methods: extract from dataset with goal to generalize.

Supervised Learning
- Training set with vectors | categorised (colours)
- Flowchart: raw data collection » pre-processing » sampling » re-processing » learning
algorithm training » hyperparameter optimisation » post-processing » final classification /
regression model

Pre-processing
Feature transformation:

- Categorical variables
o Nominal (green » [0,1,0])
o Ordinal (XL » 3)
- Normalisation and outlier removal
o Z-score (mean/SD)
o Remove outliers (depends on your goal)
- Vector normalisation
o L2-norm (√∑x²)  ○
o L1-norm (∑|x|)  ◊

Data Exploration and Visualisation (descriptive analysis)
- Sort or rearrange your data
- Goal of thesis: how well following the guidelines?

Splitting your data
- The fundamental goal is to generalize beyond the data instances used to train models
- Never touch the test data (until the end)
- Test data must belong to the same (statistical) distribution as the training data!
1. Sequential Split: for example a time series, typically train on a period, for example one 1-6
and test on 7-8. Common pitfall is cycles in the data (on different time-scales).
2. Random Split: blindly assign instances to training…….

Sampling and splitting your data
- In the case of small data, you want to check
(stratify) your data in terms of target, or at
least check if the ratios are representative.
- In the case of unbalanced data you might
want to stratify your data.

R61,08

Get access to the full document:

100% satisfaction guarantee

Immediately available after payment

Both online and in PDF

No strings attached

Get to know the seller

JHessels

2,5

(6)

Get to know the seller

JHessels Tilburg University

View profile

Sold

Member since

7 year

Number of followers

Documents

Last sold

1 year ago

2,5

6 reviews

Why students choose Stuvia

Created by fellow students, verified by reviews

Quality you can trust: written by students who passed their exams and reviewed by others who've used these notes.

Didn't get what you expected? Choose another document

No worries! You can immediately select a different document that better matches what you need.

Pay how you prefer, start learning right away

No subscription, no commitments. Pay the way you're used to via credit card or EFT and download your PDF document instantly.

“Bought, downloaded, and aced it. It really can be that simple.”

Alisha Student

Frequently asked questions

What do I get when I buy this document?

You get a PDF, available immediately after your purchase. The purchased document is accessible anytime, anywhere and indefinitely through your profile.

Satisfaction guarantee: how does it work?

Our satisfaction guarantee ensures that you always find a study document that suits you well. You fill out a form, and our customer service team takes care of the rest.

Who am I buying this summary from?

Stuvia is a marketplace, so you are not buying this document from us, but from seller JHessels. Stuvia facilitates payment to the seller.

Will I be stuck with a subscription?

No, you only buy this summary for R61,08. You're not tied to anything after your purchase.

Can Stuvia be trusted?

4.6 stars on Google & Trustpilot (+1000 reviews) 48341 documents were sold in the last 30 days Founded in 2010, the go-to place to buy summaries for 15 years now

Data Mining 2017/2018 - Short Summary

Written for

Document information

Subjects

Content preview

More courses for Tilburg University (UVT) > Data Science

Get to know the seller

Recently viewed by you

Why students choose Stuvia

Created by fellow students, verified by reviews

Didn't get what you expected? Choose another document

Pay how you prefer, start learning right away

Frequently asked questions

What do I get when I buy this document?

Satisfaction guarantee: how does it work?

Who am I buying this summary from?

Will I be stuck with a subscription?

Can Stuvia be trusted?