Summary Machine Learning for
Business
Prof: Lien Michiels
Gabriel Gonzalez-Lopez
2025 - 2026
1
, CHP1 - Intro to Machine Learning 4
The Landscape of Data Disciplines 4
Data Types 4
Modeling Philosophies: Explanatory vs. Predictive 5
Core Machine Learning Concepts & Notation 5
Learning Paradigms 5
CHP2 - CRISP-DM Process 7
CRISP-DM Process 7
Classi cation 10
CHP3 - Classi cation via Decision Trees 11
Anatomy of a Decision Tree 11
The Algorithm: Recursive Partitioning 11
CHP4 - Logistic Regression and Generalization 14
Logistic Regression 14
Over tting & Generalization 16
Solutions & Tuning 17
CHP5 - Evaluating Classi cation Models 18
Evaluating Classi cation Models 18
Visualizing Model Performance 20
CHP6 - Naive Bayes 21
Key Insights 21
Support Vector Machines (SVM) 23
Random Forests 24
CHP7 - Similarity, Neighbors & Clustering 24
Core Concept: Similarity 24
Measuring Similarity and Distance 25
k-Nearest Neighbors (k-NN) 25
Comparison: k-Means vs. Hierarchical Clustering 27
CHP8 - Text Mining and Association Rule Mining 27
Text Mining 27
Association Rule Mining 29
CHP9 - Introduction Recommender Systems 30
Recommender Systems 30
Recommender Systems as a Machine Learning Problem 30
Recommendation Algorithms 32
Deployment and Architecture 34
2
fi fi fi fi fi
,CHP10 - Neural Networks & Deep Learning 34
Key Concepts of Neural Networks 34
Pros and Cons of Neural Networks 36
Deep Learning Architectures 37
CHP11 - Ensembling 40
Key Concepts & De nition 40
Practical Information & Exam Preparation 43
I. Exam Format 43
II. Exam Content 43
III. How to Study (Bloom’s Taxonomy) 43
IV. Example Exam Questions 43
3
fi
, CHP1 - Intro to Machine Learning
The Landscape of Data Disciplines
We begins by distinguishing between several overlapping terms often used in the
industry to ensure clarity on what Machine Learning actually is compared to its siblings.
• Data Science: This is described as the broadest term. It is the process of extracting
insights and knowledge from data using a combination of statistics, computer
science, and domain expertise.
• Data Analytics: This is a more "manual" process focused on examining, cleaning,
and modeling data to support decision-making. It is typically used by business
analysts or economists.
• Data Mining: The oldest term in the group, referring to the discovery of patterns
or anomalies in large datasets using statistical methods.
• Machine Learning (ML): The core focus of this course. It is defined as the process
of building algorithms that can learn from data to make predictions or decisions
without being explicitly programmed.
• Arti cial Intelligence (AI): The "art" of creating systems that perform tasks
requiring human intelligence (reasoning, perception). The lecture also notes
Arti cial General Intelligence (AGI) as the theoretical quest for AI that can
perform any intellectual task a human can.
• Infrastructure & Governance:
◦ Big Data: Large, complex datasets requiring advanced infrastructure
(though the term is becoming outdated).
◦ Data Engineering: Designing systems to collect, store, and process data.
This is cited as a primary reason for project success or failure.
◦ Machine Learning Engineering: The practice of deploying and maintaining
ML models in production environments.
◦ Data Governance: Policies ensuring data is accurate, secure, and used
responsibly.
Data Types
Data is categorized into three specific types based on its organization:
• Structured Data: Data that fits neatly into rows and columns (like a relational
database or Excel sheet). It consists of numbers, dates, and strings and is easy to
retrieve and analyze.
• Unstructured Data: Data primarily meant for humans, such as text, images, video,
and audio. This makes up approximately 80% of the world's data and requires
significant preprocessing before it can be used in a classifier.
• Semi-structured Data: A middle ground where data has some structure but is not
rigid (e.g., HTML, XML, JSON).
4
fi
Business
Prof: Lien Michiels
Gabriel Gonzalez-Lopez
2025 - 2026
1
, CHP1 - Intro to Machine Learning 4
The Landscape of Data Disciplines 4
Data Types 4
Modeling Philosophies: Explanatory vs. Predictive 5
Core Machine Learning Concepts & Notation 5
Learning Paradigms 5
CHP2 - CRISP-DM Process 7
CRISP-DM Process 7
Classi cation 10
CHP3 - Classi cation via Decision Trees 11
Anatomy of a Decision Tree 11
The Algorithm: Recursive Partitioning 11
CHP4 - Logistic Regression and Generalization 14
Logistic Regression 14
Over tting & Generalization 16
Solutions & Tuning 17
CHP5 - Evaluating Classi cation Models 18
Evaluating Classi cation Models 18
Visualizing Model Performance 20
CHP6 - Naive Bayes 21
Key Insights 21
Support Vector Machines (SVM) 23
Random Forests 24
CHP7 - Similarity, Neighbors & Clustering 24
Core Concept: Similarity 24
Measuring Similarity and Distance 25
k-Nearest Neighbors (k-NN) 25
Comparison: k-Means vs. Hierarchical Clustering 27
CHP8 - Text Mining and Association Rule Mining 27
Text Mining 27
Association Rule Mining 29
CHP9 - Introduction Recommender Systems 30
Recommender Systems 30
Recommender Systems as a Machine Learning Problem 30
Recommendation Algorithms 32
Deployment and Architecture 34
2
fi fi fi fi fi
,CHP10 - Neural Networks & Deep Learning 34
Key Concepts of Neural Networks 34
Pros and Cons of Neural Networks 36
Deep Learning Architectures 37
CHP11 - Ensembling 40
Key Concepts & De nition 40
Practical Information & Exam Preparation 43
I. Exam Format 43
II. Exam Content 43
III. How to Study (Bloom’s Taxonomy) 43
IV. Example Exam Questions 43
3
fi
, CHP1 - Intro to Machine Learning
The Landscape of Data Disciplines
We begins by distinguishing between several overlapping terms often used in the
industry to ensure clarity on what Machine Learning actually is compared to its siblings.
• Data Science: This is described as the broadest term. It is the process of extracting
insights and knowledge from data using a combination of statistics, computer
science, and domain expertise.
• Data Analytics: This is a more "manual" process focused on examining, cleaning,
and modeling data to support decision-making. It is typically used by business
analysts or economists.
• Data Mining: The oldest term in the group, referring to the discovery of patterns
or anomalies in large datasets using statistical methods.
• Machine Learning (ML): The core focus of this course. It is defined as the process
of building algorithms that can learn from data to make predictions or decisions
without being explicitly programmed.
• Arti cial Intelligence (AI): The "art" of creating systems that perform tasks
requiring human intelligence (reasoning, perception). The lecture also notes
Arti cial General Intelligence (AGI) as the theoretical quest for AI that can
perform any intellectual task a human can.
• Infrastructure & Governance:
◦ Big Data: Large, complex datasets requiring advanced infrastructure
(though the term is becoming outdated).
◦ Data Engineering: Designing systems to collect, store, and process data.
This is cited as a primary reason for project success or failure.
◦ Machine Learning Engineering: The practice of deploying and maintaining
ML models in production environments.
◦ Data Governance: Policies ensuring data is accurate, secure, and used
responsibly.
Data Types
Data is categorized into three specific types based on its organization:
• Structured Data: Data that fits neatly into rows and columns (like a relational
database or Excel sheet). It consists of numbers, dates, and strings and is easy to
retrieve and analyze.
• Unstructured Data: Data primarily meant for humans, such as text, images, video,
and audio. This makes up approximately 80% of the world's data and requires
significant preprocessing before it can be used in a classifier.
• Semi-structured Data: A middle ground where data has some structure but is not
rigid (e.g., HTML, XML, JSON).
4
fi