Complete Revision & Study Guide
A topic-by-topic guide covering big data foundations, statistics, data cleaning, high performance
computing, and the core machine learning methods used in data analytics at scale.
Original study notes — independently written summary and explanation
,Contents
• 1. Foundations: Big Data & Data Analytics
• 2. Analytics Process Frameworks
• 3. Statistics Refresher
• 4. Data Cleaning & Preparation
• 5. Tidy Data
• 6. High Performance Computing Fundamentals
• 7. Tools for Scalable Analytics: Spark, Dask & PyTorch
• 8. Supervised Learning: Linear & Logistic Regression
• 9. K-Nearest Neighbours Classification
• 10. Naive Bayes Classification
• 11. Tree-Based Models & Random Forests
• 12. Neural Networks
• 13. Clustering: K-Means, Gaussian Mixtures & DBSCAN
,1. Foundations: Big Data & Data Analytics
1.1 Data, Data Analytics, and Data Science
Data are pieces of information gathered through observation. More formally, data are a set of
qualitative and quantitative values describing one or more people or objects.
Data Analytics is a broad field covering business intelligence and application-driven analysis of data.
Data Science, by contrast, is the study of the computational methods, principles, and systems used
to extract knowledge from data — it is more concerned with the underlying methodology than with
any single application.
1.2 The "Danger Zone"
A recurring theme in this field is what can be called the danger zone: people who can produce
output that looks like legitimate analysis, without actually understanding how that output was
produced or what it means. This typically happens when someone has strong computing skills and
domain knowledge, but lacks a real grounding in data analysis principles. The risk is confident,
plausible-looking conclusions that are quietly wrong.
1.3 Big Data and the Five Vs
Big Data refers to datasets that are too large, fast-moving, or unstructured to be handled by
conventional tools. It is commonly broken down into five defining characteristics, the "five Vs":
Characteristic What it means
Volume The sheer size of the dataset — too large for traditional
infrastructure to store or process.
Velocity The speed at which data arrives and needs to be processed.
Variety The mixture of structured and unstructured data formats.
Veracity The quality, trustworthiness, and bias present in the data.
Value Whether the effort of processing the data is actually worth the
insight gained.
Volume
Challenge: datasets are too large to be processed on conventional IT infrastructure.
Solutions: scalable storage and databases, and distributed queries that process data in parallel.
Velocity
Challenges: the rate at which data is generated can outstrip the system's ability to store it, and some
decisions need to be made in near real time.
Solutions: storing only a (possibly reduced or modified) subset of the incoming data, and using
streaming techniques to support online decision-making.
Variety
Structured data is arranged in a tabular format with a fixed schema, and is efficient to process with
traditional tools. Unstructured data covers everything else — free text, geospatial data, audio, video
— and is typically stored in its native format.
Challenge: structured data is often not immediately analysis-ready, and the volume of unstructured
data is growing rapidly.
, Solutions: general data cleaning; for unstructured data specifically, transforming it into structured
form where possible, storing it using newer technologies such as NoSQL databases, and processing it
with analytical techniques built for unstructured content.
Veracity
Veracity is essentially a measure of data quality: the degree to which data are accurate, unbiased,
and trustworthy.
Challenges: data are often collected for one purpose but later reused for a different one it wasn't
designed for, and quality can vary unevenly across a dataset.
Solutions: domain expertise is essential for spotting quality issues, alongside systematic data
cleaning and data governance policies that standardise how data is collected and processed.
Value
Challenge: understanding how a business can actually extract value from its Big Data, rather than
simply accumulating it.
Solutions: a cost/benefit analysis can help establish whether processing the data is worthwhile —
this becomes more favourable as the cost of storage continues to fall.
1.4 High Performance Computing (HPC): a first look
HPC is the use of specialised computing systems to solve complex problems at scale. The same
factors that define Big Data — volume, velocity, and variety — are what typically force analysts
beyond the limits of a standard desktop machine.
When more CPU power, memory, or storage is required than a conventional setup can provide, the
task needs "scaling" — expanding the computational resources available to it. There are two broad
approaches:
• Vertical scaling ("scaling up"): increasing the resources available on a single machine — more
CPUs, more memory, or a higher-spec virtual machine. This is usually the easier route, since it
typically requires few or no code changes.
• Horizontal scaling ("scaling out"): increasing the number of machines (nodes) working on the
problem, for example by adding more workers to a Spark cluster. This is generally harder, since
computation now has to be coordinated across multiple machines. Frameworks such as Spark
and Dask exist specifically to make this more manageable.