← SHANNONWAREARTICLE · 2026.09.28

Learning in Public: ML Zoomcamp - Module 1

As the DataTalks.Club AI Dev Tools Zoomcamp winds down, their ML Zoomcamp has begun. Today marks the end of Module 1: A whistle stop tour of the categories of the types of machine learning, and then a dive into the machine learning toolset: Pandas, Linear Algebra (vector and matrix multiplication), Numpy, Jupyter Notebooks, and GitHub Codespaces.

CompSci is Mightier than the Sword

Where as the AI Dev Tools Zoomcamp places us in the drivers seat to make full use of AI both as coding agent and software development advisor, the ML Zoomcamp is the directional opposite: beginning with a little bit of applied maths and and introduction to the tools that have made machine learning great.

Jupyter Notebook - ML Zoomcamp - GitHub Codespace

Introduction

Principle activity of machine learning is prediction.

The data set is passed to the model, and the model extracts patterns from the data.

Features are the information that we know about the subject.

Target is what we want to predict about the subject, what can be predicted by looking at all of the features. Also referred to as the "target variable".

Model encapsulates all of the patterns that we learn from the data. Then we use the features to predict the target.

Input data is of two types: Features and target. The output is the model.

ML vs Rule-Based Systems

It is a good idea to start with a rules-based system, and then use these rules as features for a machine learning system.

Supervised Machine Learning

Feature Matrix (capital X): A two dimensional array where rows are observations and columns are features.

Vector (lower case y): For each row of the matrix we have this target variable, which is practically a one dimensional array.

g(X) ≈ y

Lower case g is our model that take the matrix X and produces something that is approximately close to our target y.

The process of looking at the features and coming up with g is called training.

The Different Types of Supervised Machine Learning

Regression: This was the car price prediction example, that outputs a number from zero to positive infinity.

Classification: The output is a category. The spam example produces categories, spam or not spam.

Multiclass classification: When we want to classify something into multiple different categories.

Binary classification: There are two classification categories. This is the spam example as well.

Ranking: For example on an e-commerce website, how do we select encourage the user to select the item that they are most likely interested in? By presenting items as a ranked list.

CRISP-DM

CRISP-DM stands for Cross Industry Standard Process for Data Mining, and involves six steps:

  1. Business Understanding
  2. Data Understanding
  3. Data Preparation
  4. Modelling
  5. Evaluation
  6. Deployment

Numpy

It might be said that Numpy is to machine learning as cement is to concrete. Numpy is a library of Python functions that allows for very fast, very efficient mathematical operations on arrays of data, especially matrices.

Linear Algebra

The minimum linear algebra required to begin in machine learning are:

  1. How to multiply two vectors
  2. How to multiply a vector and a matrix
  3. How to multiply two matrices
  4. Cross product, dot product, etc.

Pandas

Built on top of Numpy, Pandas is a library of Python functions that is optimised for performing operations on tables of data, especially the columns of tables. Imagine adding, multiplying, dividing, finding the mean or mode of all of one particular column in a very large table, for starters.

ImpressumContactTermsPrivacy