About Course
Welcome to the most comprehensive course on Machine Learning and Data Science!
This course is the best way to start from scratch and become a data scientist and machine learning expert using Python.
This voluminous course can replace you with a whole set of other courses that can cost tens of times more. In this course you will study the following topics:
- Programming in Python (express course)
- NumPy in Python
- Deep dive into Pandas for data analysis and preprocessing
- Detailed exploration of Seaborn for data visualization (including Matplotlib for customizing plots)
- Machine learning with SciKit Learn, including the following topics:
- Linear Regression – Linear Regression
- Regularization – Regularization
- Lasso Regression – Lasso Regression
- Ridge Regression – Ridge Regression
- Elastic Net Regularization
- Logistic Regression – Logistic regression
- K Nearest Neighbors – K-nearest neighbors method
- Decision Trees – Decision Trees
- Random Forests – Random Forests
- AdaBoost, GradientBoosting – Adaptive boosting, Gradient boosting
- Natural Language Processing – Processing of language data
- K Means Clustering
- Hierarchical Clustering – Hierarchical clustering
- DBSCAN (Density-based spatial clustering of applications with noise) – Clustering based on data density
- PCA – Principal Component Analysis – Principal component method
- And many many others!
What Will You Learn?
- Building supervised machine learning models (Supervised Learning)
- Using NumPy to work with numbers in Python
- Using Seaborn to Create Beautiful Data Visualization Graphs
- Using Pandas for Data Manipulation in Python
- Using Matplotlib to fine-tune data visualizations in Python
- Feature Engineering using Realistic Examples
- Regression algorithms for predicting continuous variables
- Skills in preparing data for machine learning
- Classification algorithms for predicting categorical variables
- Creating a portfolio of machine learning and data science projects
- Working with Scikit-Learn to apply various machine learning algorithms
- Quickly set up Anaconda for machine learning work
- Understanding the full cycle of stages of machine learning work
Course Content
Introductory part of the course
-
Welcome to the course!
-
COURSE OVERVIEW – DON’T SKIP THIS LECTURE
-
Download slides for presentations (OPTIONAL)
-
Installing Anaconda, Python, Jupyter Notebook
-
Read this article – A note on setting up your development environment
-
Setting up the development environment
-
FAQ
OPTIONAL: Python crash course
-
Optional: Python crash course
-
Python Crash Course – Part 1
-
Python Crash Course – Part 2
-
Python Crash Course – Part 3
-
Python Test Exercises
-
Solutions for Python Test Exercises
Stages of machine learning work
-
Stages of machine learning work
NumPy
-
Overview of the NumPy section
-
NumPy Arrays
-
Indexing and selecting data from NumPy arrays
-
Operations in NumPy
-
NumPy Test Exercises
-
Solutions for NumPy Test Exercises
00:00
Pandas
-
Overview of the Pandas section
-
Series – Part 1
-
Series – Part 2
-
Dataframes – Part 1 – Creating dataframes
-
Dataframes – Part 2 – Basic Attributes
-
Dataframes – Part 3 – Working with columns
-
Dataframes – Part 4 – Working with Strings
-
Conditional Filtering
-
Useful methods – Apply for one column
-
Useful Methods – Apply to Multiple Columns
-
Useful Techniques – Statistical Information and Data Sorting
-
Missing data – Overview
-
Missing data – Operations in Pandas
-
GROUP BY Data Aggregation – Part 1
-
GROUP BY Data Aggregation – Part 2 – Multi-Index
-
Combining dataframes – Concatenation
-
Merging dataframes – Inner Merge
-
Merging dataframes – Left and Right Merge
-
Merging dataframes – Outer Merge
-
Pandas Methods for Text
-
Pandas Methods for Date and Time
-
Input/Output in Pandas – CSV files
-
Input/Output in Pandas – HTML tables
-
Input/Output in Pandas – Excel Files
-
Input/Output in Pandas – SQL Database
-
Pivot tables in Pandas
-
Pandas Test Exercises
-
Solutions for Pandas Test Exercises
-
Overview of the section about Matplotlib
Matplotlib
-
Overview of the section about Matplotlib
-
Matplotlib Basics
-
Figure object – how it works
-
Figure Object – Code in Python
-
Figure object – parameters
-
Subplots – multiple plots next to each other
-
Styling Matplotlib: Legends
-
Styling Matplotlib: Colors and Styles
-
Additional materials on Matplotlib
-
Matplotlib test exercises
-
Solutions for Matplotlib Test Exercises
Seaborn
-
Scatterplots – Scatter plots (scatter plots)
-
Distribution Plots – Part 1 – Types of Plots
-
Distribution Plots – Part 2 – Code in Python
-
Categorical Plots – Categorical Statistics – Plot Types
-
Categorical Plots – Categorical Statistics – Code in Python
-
Categorical Plots – Categorical Distributions – Plot Types
-
Categorical Plots – Categorical distributions – Code in Python
-
Seaborn Grid
-
Matrix graphs
-
Seaborn Test Exercises
-
Solutions for Seaborn Test Exercises
Large Data Visualization Project
-
Machine Learning Overview
00:00 -
Machine Learning Overview
00:00 -
Overview of the Data Visualization Project
-
Analysis of project solutions – Part 1
-
Analysis of project solutions – Part 2
-
Analysis of project solutions – Part 3
Machine Learning Overview
-
Section overview
-
Why machine learning is needed
-
Types of Machine Learning Algorithms
-
Process for supervised learning
-
(OPTIONAL) Additional Reading Book – ISLR
Linear Regression
-
Overview of the linear regression section
-
Linear Regression – History of the Algorithm
-
Least squares
-
Cost Function
-
Gradient Descent
-
Simple Linear Regression
-
Scikit-Learn Review
-
Scikit-Learn – Train Test Split
-
Scikit-Learn – evaluation of model performance
-
Residual Plots
-
Implementation of the model and interpretation of coefficients
-
Polynomial regression – theory
-
Polynomial Regression – Feature Generation
-
Polynomial Regression – Model Training and Evaluation
-
Bias-Variance Trade-Off Dilemma
-
Polynomial regression – choosing the degree of the polynomial
-
Polynomial Regression – Model Implementation
-
Regularization – overview
-
Feature scaling
-
Cross Validation – Overview
-
Regularization – data preparation
-
L2 Regularization – Ridge regression – theory
-
L2 Regularization – Ridge Regression – Code in Python
-
L1 Regularization – Lasso Regression – theory and code in Python
-
L1 and L2 Regularization – Elastic Net
-
Data Review for Linear Regression Validation Project
-
Dealing with missing data – Part 1 – Assessing the situation
00:00
Feature Engineering and Data Preparation
-
Feature Engineering Overview
-
Working with outliers
-
Dealing with missing data – Part 1 – Assessing the situation
-
Working with missing data – Part 2 – Working by row
-
Working with missing data – Part 3 – Working by columns
-
Working with Categorical Variables
Cross Validation and Linear Regression Validation Project
-
Working with outliers
-
Working with missing data – Part 2 – Working by row
-
Working with missing data – Part 3 – Working by columns
-
Working with Categorical Variables
Logistic regression
-
Overview of the section on logistic regression
-
Logistic Regression Theory – Part 1 – Logistic Function
-
Logistic Regression Theory – Part 2 – Transition from Linear to Logistic
-
Logistic Regression Theory – Part 3 – Mathematics of Transition
-
Logistic Regression Theory – Part 4 – Finding the Best Graph
-
Logistic Regression in Scikit-Learn – Part 1 – Data Exploration
-
Logistic Regression in Scikit-Learn – Part 2 – Creating and Training a Model
-
Classification metrics – Confusion Matrix and Accuracy
-
Classification metrics – Precision, Recall and F1-Score
-
Classification Metrics – ROC Curves
-
Logistic Regression in Scikit-Learn – Part 3 – Evaluating Model Performance
-
Multi-class classification – Logistic regression – Data mining
-
Multi-class classification – Logistic regression – Model
-
Logistic Regression Validation Project
-
Solutions for Logistic Regression Validation Project
K-Nearest Neighbors method (KNN)
-
Overview of the section on the K-nearest neighbors method
-
K-nearest neighbors theory
-
KNN: writing code in Python – Part 1
-
KNN: writing code in Python – Part 2
-
KNN Test Exercises
-
Solutions for KNN Test Exercises
Support Vector Machines (SVM)
-
Overview of the support vector machine section
-
History of support vector machines
-
Support Vector Machine Theory – Hyperplanes and Margins
-
Theory of support vector machine – kernels
-
Support vector machine theory – “kernel trick” and mathematics (optional)
-
SVM in Scikit-Learn for Classification Problems – Part 1
-
SVM in Scikit-Learn for Classification Problems – Part 2
-
SVM in Scikit-Learn for regression problems
-
Support Vector Machine Testing Exercises
-
Solutions for Support Vector Machine Verification Exercises
Decision Trees
-
Overview of the section on decision trees
-
Decision Trees – History
-
Decision Trees – Terminology
-
Decision Trees – Gini Impurity Metric”
-
Building Decision Trees Using Gini Impurity – Part 1
-
Building Decision Trees Using Gini Impurity – Part 2
-
Python Code for Decision Trees – Part 1 – Data
-
Python Code for Decision Trees – Part 2 – Model
Random Forests
-
Overview of the section about random forests
-
History and motivation for creating random forests
-
Random Forest Hyperparameters – Overview
-
Random Forest Hyperparameters – Number of Trees and Number of Features
-
Random Forest Hyperparameters – Bootstrapping and oob_score
-
Data classification using RandomForestClassifier – Part 1
-
Data classification using RandomForestClassifier – Part 2
-
Regression with RandomForestRegressor – Part 1 – Data Overview
-
Regression with RandomForestRegressor – Part 2 – Basic Models
-
Regression with RandomForestRegressor – Part 3 – Polynomial Model
-
Regression with RandomForestRegressor – Part 4 – Other models
Boosted Trees
-
Boosting section overview
-
History of boosting
-
AdaBoost – Theory – How adaptive boosting works
-
AdaBoost – Code in Python – Data
-
AdaBoost – Code in Python – Model
-
Gradient boosting – Theory
-
Gradient Boosting – Writing Code in Python
Supervised Learning Test Project
-
Review of the verification project
-
Decision Analysis – Part 1 – Exploratory Data Analysis
-
Analysis of solutions – Part 2 – Churn analysis
-
Decision Analysis – Part 3 – Decision Tree Models
NLP (Natural Language Processing) and Naive Bayes Classifier
-
Overview of the section on NLP and Naive Bayes
-
Naive Bayes – Part 1 – Bayes’ Theorem
-
Naive Bayes – Part 2 – the algorithm itself
-
Extracting features from text – Theory
-
Extracting features from text – “Bag of words” – writing the code manually
-
Extracting features from text using Scikit-Learn
-
Text classification – Part 1
-
Text classification – Part 2
-
Test exercises on text classification
-
Solutions for Test Exercises on Text Classification
Machine learning without a teacher – Unsupervised Learning
-
Unsupervised Learning Review – Unsupervised Learning
K-Means Clustering
-
Overview of the section on K-means clustering
-
Principles of data clustering (without reference to a specific algorithm)
-
K-means clustering theory
-
K-Means Clustering – Writing Code – Part 1
-
K-Means Clustering – Writing Code – Part 2
-
Selecting the number of clusters K – Theory
-
Selecting the number of clusters K – Writing code in Python
-
Color quantization – Theory
-
Color Quantization – Writing Code in Python
-
K-Means Clustering Test Exercises
-
Solutions for K-Means Clustering Test Exercises – Part 1
-
Solutions for K-Means Clustering Test Exercises – Part 2
-
Solutions for K-Means Clustering Test Exercises – Part 3
Hierarchical data clustering
-
Overview of the section on hierarchical clustering
-
Theory and intuition of hierarchical clustering
-
Hierarchical Clustering – Writing Code, Part 1 – Data
-
Hierarchical Clustering – Writing Code, Part 2 – Scikit-Learn
DBSCAN – Data Density Based Clustering
-
Overview of the DBSCAN clustering section
-
Theory of the DBSCAN algorithm
-
Comparing DBSCAN and K-Means Clustering
-
Key hyperparameters of DBSCAN – Theory
-
DBSCAN Key Hyperparameters – Code in Python
-
DBSCAN Test Exercises
-
Solutions for DBSCAN Test Exercises
Principal component analysis (PCA – Principal Component Analysis)
-
Review of the section on principal component analysis
-
Principal Component Analysis Theory – Part 1 – History and Intuition
-
Principal Component Analysis Theory – Part 2 – Mathematics
-
Implementing a Principal Component Method Manually
-
Principal Component Method in Scikit-Learn
-
Test exercises using the principal component method
-
Solutions for Principal Component Analysis Test Exercises
Student Ratings & Reviews
No Review Yet