Machine Learning Basics
Machine learning is a subfield of artificial intelligence (AI) that involves developing algorithms and models that enable computer systems to learn and improve from experience without being explicitly programmed. In other words, machine learning algorithms use statistical analysis to identify patterns and relationships in large datasets and use these patterns to make predictions or decisions.
In order to familiarize with machine learning 4 type field we need to cover.
- Supervised learning
- Unsupervised learning
- Reinforcement learning
- Deep learning
It’s important to note that machine learning is a complex and rapidly evolving field, so it’s important to be patient and persistent in your learning journey. As you gain more experience, you may want to explore more advanced topics such as natural language processing also called NLP, computer vision, and deep reinforcement learning.
We’re starting with Supervised Learning concept today.
Supervised Learning
Supervised learning is a type of machine learning in which a model is trained on labeled data. Labeled data means that each example in the dataset has both input features and an associated output label. The goal of supervised learning is to learn a mapping between the input features and the output labels so that the model can make accurate predictions on new, unseen data.
In supervised learning, the model is typically trained using an optimization algorithm to minimize the difference between its predicted output and the actual output labels in the training data. Once trained, the model can be used to make predictions on new input data by computing the model’s output for that data.
Examples of supervised learning problems include image classification, sentiment analysis, and speech recognition. Some popular algorithms used in supervised learning include;
- Linear regression
- Logistic regression
- Decision trees
- Neural networks.
Today’s example, we’ll use two different algorithms on same data set. These are;
- DecisionTreeClassifier
- Linear Regression
If you’re ready let’s start with learning the DecisionTreeClassifer algorithm and behind it’s logic.
Decision Tree Classifier
DecisionTreeClassifier is a classification algorithm in scikit-learn that builds a decision tree from the training data to predict the class label of new samples. A decision tree is a flowchart-like structure where each internal node represents a test on an attribute, each branch represents the outcome of the test, and each leaf node represents a class label.
The DecisionTreeClassifier algorithm works by recursively partitioning the training data into subsets based on the values of the features, such that each subset contains as much pure classes (i.e., samples of the same class label) as possible. The purity of a subset is measured using a criterion such as Gini impurity or entropy, which measures the degree of homogeneity within the subset. The algorithm continues to recursively partition the subsets until a stopping criterion is met, such as reaching a maximum depth or minimum number of samples in a leaf node.
Once the decision tree is constructed, it can be used to predict the class label of new samples by following the path from the root node to a leaf node based on the values of the features of the new sample. The class label of the leaf node that the sample reaches is then returned as the predicted class label.
The DecisionTreeClassifier in scikit-learn offers various parameters that can be tuned to optimize the performance of the model, such as the maximum depth of the tree, the criterion for measuring the impurity of the subsets, and the minimum number of samples required to split an internal node.
In summary, DecisionTreeClassifier is a classification algorithm that builds a decision tree from the training data to predict the class label of new samples by recursively partitioning the subsets of the training data based on the values of the features.

Linear Regression
Linear regression is a supervised learning algorithm in machine learning that models the relationship between a dependent variable (target) and one or more independent variables (features) as a linear equation. The goal of linear regression is to find the best-fit line that describes the relationship between the input features and the target variable.
In a simple linear regression model, there is only one independent variable, while in a multiple linear regression model, there are multiple independent variables. The model assumes a linear relationship between the features and the target, which means that the target variable can be expressed as a weighted sum of the input features, where each weight represents the contribution of the corresponding feature.
The main benefits of using a linear regression model are:
-
Interpretability: The model provides a clear understanding of how the target variable is related to the input features. The coefficients of the linear equation can be interpreted as the contribution of each feature to the target variable.
-
Simplicity: The model is simple and easy to understand, which makes it a good choice for solving simple regression problems.
-
Efficiency: The model is computationally efficient and can be trained quickly on large datasets.
-
Flexibility: Linear regression can be extended to handle more complex relationships by adding polynomial terms, interaction terms, or by applying other techniques such as regularization.
-
Versatility: Linear regression can be used for both continuous and categorical target variables, making it a versatile model for a wide range of regression problems.
In summary, linear regression is a popular and widely used machine learning algorithm that models the relationship between the input features and the target variable as a linear equation. Its benefits include interpretability, simplicity, efficiency, flexibility, and versatility.

Hands-on Python Example
Here is the python example of UCI student performance dataset you can download dataset from this link. We’ll be using two model we talked about and also calculate MSE for best profit.
1import pandas as pd
2from sklearn.tree import DecisionTreeClassifier
3from sklearn.linear_model import LinearRegression
4from sklearn.model_selection import train_test_split
5from sklearn.metrics import accuracy_score, mean_squared_error
6
7# Load the student-mat.csv dataset
8df = pd.read_csv('./student/student-mat.csv', delimiter=';')
9
10#Preprocessing dataset for our models
11featuresForPreprocessing = []
12for column in df.columns:
13 if isinstance(df[column][0], str):
14 featuresForPreprocessing.append(column)
15
16# Convert categorical variables to dummy variables
17df = pd.get_dummies(df, columns=featuresForPreprocessing)
18# Convert target variables to binary values
19df['pass'] = df['G3'].apply(lambda x: 1 if x >= 10 else 0)
20
21# Select features and target variable
22features = df.drop(['G1', 'G2', 'G3', 'pass'], axis=1)
23target = df['pass']
24
25# Split the data into training and testing sets
26X_train, X_test, y_train, y_test = train_test_split(features, target, test_size=0.2, random_state=42)
27
28LRmodel = LinearRegression()
29LRmodel.fit(X_train, y_train)
30
31predictonLR = LRmodel.predict(X_test)
32
33DTCmodel = DecisionTreeClassifier(random_state=42)
34DTCmodel.fit(X_train, y_train)
35
36predictonDTC = DTCmodel.predict(X_test)
37
38# Calculate MSE for both models
39LRmse = mean_squared_error(y_test, predictonLR)
40DTCmse = mean_squared_error(y_test, predictonDTC)
41
42# Print the results
43print("Linear Regression MSE:", LRmse)
44print("Decision Tree Classifier MSE:", DTCmse)
Thank you for visitig for this blog. I hope you liked that post.