Machine Learning
40.7K subscribers
3.64K photos
31 videos
47 files
670 links
Real Machine Learning โ€” simple, practical, and built on experience.
Learn step by step with clear explanations and working code.

Admin: @HusseinSheikho || @Hussein_Sheikho
Download Telegram
Cheat sheet for Scikit-learn: ๐Ÿ“š Scikit-learn is a Python library for machine learning.

๐Ÿ“ฅ Loading Data - downloading and preparing data.
๐Ÿงผ Preprocessing - standardization, normalization, and feature processing.
๐Ÿ—๏ธ Create Your Model - creating models for classification, regression, and clustering.
๐ŸŽฏ Model Fitting - training the model on data.
๐Ÿ”ฎ Prediction - obtaining forecasts.
๐Ÿ“Š Evaluate Performance - assessing the quality of the model using various metrics.
๐Ÿ”„ Cross-Validation - checking the model on different samples.
โš™๏ธ Tune Your Model - optimizing parameters using Grid Search and Randomized Search.

#ScikitLearn #MachineLearning #Python #DataScience #AI #MLOps

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค6๐Ÿ‘2
๐Ÿ”– The Legendary MIT Textbook on Mathematics for Computer Science

Mathematics for Computer Science is one of the best free textbooks for developers, ML engineers, and data scientists.

It contains over 1000 pages covering discrete mathematics, logic, graphs, probability, combinatorics, recurrence relations, and other fundamental topics.

โ›“๏ธ Link to the textbook:
https://people.csail.mit.edu/meyer/mcs.pdf

#ComputerScience #Mathematics #MachineLearning #DataScience #MIT #OpenSource

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค6
Combining Plots in Matplotlib ๐Ÿ“Š

In Matplotlib, you can easily combine multiple plots in a single window using the subplot() function. Simply create the necessary plots, specify their layout, add titles, and you'll get a clear visualization for easy data comparison.

#Matplotlib #DataVisualization #Python #DataScience #Coding #Plotting

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค6๐Ÿ‘1
Reinforcement Learning Methods and Tutorials ๐Ÿง ๐Ÿ“š

In these tutorials for reinforcement learning, it covers from the basic RL algorithms to advanced algorithms developed recent years.

Learning Resources: https://github.com/MorvanZhou/Reinforcement-learning-with-tensorflow ๐Ÿš€

Here's a collection of simple materials on methods and practical guides, covering both basic reinforcement learning algorithms and modern, recently developed, and updated advanced algorithms. ๐Ÿ“–โœจ

#ReinforcementLearning #MachineLearning #AI #DeepLearning #TechTutorials #DataScience

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค5
Feature Scaling: Why Feature Scaling Affects Model Training

Feature scaling is often overlooked because it seems like just another data preprocessing step. However, in practice, it often helps models train faster and more stably. Imagine one feature has values ranging from 0 to 1, while another has values ranging from 0 to 10,000. Although both features may be equally important for prediction, it's more difficult for the optimizer to work with such data.

This means it has to take more steps to find a good solution. Additionally, regularization becomes less effective because features with different scales require coefficients of different magnitudes. Let's look at how this looks in a simple example.

Install dependencies:
pip install numpy scikit-learn

Import libraries:
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score

Let's create a small synthetic dataset. It will have two features: the first has a normal scale, and the second is about a thousand times larger.

Importantly, both features actually influence the target variable. That is, the only difference between them is the scale.
np.random.seed(42)
x_small = np.random.normal(0, 1, 300)
x_large = np.random.normal(0, 1000, 300)

X = np.vstack([x_small, x_large]).T

y = (x_small + 0.001 * x_large > 0).astype(int)

Now, let's split the data into training and testing sets. We won't scale anything yetโ€”first, let's see how the model behaves on the original data.
X_train, X_test, y_train, y_test = train_test_split(
X, y,
test_size=0.3,
random_state=42,
stratify=y
)

Let's train a logistic regression model without scaling.

In addition to the model's quality, let's also look at the number of iterations (n_iter_). This metric shows how much work the optimizer had to do to find the coefficients.
model = LogisticRegression()
model.fit(X_train, y_train)

pred = model.predict_proba(X_test)[:, 1]

print("ROC-AUC:", roc_auc_score(y_test, pred))
print("Iterations:", model.n_iter_)

Now, let's scale the features to the same scale using StandardScaler.

It calculates the mean and standard deviation only for the training set and then uses the same values for the test set. This is important because the model should not "peek" at the test data during training.

After this transformation, both features are approximately on the same scale, and it becomes easier for the optimizer to work with them.
scaler = StandardScaler()

X_train_scaled = scaler.fit_transform(X_train)
X_test_scaled = scaler.transform(X_test)

Now, let's retrain the model.

We're using the same model, the same data, and the same parameters. The only difference is that the features are now scaled.
model = LogisticRegression()
model.fit(X_train_scaled, y_train)

pred = model.predict_proba(X_test_scaled)[:, 1]

print("ROC-AUC (scaled):", roc_auc_score(y_test, pred))
print("Iterations (scaled):", model.n_iter_)

Most often, the ROC-AUC doesn't change much. However, the number of iterations becomes smaller. This means that the optimizer found a solution faster, and the training was more stable.

๐Ÿ”ฅ Feature scaling is a simple data preprocessing step that, in many cases, allows the model to train faster and more stably. For logistic regression, SVMs, neural networks, and other algorithms that use numerical optimization, it's best not to skip it.

โœจ #DataScience #MachineLearning #Python #Coding #Tech #AI

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค5๐Ÿ‘1
Foundations of Applied Mathematics is a free series of four textbooks created for the applied and computational mathematics program at Brigham Young University. ๐Ÿ“š

The series includes four volumes:
*   Mathematical Analysis
*   Algorithms, Approximation, and Optimization
*   Uncertainty and Data
*   Dynamics and Control

The series is suitable for upper-level undergraduate and introductory graduate students. It also includes Python lab exercises and practical assignments, connecting mathematical theory with numerical computation, algorithms, data analysis, and scientific applications. ๐Ÿ

I particularly appreciate that these are not just theoretical textbooks. The accompanying Python materials help to illustrate how these concepts are applied to real-world computational problems. ๐Ÿ’ป

https://foundations-of-applied-mathematics.github.io

#Mathematics #Python #Education #DataScience #Algorithms #Learning

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค4
This media is not supported in your browser
VIEW IN TELEGRAM
A Powerful Alternative to Pandas ๐Ÿš€

This is an optimized replacement for Pandas that can significantly speed up data processing without requiring major changes to your code. โš™๏ธ

To get started, simply replace a single import:

import fireducks.pandas as pd

Performance Benchmarks demonstrate speed improvements in various use cases. ๐Ÿ“ˆ

More: https://colab.research.google.com/drive/1UIokuJ4cytoiVSabRDqcziDXOan8bVua?usp=sharing

#Pandas #Python #DataScience #Performance #Fireducks #BigData

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค2
Day 7 of self-studying Berkeley CS189 โ€” stochastic gradient descent notes ๐Ÿ“š๐Ÿ“

๐Ÿ”ฅ *Stochastic Gradient Descent (SGD)* is a powerful optimization algorithm used to minimize loss functions in machine learning. Unlike batch gradient descent, which uses the entire dataset to compute gradients, SGD updates parameters using a single training example (or a small mini-batch) at a time.

๐Ÿš€ Key Benefits:
- Faster convergence on large datasets
- Escapes local minima more easily
- Suitable for online learning scenarios

๐Ÿ“Š The Update Rule:
ฮธ = ฮธ - ฮฑ * โˆ‡J(ฮธ; xโฝโฑโพ, yโฝโฑโพ)
Where ฮฑ is the learning rate and (xโฝโฑโพ, yโฝโฑโพ) is a single training example.

๐Ÿ“Œ Challenges:
- High variance in updates
- Requires careful tuning of the learning rate

๐Ÿง  *Tip:* Use momentum or adaptive learning rates (like Adam) to stabilize training!

#MachineLearning #CS189 #SGD #DeepLearning #DataScience #Algorithms

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค5
This media is not supported in your browser
VIEW IN TELEGRAM
๐Ÿ”– Learning Data Science through interactive examples

One of the most useful repositories for those who want to better understand machine learning.

It transforms complex concepts into visual experiments: you can study models, change parameters, and immediately see the results.

โ›“ Link to GitHub
https://github.com/GeostatsGuy/DataScienceInteractivePython

#DataScience #MachineLearning #Python #Learning #Tech #GitHub

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค5
๐Ÿ”– Over 300 real-world case studies of ML systems from top companies. ๐Ÿค–

We found a repository that collects genuine ML engineering experience โ€“ not theory from textbooks, but real stories of implementing models in production. ๐Ÿ“š

Inside, you'll find case studies from Uber, Netflix, Google, and other companies: how they built the architecture, what problems arose, where the systems failed, and what solutions helped them recover. ๐Ÿ—๏ธ

โ›“ Link to GitHub
https://github.com/Engineer1999/A-Curated-List-of-ML-System-Design-Case-Studies

#MachineLearning #MLCaseStudies #DataScience #Engineering #Uber #Netflix

โœจ Join Best TG Channels https://t.me/addlist/0f6vfFbEMdAwODBk

โญ๏ธ Join Our WhatsApp Channel https://whatsapp.com/channel/0029VaC7Weq29753hpcggW2A
โค2