Epython Lab
6.15K subscribers
679 photos
31 videos
104 files
1.3K links
Welcome to Epython Lab, where you can get resources to learn, one-on-one trainings on machine learning, business analytics, and Python, and solutions for business problems.

Buy ads: https://telega.io/c/epythonlab
Download Telegram
๐๐ž๐Ÿ๐จ๐ซ๐ž ๐œ๐ก๐š๐ง๐ ๐ข๐ง๐  ๐ฒ๐จ๐ฎ๐ซ ๐Œ๐‹ ๐ฆ๐จ๐๐ž๐ฅ, ๐œ๐ก๐ž๐œ๐ค ๐ฒ๐จ๐ฎ๐ซ ๐๐š๐ญ๐š.
When a model performs badly, the first thing we often do is try a different algorithm.

Sometimes that works.

But before doing that, I usually look at the dataset. ๐Ÿ”

I check things like:

๐Ÿ”น Missing values
๐Ÿ”น Duplicate records
๐Ÿ”น Outliers
๐Ÿ”น Wrong data types
๐Ÿ”น Class imbalance
๐Ÿ”น Data leakage
๐Ÿ”น High-cardinality columns
๐Ÿ”น Features with little useful information

There is no point spending hours tuning a model if the dataset itself has problems. โš ๏ธ

A simple workflow I prefer is:

๐Ÿ“ฅ ๐‘น๐’‚๐’˜ ๐‘ซ๐’‚๐’•๐’‚
โ†“
๐Ÿ”Ž ๐‘ช๐’‰๐’†๐’„๐’Œ ๐‘ธ๐’–๐’‚๐’๐’Š๐’•๐’š
โ†“
๐Ÿงน ๐‘ช๐’๐’†๐’‚๐’
โ†“
๐Ÿ“Š ๐‘จ๐’๐’‚๐’๐’š๐’›๐’†
โ†“
๐Ÿค– ๐‘ป๐’“๐’‚๐’Š๐’
โ†“
๐Ÿ“ˆ ๐‘ด๐’๐’๐’Š๐’•๐’๐’“

Data quality is not just something to deal with before machine learning. It affects every step that comes after it.

So when a model is not performing as expected, don't immediately change the model.

๐Ÿ” Take another look at the data first. https://youtube.com/playlist?list=PL0nX4ZoMtjYHTtowSzzB2gVH2AuuoF9WW&si=EhLKvJCVlYQXknOs
Use Data quality checker tool: https://datasetdoctor.fastapicloud.dev
๐Ÿ‘3โค1
Your model can look excellent and still be wrong.

One of the first things I check when evaluating an ML dataset is data leakage. ๐Ÿ”

Data leakage happens when information that would not actually be available at prediction time gets into the training data.

For example:

๐Ÿฅ Healthcare
You are predicting whether a patient will be admitted, but your dataset includes a field recorded after admission.

๐Ÿ’ณ Fraud detection
You are predicting fraud, but one of the features is created after the transaction has already been investigated.

๐Ÿ“ฆ Customer churn
You are predicting who will leave, but the training data contains information that only becomes available after the customer leaves.

The result?

Your model may show:

๐Ÿ“ˆ 98% accuracy
๐Ÿ“ˆ Excellent validation results
๐Ÿ“ˆ Great performance during testing

Then you put it into production...

And the performance drops.

The problem was not necessarily the model.

The model had access to information it would never have in the real world.

That is why I don't look at model performance alone.

I also ask:

๐Ÿ”Ž Where did each feature come from?
โฑ๏ธ When was it created?
๐ŸŽฏ Would this information actually be available when making the prediction?

A high score is not always a good score.

Sometimes, it is a warning sign.

Check out data quality issues
https://youtube.com/playlist?list=PL0nX4ZoMtjYHTtowSzzB2gVH2AuuoF9WW&si=EhLKvJCVlYQXknOs

Also checkout data quality checker tool https://datasetdoctor.fastapicloud.dev

#MachineLearning #DataScience #AI #DataLeakage #MLOps #Python
The hardest part of building AI applications isn't writing the prompt or calling the model. In the last two weeks, I learned that keeping the backend from turning into spaghetti code once you move past the tutorial phase.

When you're wiring up an AI document pipeline in FastAPI, a few things quickly become non-negotiable:



โ€ข Payload Guardrails: If your Pydantic schemas aren't catching malformed JSON, missing nested fields, or bad Enums at the door, your AI service will fail unpredictably downstream.



โ€ข Route Isolation: Mixing your raw API endpoints with validation logic and business rules makes refactoring a nightmare by week three.



โ€ข The Persistence Gap: Transitioning from mock in-memory data structures to a real relational database and a vector store for RAG is where most clean prototypes start to break down.



If you're building production backends for AI and ML features, where do you usually draw the line between keeping things simple and over-engineering your architecture?



You can learn about FastAPI: https://www.youtube.com/playlist?list=PLQNCas8_eikM



#FastAPI #Python #BackendEngineering #SoftwareArchitecture #APIs #Pydantic #ArtificialIntelligence #MachineLearning #RAG
FastAPI Episode 7: Advanced Pydentic Model Design
https://www.youtube.com/watch?v=j4lLM6tWKKk
โค4
When building a FastAPI application, I pay close attention to how data is validated before it reaches the business logic.

That is where Pydantic becomes particularly useful.

Pydantic lets us define the structure and rules for the data our application accepts, instead of scattering validation checks throughout the codebase.

I use it to define and enforce the data contract at the application boundary.

For example:

๐‘๐‘™๐‘Ž๐‘ ๐‘  ๐‘ƒ๐‘Ÿ๐‘œ๐‘๐‘’๐‘ ๐‘ ๐‘–๐‘›๐‘”๐ถ๐‘œ๐‘›๐‘“๐‘–๐‘”(๐ต๐‘Ž๐‘ ๐‘’๐‘€๐‘œ๐‘‘๐‘’๐‘™):

๐‘โ„Ž๐‘ข๐‘›๐‘˜_๐‘ ๐‘–๐‘ง๐‘’: ๐‘–๐‘›๐‘ก = ๐น๐‘–๐‘’๐‘™๐‘‘(๐‘”๐‘ก=0)

๐‘œ๐‘ฃ๐‘’๐‘Ÿ๐‘™๐‘Ž๐‘: ๐‘–๐‘›๐‘ก = ๐น๐‘–๐‘’๐‘™๐‘‘(๐‘”๐‘’=0)

@๐‘š๐‘œ๐‘‘๐‘’๐‘™_๐‘ฃ๐‘Ž๐‘™๐‘–๐‘‘๐‘Ž๐‘ก๐‘œ๐‘Ÿ(๐‘š๐‘œ๐‘‘๐‘’="๐‘Ž๐‘“๐‘ก๐‘’๐‘Ÿ")

๐‘‘๐‘’๐‘“ ๐‘ฃ๐‘Ž๐‘™๐‘–๐‘‘๐‘Ž๐‘ก๐‘’_๐‘๐‘œ๐‘›๐‘“๐‘–๐‘”(๐‘ ๐‘’๐‘™๐‘“):

๐‘–๐‘“ ๐‘ ๐‘’๐‘™๐‘“.๐‘œ๐‘ฃ๐‘’๐‘Ÿ๐‘™๐‘Ž๐‘ >= ๐‘ ๐‘’๐‘™๐‘“.๐‘โ„Ž๐‘ข๐‘›๐‘˜_๐‘ ๐‘–๐‘ง๐‘’:

๐‘Ÿ๐‘Ž๐‘–๐‘ ๐‘’ ๐‘‰๐‘Ž๐‘™๐‘ข๐‘’๐ธ๐‘Ÿ๐‘Ÿ๐‘œ๐‘Ÿ("๐‘œ๐‘ฃ๐‘’๐‘Ÿ๐‘™๐‘Ž๐‘ ๐‘š๐‘ข๐‘ ๐‘ก ๐‘๐‘’ ๐‘ ๐‘š๐‘Ž๐‘™๐‘™๐‘’๐‘Ÿ ๐‘กโ„Ž๐‘Ž๐‘› ๐‘โ„Ž๐‘ข๐‘›๐‘˜_๐‘ ๐‘–๐‘ง๐‘’")

๐‘Ÿ๐‘’๐‘ก๐‘ข๐‘Ÿ๐‘› ๐‘ ๐‘’๐‘™๐‘“

Both fields can be individually valid, while their combination is not.

That is the difference between:

Field validation โ†’ Is this value valid?

Model validation โ†’ Is this combination valid?

With Pydantic, ๐™๐™ž๐™š๐™ก๐™™(), ๐™›๐™ž๐™š๐™ก๐™™_๐™ซ๐™–๐™ก๐™ž๐™™๐™–๐™ฉ๐™ค๐™ง(), ๐™–๐™ฃ๐™™ ๐™ข๐™ค๐™™๐™š๐™ก_๐™ซ๐™–๐™ก๐™ž๐™™๐™–๐™ฉ๐™ค๐™ง() let us keep these rules close to the schema.

The result is a cleaner boundary:

๐™๐™š๐™ฆ๐™ช๐™š๐™จ๐™ฉ โ†’ ๐™‘๐™–๐™ก๐™ž๐™™๐™–๐™ฉ๐™ž๐™ค๐™ฃ โ†’ ๐˜ฝ๐™ช๐™จ๐™ž๐™ฃ๐™š๐™จ๐™จ ๐™‡๐™ค๐™œ๐™ž๐™˜ โ†’ ๐˜ฟ๐™–๐™ฉ๐™–๐™—๐™–๐™จ๐™š / ๐˜ผ๐™„

For production AI applications, this matters. Document metadata, processing parameters, search filters, and structured AI outputs all need predictable contracts.

Good schemas do more than describe data. They protect the rest of the system.


You can explore more: https://www.youtube.com/playlist?list=PLQNCas8_eikM


#Python #Pydantic #PydanticV2 #FastAPI #BackendEngineering #AIEngineering