Data eXplore : Data Science, ML, Big Data, LLMs and AI Security
583 subscribers
845 photos
446 videos
1 file
675 links
Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Download Telegram
🧠 Intelligent router for LLM

Semantic Router directs requests to the OpenAI API based on semantic understanding, selecting the most suitable models from the pool. It uses BERT classification to improve inference accuracy and offers security features such as PII detection and jailbreak protection.

🟢 Highlights:-
- Auto-selection of models to optimize requests
- Tools for selection based on request context
- PII detection and protection
- Caching of semantic representations to speed up processing

GitHub

#python

🤖 Data Science, ML & Big Data with @DataXplore
🚀 Powerful multimodal models LLaVA-OneVision-1.5

An open platform for training multimodal models, demonstrating outstanding results at low cost. The models are trained on high-quality data and provide excellent efficiency.

🟢 Features:
- Fully open source code and training data
- High quality and diversity of training data
- Efficient architecture for economical training
- Support for modern technologies such as MoE and FP8
- Optimized code for scalability


GitHub

#python

🤖 Data Science, ML & Big Data with @DataXplore
🧠 LIMIT: A Study of Retrieval Limits Based on Embeddings

The repository contains the LIMIT dataset, created to test embedding models on theoretical principles. The study shows that even modern models cannot retrieve certain documents, highlighting the limitations of the current approach using single-vector embeddings.

🚀Key points:
- Dataset for testing embedding models.
- Includes 50k documents and 1000 queries.
- Highlights theoretical limitations of information retrieval.
- Code for data generation and experiments is available in the repository.


GitHub

#python

🤖 Data Science, ML & Big Data with @DataXplore
🤖 Tongyi DeepResearch: A powerful language model for deep search specially designed for deep information-oriented tasks. It

🟢 Features:
- High performance on complex search tasks with 30.5 billion parameters.

- Demonstrates outstanding results on various benchmarks, including Humanity's Last Exam and WebWalkerQA,

- Fully automated data synthesis process and advanced reinforcement learning methods.

- Compatibility with multiple inference paradigms.

- Efficient training using agent interaction data.


GitHub

#python

🤖 Data Science, ML & Big Data with @DataXplore
ShinkaEvolve: Evolution of programs with AI

ShinkaEvolve is a framework that combines large language models with evolutionary algorithms to automate scientific discoveries. It enables improving scientific code by leveraging the creative capabilities of AI and optimization through evolution, supporting parallel evaluation of candidates.

🟢 Key points
- Combines LLM and evolutionary algorithms.
- Supports parallel evaluation on local machines and clusters.
- Stores an archive of successful solutions for knowledge transfer.
- Optimizes performance while maintaining code correctness.
- Ideal for scientific tasks with available verifiers.


GitHub

#python

🤖 Data Science, ML & Big Data with @DataXplore
Start Learning AI & ML - From Python To Harvard Level

1️⃣ FREE mini-courses on
Python, DS and ML
What's inside:
• Completely and highly practical.
• Python, Pandas, visualization
• Basics of machine learning and feature engineering
• Data preparation and working with models

Practice without unnecessary theory: you learn and immediately apply.
👉 Course


2️⃣ A huge collection of the 17 best GitHub repositories for learning Python
☞ 30-Days-Of-Python
— covers python basics

☞ Python Basics — simple-clear Python basics for beginners

☞ Learn Python — guide with examples and code

☞ Python Guide — best practices, tools, advanced topics

☞ Learn Python 3 — An easy-to-understand guide to Python 3 with practice

☞ Python Programming Exercises — 100+ Python problems

☞ Coding Problems — algorithmic problems, perfect for interview prep

☞ Project-Based-Learning — learn Python through real projects

☞ Projects — ideas for practical skill improvement

☞ 100-Days-Of-ML-Code — a step-by-step guide to ML in Python

☞ TheAlgorithms/Python — a huge collection of algorithms in Python

☞ Amazing-Python-Scripts — useful scripts from automation to advanced utilities

☞ Geekcomputers/Python — a collection of practical scripts: networking, files, automation

☞ Materials — code, exercises, projects from Real Python

☞ Awesome Python — a top list of best frameworks and libraries

☞ 30-Seconds-of-Python — short snippets for quick solutions

☞ Python Reference — life hacks, tutorials, and useful scripts


3️⃣ Harvard Machine Learning Course
Iconic CS 249 track has been turned into an interactive textbook - arguably one of best starting points for engineers who want to build real ML systems, not just play with models.

• Complete ML foundation: explains fundamentals from scratch, only Python knowledge is required
• System design and data engineering
• Dataset preparation, MLOps, monitoring
• AI deployment in IoT and production

Its a practical course: not about formulas, but about how to implement ML so that it brings business profit.
If you want to understand how models operate in production - an ideal start.
👉 Course


#Python #AI #DataScience #ML #freecourses

🤖 Data Science, ML & Big Data with @DataXplore
Build Agents, Scale Systems, Ship Projects

1️⃣ Fresh series on Python and AI by Microsoft
• Content: Includes 9 lectures supplemented with videos, detailed presentations, and code examples. series - training in AI agent development - accessible, clearly written, even for programming beginners.
• Topics: cover topics such as RAG (Retrieval-Augmented Generation), embeddings, agents, and the MCP protocol.
👉 Course


2️⃣ Ready-made guide to train & host LLM from scratch by HuggingFace
– Architectures, their features, and hyperparameter optimization
– Working with data
– Pretraining and the pitfalls involved
– Post-training: all modern approaches and how to apply them
– Infrastructure, how to build and optimize it properly
Guide with 200+ pages, 7 big chapters, read + lots of diagrams and examples with Simple English.


3️⃣ Database sharding guide from PlanetScale
Learn how to scale databases through sharding - splitting data across servers to increase performance and fault tolerance.

• Sharding is needed when a single database can no longer handle the load.
• There are two popular approaches — range-based and hash-based.
• It is important to choose a stable key (e.g., user_id) and avoid cross-shard queries.
• A proxy layer slightly increases latency but provides scalability.

Excellent material if you want to understand how systems at YouTube scale. And here is a lot of SQL basics Read


#Python #AI #DataScience #ML #freecourses

🤖 Data Science, ML & Big Data with @DataXplore
Reconstruction of Visual Space from Any Views

Depth Anything 3 (DA3) is a model that predicts spatially consistent geometry from arbitrary visual inputs.

🟢 Key Features:
- The DA3 model outperforms previous versions in depth estimation.
- Supports monocular and multi-view depth estimation.
- High-accuracy pose estimation.
- User-friendly interface and export options to various formats.
- Specialized models for metric depth evaluation.
- Uses a simple transformer and a unique depth representation, enabling high performance in depth and pose estimation.


GitHub #python

•••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
🎤Fun-ASR: speech recognition system

Fun-ASR is a powerful speech recognition model, trained on millions of hours of real data.

It supports 31 languages and is optimized for accurate recognition in noisy environments and various dialects. Ideal for educational and financial applications.

🚀 Key features:
- High recognition accuracy in noisy conditions (up to 93%)
- Support for 7 Chinese dialects and 26 regional accents
- Multilingual support with the ability to freely switch between languages
- Recognition of song lyrics against music backgrounds

GitHub #python

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Autonomous AI research with autoresearch

This repository proposes concept of autonomous AI learning, where agent modifies code and conducts experiments on its own.

Key points:
- The autonomous agent modifies train.py to optimize the model.

- learning process takes place in a fixed timeframe of 5 minutes, after which it evaluates the results and continues the iterations.

- Using a simple interface program.md, users can configure the agent to optimize models without directly interfering with the code.
- Support for only one NVIDIA GPU.


GitHub #python

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Practice for PyTorch Interviews

TorchCode offers a structured environment for training programming skills needed for ML interviews. Solve problems implementing operators and architectures, receiving instant feedback and hints.

➡️ Key features:
- 40 tasks frequently encountered in interviews
- Automatic verification of correctness and performance
- Instant feedback on each test
- Hints and reference solutions for study
- Ability to run in the browser without installation


GitHub #Python

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore
Media is too big
VIEW IN TELEGRAM
One of the most powerful open-source models for computer vision SAM 3.1 released

Model understands what is happening in an image or video and is able to find objects based on a text description. You can literally write "a person in a red T-shirt" and it will find necessary people.

➡️ How it works?
It works not only with pictures, but also with videos. The object can be specified once, and then the model will track it between frames.

The key idea - open-vocabulary. The model is not limited to fixed classes, like old systems. It operates with a huge number of concepts and can find almost any object.

Another important point is that you can combine control methods: text, clicks, frames, masks. This gives much more control and accuracy.

Under the hood a new architecture, where the tasks of object search and tracking are solved separately. Due to this, the model better distinguishes similar things and works more stably on video.

The repository already has everything for getting started: ready weights, code, examples, and notebooks.

In fact, this is no longer just a tool for labeling, but a full-fledged vision engine that can be integrated into real products from video analytics to data labeling automation.

Now the model can track up to 16 objects in one pass.

With multiplexing, all objects are processed simultaneously:

• fewer unnecessary calculations
• no memory bottlenecks


RESULT: video processing speed increases by approx 2 times, from 16 to 32 FPS on a single NVIDIA H100!

On new SA-CO benchmark, which includes 270 thousand unique concepts, SAM 3 achieves 75–80% of human level.

GitHub | #AI #ML #LLM #CV #python

••••••••••••••••••••••••••••••••••••••
🤖 Data & ML | @DataXplore