Data eXplore : Data Science, ML, Big Data, LLMs and AI Security
583 subscribers
845 photos
446 videos
1 file
675 links
Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Download Telegram
Cursor is completely switching to dynamic context for all models

Means the agent (based on any model) will now primarily collect context on its own, rather than using what has been provided.

🟢 How is Dynamic Context different from "Classic" approach?

Static context is a classic approach. You dump all the logs, documentation, chat history, descriptions of all toolboxes, MCP, etc. into the agent's context at once. In general, this works, but the context ends up being filled with a lot of irrelevant information and is constantly overflowing.

Now, Cursor is offering Dynamic context discovery. This involves placing a conditional "table of contents" and links in the context, while the rest is scattered across files, and the agent can add information to itself as needed. For example:

➖ Everyone remembers that when the context overflows, Cursor performs summarization and updates the window, right? Now, in addition to this, Cursor stores chat history as a file. After summarization, the agent receives a link to this file, and if some necessary detail was lost in the summary, he can search the history and supplement himself.

➖ Long responses from tool calls are now also recorded in files, rather than being sent directly into the context. Only a link to the necessary output appears in the context, while the gigantic JSON file sits waiting for the agent to access it and search for what he needs using conditional grep or tail.

➖ The same applies to MCP, Agent Skill, and terminal sessions. Bulky tool descriptions and terminal outputs are stored not in the context, but in files. The context simply says "MCP is available for jira, datadog, figma", and the agent, if he needs something, goes to the detailed description and invokes the tool.

It turns out to be quite nice and practical. On A/B tests, the overall token consumption has decreased by ~46.9%.


And it's also scalable, because here the context transforms from a place where knowledge is stored into an instruction on how to retrieve it. Find details

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
How to use LLM without losing quality?

DFlash is a way to speed up text generation for large models.

HOW and WHY to use?

Works like this: one model quickly creates a draft, and another one checks it and corrects errors.

- 6.2× faster without losing quality on Qwen3-8B
- 2.5 times faster than EAGLE-3

The idea is simple:

• Diffusion models - generate quickly, but sometimes make mistakes
• Autogenerative (AR) - very accurate, but work slowly
• DFlash combines both approaches:
diffusion - draft → AR - checking and confirmation


Both quickly and accurately, instead of choosing one or the other.

Blog, Code, Models

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
SleepFM model diagnoses 130 diseases by analyzing a single night's sleep.

Stanford trained SleepFM model (fundamental for predicting a range of pathologies) from atrial fibrillation and myocardial infarction to dementia and Parkinson's disease.

🔴 Why traditional ML models fail despite having gigabytes of data?
Polysomnography is the "gold standard" for studying sleep: a person is fitted with sensors (EEG, ECG, respiration, muscles) and gigabytes of raw signals are recorded.

But in the ML world, this data is used ineffectively. Existing models were trained on small datasets for specific tasks (finding apnea, determining sleep phases).

A huge amount of physiological information about a patient's health was simply ignored, because it's impossible to manually label hundreds of hours of recordings for each disease.

Moreover, if the EEG sensor was mounted slightly differently in one clinic or fell off, the usual model would break down.


🟢 How training on 585k hours possible without human labels?
At the university, they realized that they didn't need human labelers, they needed volumes. They collected a huge dataset of 585,000 hours of sleep recordings from more than 65,000 patients and invented a unique SSL learning algorithm for the future model.

1️⃣ LOO-CL (Leave-One-Out Contrastive Learning)

Instead of teaching the model to predict a diagnosis, they made it solve a puzzle: the system receives input signals from 3 modalities (heart, muscles, respiration) and must predict the embedding of the fourth (brain waves).

This forces the neural network based on 1D CNN and Transformers to learn deep, hidden connections between physiological processes.

2️⃣ The second feature is Channel-Agnostic Attention.

The models don't care about which sensors are connected and in what order. If a channel fails or is absent, attention pooling simply redistributes weights, and inference continues.

3️⃣ SleepFM has learned to read sleep not just for insomnia.

Having received a single night of recordings as input, the model predicts the risk of 130 diseases, and it does this more accurately than specialized models trained with a teacher: the risk of Parkinson's disease is detected in 89% of cases, dementia in 85%, and the probability of a heart attack in 81%.


Such diagnostics could move from labs to smartwatches with development of wearable electronics, and tests shown that noise of sleep signals can hide a patient's entire medical record.

Details #news #AI #ML

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Convert PDF files into clean data, ready for LLM.

Dolphin is a document parsing framework that converts PDFs into structured formats: Markdown, HTML, LaTeX, and JSON.

Works and Features:

Stage 1️⃣ Detailed analysis of the layout at the page level. Elements and their order are determined according to the natural reading order.

Stage 2️⃣ Parallel parsing of elements using different types of anchors and task-specific prompts.

Key features:

» Open-source
» A two-stage approach of analyze-then-parse based on a single VLM
» Encouraging performance on document parsing tasks
» Generation of a sequence of elements in the natural reading order
» Heterogeneous anchor prompts for different types of document elements
» An efficient parallel parsing mechanism


GitHub

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
On subject of progress: An agent from SakanaAI took a confident first place in a coding competition.

🟢 How did a "wrapper" beat 800 human experts?

Last year, an agent from OpenAI only took second place in the same competition.

This year, about 800 people participated in the AtCoder Heuristic Contest. The ALE-Agent from the Japanese laboratory outperformed everyone and took the top spot with a significant lead. The cost of solution was approximately £1,300.

Interestingly, authors of this year's optimization task themselves expected a classic approach using annealing and constructive heuristics, but the Sakana agent took a different path. He suddenly implemented the virtual power heuristic, which allowed him to escape local optima even better than human experts.

Agent is a rather clever wrapper over (in this case) GPT-5.2 high and Gemini 3 Pro high.


Sakana themselves never really shone in terms of models, but they learned to work competently with inference time scaling and here's the result.

In a word, well done! Read Here

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Binary search + rescoring in int8.

The simple strategy to search through 40 million texts in ~200 ms using only a CPU server, 8GB of RAM, and 45GB of disk space.

If you want to try it out immediately, there's a demo of 40 million texts from Wikipedia. No login or other hassles required.

🟢 The Inference Strategy:

☞ Embed the query with a dense model into a regular fp32 vector

☞ Quantize the fp32 embedding into a binary format, which is 32 times smaller

☞ Retrieve, for example, 40 documents (about 20 times faster than an fp32 index) using an approximate or exact binary index

☞ Load the int8 embeddings for these top-40 documents from the disk

☞ Rescoring: fp32 embedding of the query × 40 int8 embeddings

☞ Sort these 40 documents by the new score and take the top-10

☞ Load the titles and texts of the top-10 documents

The documents are embedded once, and then these embeddings are used in two representations:

1️⃣ A binary index (I used IndexBinaryFlat for exact search and IndexBinaryIVF for approximate search)

2️⃣ An int8 view, i.e., a way to quickly read int8 embeddings from the disk by document ID

➡️ In end, instead of fp32 embeddings, you store:

- a binary index (32 times smaller)
- int8 embeddings (4 times smaller)

Plus, Only the binary index is kept in memory, so the RAM savings are also x32 compared to fp32 search.
For comparison: a regular fp32 retrieval on such a task would require about 180GB of RAM, 180GB of disk for embeddings, and would be 20–25 times slower.

A binary retrieval with int8 rescoring fits into about 6GB of RAM and ~45GB of disk for embeddings.

For example, if you load 4 times more documents through the binary index and then rescoring them in int8, you can recover about 99% of the quality of fp32 search (compared to ~97% for pure binary search)


HFblog
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security
Live stream scheduled for
Voicechat 2 of Saturday series with industry pros.

Q/A with QA Tester

📅 Time: TODAY, Jan 10, 2026 | 4:30 PM UTC (10 PM IST)
🎧 Listen Only: Join Livestream

🎤 Want to Speak or Ask?
Comment for Speaker Link

#DataXplore #AI #ML #VoiceChat

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Ralph Mode for Deep Agents

What if we give the agent a task and let it run endlessly?

Developed Ralph Mode based on Deep Agents specifically for such an experiment.

Ralph Mode cycles the agent repeatedly, with each pass using a clean context, and the file system is used as memory. You can start it, step away, and then stop it with Ctrl+C when you're done (or set limits in advance).

This video shows how to run Ralph Mode together with Deep Agents and automatically compile an entire Python course.


Video, Repo

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
The most comprehensive review of RL that I've seen.

It was written by Kevin Murphy from Google DeepMind, who has over 128k citations.

How it differs from other materials on RL:

→ There's a bridge between classical RL and the current era of LLMs:

A separate chapter on LLMs and RL, which discusses:

RLHF, RLAIF, and reward modeling
PPO, GRPO, DPO, RLOO, REINFORCE++
Training reasoning models
Multi-turn RL for agents

Scaling computations for inference (test-time compute scaling)

→ The basics are explained very clearly

All major algorithms like value-based methods, policy gradients, and actor-critic are explained with mathematical rigor.

→ Model-based RL and world models are also well-covered

There's Dreamer, MuZero, MCTS, and more on the list - this is exactly where the field is heading now.

→ A section on multi-agent RL

Game theory, Nash equilibrium, and MARL for LLM agents.


••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Everyone is talking about n8n, but it's worth taking a closer look at Sim.

This is an open-source platform for building AI agents:

✓ Next.js + Bun + PostgreSQL + Zustand stack
✓ You can connect any AI model
✓ You can deploy it on your own server

GitHub

•••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Tencent introduced a diffusion language model: 6X faster than classic LLMs

WeDLM-8B Instruct does not use autoregression like regular LLMs,
but a diffusion method for text generation.

What does this provide?
🚀 In mathematical reasoning tasks, the model works 3–6 times faster
than Qwen3-8B even with vLLM optimizations - while maintaining quality.

This release breaks the old myth that "diffusion models are not suitable for precise text tasks".

In practice, WeDLM shows that such an approach can compete
and even outperform transformers in inference speed.


The model is open and available under the Apache 2.0 license:

GitHub, HuggingFace

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
30 terms from the field of agent-based AI that AI development engineers should

🤖 @DataXplore
Do you want to learn AI on real projects?
In this repository, there are 29 projects with Generative AI, Machine Learning, and Deep Learning.

With full code for each one. This is pure gold: https://github.com/KalyanM45/AI-Project-Gallery

🤖 @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Stokes' theorem is a classic of vector analysis.

Essentially, it states that the linear integral of a vector field over a closed contour is equal to the surface integral of the rotor of this field over the surface bounded by this contour.

🤖 @DataXplore
🎤Fun-ASR: speech recognition system

Fun-ASR is a powerful speech recognition model, trained on millions of hours of real data.

It supports 31 languages and is optimized for accurate recognition in noisy environments and various dialects. Ideal for educational and financial applications.

🚀 Key features:
- High recognition accuracy in noisy conditions (up to 93%)
- Support for 7 Chinese dialects and 26 regional accents
- Multilingual support with the ability to freely switch between languages
- Recognition of song lyrics against music backgrounds

GitHub #python

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Connect Telegram notifications in Claude Code

Hapi.run, a wrapper for CLI agents (Claude Code, Codex, Gemini-cli). clone of happy.engineering (which is buggy and glitchy, but has its advantages).

Consists of three parts:
- Server
- Client
- Daemon (but we don't need it today)

➡️ Quick launch:
1️⃣ Install the project:

npm install -g @twsxtd/hapi

2️⃣ Enter the bot token (first, you need to create it in @botfather):

export TELEGRAM_BOT_TOKEN=111:TOKEN_BOTA

3️⃣ Start the server:

hapi server

In the screenshot, I showed how the server generates a token (a password). If everything is done correctly, you'll see at the bottom that your bot has started.

4️⃣ However, the bot won't work without a tunnel (it needs to attach buttons from the mini app):
• cloudflare (without registration)
• ngrok (with registration)
• or something else.

brew install cloudflare

cloudflare tunnel --url http://localhost:3006

You'll see the tunnel address. Which you need to export:

export WEBAPP_URL="https://your-public-url"

As you've probably guessed, the client will connect to the server. To start the client, enter:

hapi

The client will ask you to enter the token, which you already received in the previous step. If you forgot to start the server - it will be started automatically.

5️⃣ In your bot, be sure to press /start

The bot will respond with a welcome message with a link to the mini app - enter the token in the mini app. Now you can manage your coding agents from your desktop and mobile simultaneously and receive all notifications via SMS in the Telegram bot.

All settings are stored in ~/.hapi/settings.json - edit them if something has changed, for example the tunnel address, or check the token there if you've forgotten it.


🦾 @PromptXplore
Holidays are over, Hope everyone enjoyed!!

I'm back! content returns to its regular schedule today now.

– InXplore