Data eXplore : Data Science, ML, Big Data, LLMs and AI Security
583 subscribers
845 photos
446 videos
1 file
675 links
Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
★ @DataML
Download Telegram
Ralph Mode for Deep Agents

What if we give the agent a task and let it run endlessly?

Developed Ralph Mode based on Deep Agents specifically for such an experiment.

Ralph Mode cycles the agent repeatedly, with each pass using a clean context, and the file system is used as memory. You can start it, step away, and then stop it with Ctrl+C when you're done (or set limits in advance).

This video shows how to run Ralph Mode together with Deep Agents and automatically compile an entire Python course.


Video, Repo

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
The most comprehensive review of RL that I've seen.

It was written by Kevin Murphy from Google DeepMind, who has over 128k citations.

How it differs from other materials on RL:

→ There's a bridge between classical RL and the current era of LLMs:

A separate chapter on LLMs and RL, which discusses:

RLHF, RLAIF, and reward modeling
PPO, GRPO, DPO, RLOO, REINFORCE++
Training reasoning models
Multi-turn RL for agents

Scaling computations for inference (test-time compute scaling)

→ The basics are explained very clearly

All major algorithms like value-based methods, policy gradients, and actor-critic are explained with mathematical rigor.

→ Model-based RL and world models are also well-covered

There's Dreamer, MuZero, MCTS, and more on the list - this is exactly where the field is heading now.

→ A section on multi-agent RL

Game theory, Nash equilibrium, and MARL for LLM agents.


••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Everyone is talking about n8n, but it's worth taking a closer look at Sim.

This is an open-source platform for building AI agents:

✓ Next.js + Bun + PostgreSQL + Zustand stack
✓ You can connect any AI model
✓ You can deploy it on your own server

GitHub

•••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Tencent introduced a diffusion language model: 6X faster than classic LLMs

WeDLM-8B Instruct does not use autoregression like regular LLMs,
but a diffusion method for text generation.

What does this provide?
🚀 In mathematical reasoning tasks, the model works 3–6 times faster
than Qwen3-8B even with vLLM optimizations - while maintaining quality.

This release breaks the old myth that "diffusion models are not suitable for precise text tasks".

In practice, WeDLM shows that such an approach can compete
and even outperform transformers in inference speed.


The model is open and available under the Apache 2.0 license:

GitHub, HuggingFace

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
30 terms from the field of agent-based AI that AI development engineers should

🤖 @DataXplore
Do you want to learn AI on real projects?
In this repository, there are 29 projects with Generative AI, Machine Learning, and Deep Learning.

With full code for each one. This is pure gold: https://github.com/KalyanM45/AI-Project-Gallery

🤖 @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Stokes' theorem is a classic of vector analysis.

Essentially, it states that the linear integral of a vector field over a closed contour is equal to the surface integral of the rotor of this field over the surface bounded by this contour.

🤖 @DataXplore
🎤Fun-ASR: speech recognition system

Fun-ASR is a powerful speech recognition model, trained on millions of hours of real data.

It supports 31 languages and is optimized for accurate recognition in noisy environments and various dialects. Ideal for educational and financial applications.

🚀 Key features:
- High recognition accuracy in noisy conditions (up to 93%)
- Support for 7 Chinese dialects and 26 regional accents
- Multilingual support with the ability to freely switch between languages
- Recognition of song lyrics against music backgrounds

GitHub #python

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Connect Telegram notifications in Claude Code

Hapi.run, a wrapper for CLI agents (Claude Code, Codex, Gemini-cli). clone of happy.engineering (which is buggy and glitchy, but has its advantages).

Consists of three parts:
- Server
- Client
- Daemon (but we don't need it today)

➡️ Quick launch:
1️⃣ Install the project:

npm install -g @twsxtd/hapi

2️⃣ Enter the bot token (first, you need to create it in @botfather):

export TELEGRAM_BOT_TOKEN=111:TOKEN_BOTA

3️⃣ Start the server:

hapi server

In the screenshot, I showed how the server generates a token (a password). If everything is done correctly, you'll see at the bottom that your bot has started.

4️⃣ However, the bot won't work without a tunnel (it needs to attach buttons from the mini app):
• cloudflare (without registration)
• ngrok (with registration)
• or something else.

brew install cloudflare

cloudflare tunnel --url http://localhost:3006

You'll see the tunnel address. Which you need to export:

export WEBAPP_URL="https://your-public-url"

As you've probably guessed, the client will connect to the server. To start the client, enter:

hapi

The client will ask you to enter the token, which you already received in the previous step. If you forgot to start the server - it will be started automatically.

5️⃣ In your bot, be sure to press /start

The bot will respond with a welcome message with a link to the mini app - enter the token in the mini app. Now you can manage your coding agents from your desktop and mobile simultaneously and receive all notifications via SMS in the Telegram bot.

All settings are stored in ~/.hapi/settings.json - edit them if something has changed, for example the tunnel address, or check the token there if you've forgotten it.


🦾 @PromptXplore
Holidays are over, Hope everyone enjoyed!!

I'm back! content returns to its regular schedule today now.

– InXplore
Find rows with min/max values in another column in a single line in Polars v1.37.0

Previously, to extract a row with the minimum or maximum value relative to another column, you usually had to do sorting, groupby, or set up more complex filters.

In Polars v1.37.0, the min_by and max_by expression methods have been introduced. They find the minimum or maximum value for any column with a single, understandable expression.

To update and get min_by/max_by:
pip install -U polars

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Flow-generative models, trained through flow matching, typically learn curved trajectories, and they are difficult to approximate in a few steps.

Rectified flows attempt to learn straight trajectories, which are easier to simulate and require fewer computations.


Here's an interactive article that explains the geometric intuition behind Rectified Flows.

The code can also be found here.

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
How to let LLMs sort context by importance themselves?

Conventional language models read text as one long strip.

What's closer to the beginning of attention - that's "more important".
What's further - the model sees it worse. And here a problem arises: if an important fact is hidden somewhere far away in the noise, the model may simply not use it.

It spends attention on everything, instead of focusing on the main thing.

🟢 What Sakana AI has figured out?
Sakana AI proposed a solution - RePo (Context Re-Positioning).

The idea is very clear: the model gets a module that allows to dynamically "re-position" the context.

Like a person:
you read a long document, realize that the important part was 20 pages back - and mentally re-read it, ignoring the rest.

What RePo does
- pulls important pieces of information closer
- pushes away noise and excess text
- helps the model's attention focus on what's needed

In the model, there's a trainable module that re-assigns the positions of tokens according to meaning, not order

✅ important = what helps reduce the model's error and solve the task correctly
❌ secondary = what doesn't help (noise), so it's "pushed away" in terms of positions

As a result, the model with such a memory starts to work better where LLMs usually suffer:
- when the context is long
- when there's a lot of noise
- when important details are scattered far from each other
- when the data is structured (tables, lists, rules)

The authors show that RePo gives a noticeable increase in robustness, without worsening the overall quality.

▶️ Robustness to noise (Noisy Context)
Average result on 8 noisy benchmarks:

- Regular RoPE: 21.07
- RePo: 28.31

🟡 Increase: +7.24 points (strongly)

The authors separately note a key figure:
on noisy-eval (4K context) RePo is better than RoPE by +11.04 points.

🔥 Examples of increase on specific tasks
(everywhere RePo > RoPE)

- TriviaQA: 61.47 → 73.02 (+11.55)
- GovReport: 6.23 → 16.80 (+10.57)
- 2WikiMultihopQA: 23.32 → 30.86 (+7.54)
- MuSiQue: 7.24 → 13.45 (+6.21)

This is a step towards models that don't just "read what they're given", but are able to organize their working memory themselves.


Details, Article • #RePo #SakanaAI #LLM #AI #AIAgents #Context #LongContext #Attention

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Advice for AI engineers

You can run production-level LLM inference on a laptop CPU or even on a phone.

No cloud accounts. No API keys. No internet.

LFM2.5-1.2B-Instruct from liquidai gives:

239 tokens/s on an AMD CPU
82 tokens/s on a mobile NPU
less than 1 GB of RAM

Get the link

••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
Contextual Personalization of the assistant

OpenAI added a guide on Context Engineering for the Agents SDK to their cookbook, and this is probably the most competent approach to memory management.

Instead of rummaging through thousands of old messages, the agent maintains a structured user profile and a "notebook".

🟢 How it works? benifits & Pitfalls
☞ State Object: a centralized information hub in the form of a JSON object, which is stored locally. It includes a profile (hard facts: name, ID, loyalty status) and notes (unstructured notes: "likes hotels in the center").

⁠☞ Injection: before each launch, this state is fed into the system prompt in YAML format: for the profile and Markdown for the notes. Not everything at once, of course, but only what is needed at the moment.

⁠☞ Distillation: the most interesting part. The agent doesn't just chat, it has a tool save_memory_note. If you said in the conversation: "I don't eat meat", the agent calls this tool and saves the Session Note (temporary note) in real time.

⁠☞ Consolidation: garbage collection for memory. After the session ends, a separate process is launched, which takes the temporary notes, compares them with the global ones, removes duplicates and resolves conflicts according to the principle "the newer overrides the older".

BENEFITS:

⁠☞ agent starts behaving like a personal assistant without retraining.
⁠☞ There are clear rules: what the user said now > session notes > global settings.
⁠☞ We don't mix everything together, but separate hard data (for example, from CRM) and soft data (preferences from the chat).

OpenAI's approach of dividing into Session Memory and Global Memory looks reliable, but requires direct hands in writing the consolidation logic. Without this, your agent will quickly turn into a demented grandfather who remembers what never happened.

PITFALLS:

You need to make a separate call to the LLM after each dialogue to tidy up the memory. If the model glitches at this stage, it may write a hallucination into the "long memory" or delete something important. Here, strict frameworks are needed.

The context window is not elastic. Although models have a huge context, dragging "War and Peace" from the user's notes is costly in terms of money and timing. You will have to periodically trim the history, leaving only the essence.


Guide • #AI #ML #LLM #Guide #OpenAI

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
🌟 NVIDIA reinvents memory: LLMs that continue learning during inference

Contextual windows are growing, but there are two chairs: either classic attention, which feeds on memory and computes like crazy, or RNN-like Mamba, DeltaNet, which work quickly but start drifting and losing details in long contexts.

🟢 What NVDIA Proposed?

NVIDIA proposes a solution that tries to sit on both chairs at once - Test-Time Training with End-to-End formulation (TTT-E2E):

Usually, the model's weights are frozen after training. When you feed it data, it just holds it in the KV cache. In TTT, everything is different: the context is the training dataset itself. While the model reads your prompt (the context), it updates its weights (more precisely, performs gradient descent on the fly), thereby embedding the context's information into the model itself. This allows you to compress gigantic volumes into a fixed state size without bloating the KV cache to the sky.

⁠➜ result - beauty and magic:

⁠☞ Latency of inference becomes constant. It doesn't matter if there are 100 tokens in the context or a million - the time to generate the next token is the same.

⁠☞ On a context of 128k tokens - a 2.7x speedup compared to Attention (on H100). On 2M tokens - a 35x speedup.

☞ Unlike Mamba and other RNNs, the quality doesn't degrade over long distances. TTT maintains the same level as full attention.

⁠➜ Of course, there are a bunch of points with an asterisk

⁠☞ Training is complex. To allow the model to learn on the fly so skillfully, it needs to be pretrained specially. This process is currently 3.4x slower than regular training.

⁠☞ method requires calculating gradients from gradients during training. FlashAttention currently doesn't support this out of the box, requiring custom kernels or workarounds.

⁠☞ process of consuming context during inference itself requires computations during the prefill phase.

In the end, NVIDIA compares RAG to a notebook, and its TTT to the real updating of neural connections in the brain. If you want to delve into the methodology and grasp the idea - the code and the paper are publicly available.


GitHub | Paper • #AI #ML #LLM #TTTE2E #NVIDIA

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
KVzap: Nvidia learned to use memory 3–4 times more efficiently in inference

The KV-cache is today the main Achilles' heel of transformers when scaling the context. It grows linearly with the length of the sequence and is stored for each layer and each head.

For example, for a LLaMA-like model with 65B parameters, the KV-cache at 128k tokens occupies ~335 GB of memory. And it's also a pain in terms of time.

🟢 How NVIDIA solved Pain?

However, most optimizations reduce the KV-cache by layers or by heads. Although the main potential is precisely along the token axis: not all of them are really needed by the model.

The first working method of reducing KV by tokens was invented by the authors of KVzip: up to 4× compression with zero quality loss. But in practice, the method turned out to be too slow.

Nvidia took this idea, modified it a bit, and got almost the same result, but practically for free.

They simply train a small model that predicts, based on the hidden state of a token, how important its KV is. It's different for each layer, but it's either a linear model or a two-layer MLP - a maximum of 1–2 matrix multiplications.

And that's it, no expensive operations and recalculations (for comparison: in KVzip, the prompt essentially had to be run twice). Next, the KV pairs with a significance below a specified threshold are simply discarded.

The compute overhead is about 0.02% FLOPs for linear models. On a long context, this is noise against the quadratic attention.

The degradation on benchmarks is about zero, the compression is 3–4×. It's just a fairy tale (though, of course, much still depends on the engine).


Hats off to Nvidia for their excellent work. Everything is in the open source on GitHub | Paper

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
This repository collects everything you need to use AI and LLM in your projects.

120+ libraries, organized by development stages:

→ Model training, fine-tuning, and evaluation
→ Deploying applications with LLM and RAG
→ Fast and scalable model launch
→ Data extraction, crawlers, and scrapers
→ Creating autonomous LLM agents
→ Prompt optimization and security


Get Here

••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
An excellent tool to estimate how much VRAM your LLM really needs.

You change the hardware configuration, quantization, etc. and immediately see:

generation speed (tokens/sec)
exact memory allocation
system throughput and more

Try Here

••••••••••••••••••••••••••••••••••••••
🤖 @DataXplore
Document Index for Vectorless RAG Based on Reasoning

PageIndex is an open-source RAG framework that eliminates vector databases and chunking from the pipeline when searching for documents.

🟢 How it works?

Most RAG systems rely on semantic similarity: they cut a document into pieces, build embeddings, and then retrieve fragments that "resemble the query".

But similarity does not equal relevance.

In professional documents such as financial reports, legal documents, and technical manuals, multi-step parsing and domain-specific logic are often needed. Vector-based search easily gets stuck when almost every section uses the same terminology.

PageIndex does it differently.

It builds a hierarchical tree from the document, similar to a table of contents, but tailored for LLMs. Then, it uses reasoning-based tree search to "navigate" the structure in the same way a human expert would.

The two-step process:

1. Generate a tree-based index of the document structure
2. Retrieve the needed information through reasoning-based tree search

The LLM can "think" about the document structure. Instead of matching embeddings, it reasons like: "Trends in debt are usually in the financial summary or Appendix G, let's look there."

Key features:

• No vector database or embedding pipeline
• No artificial chunking that breaks context at boundaries
• Traceable retrieval with precise references down to the page level
• Navigation based on reasoning, mirroring human document analysis

PageIndex is used in Mafin 2.5 and claims 98.7% accuracy on FinanceBench for financial document analysis.


And yes, it's completely open source.

••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore