Data eXplore : Data Science, ML, Big Data, LLMs and AI Security
583 subscribers
845 photos
446 videos
1 file
675 links
Exploring Data Science, Big Data Analytics & Visualization, ML/DL, Neural Networks, LLMs with GitHub, Kaggle, HuggingFace and some white papers by big institutions.

Not just data, but science behind data

Paid project? premodi@zohomail.in
β˜… @DataML
Download Telegram
This media is not supported in your browser
VIEW IN TELEGRAM
Vector Index vs Vector Database, in simple terms

Developers often use these terms as synonyms but if you don't understand the difference, it can lead to problems in production.

🟒 How to think about it?
A vector index is essentially a search algorithm.

You feed it with vectors, it organizes them into a structure that allows you to quickly search for similar ones (e.g., HNSW), and finds nearest neighbors. FAISS is also an example of this.

But the nuance is that's all.

It's not responsible for storage, it can't properly filter by metadata, and it doesn't scale on its own. It just does the search.

A vector DB is a wrapper around the index, plus everything else that's actually needed in production.

It includes distributed storage, persistence, filtering by metadata, concurrent access, and more. A good open-source example: Milvus.

From this, it's clear: one is a component, the other is a system.

🟑 Why this matters?
Once, a company with autopilots was building a search for road videos, and the scale was huge.

Each trip generated frames, each frame turned into an embedding.

Engineers needed to ask something like "night crosswalks with pedestrians" based on months of data.

At first, FAISS looked perfect: fast, lightweight, easy to set up.

But as the data grew, the embeddings of each day became a separate index file.

After a couple of months, they had hundreds of thousands of scattered files.

Searching for several days meant digging through tons of files at once.

Complex queries required custom DBs, query planners, and filtering built around FAISS.

In the end, they had billions of vectors and no clear path forward.

This is where vector DBs come in. The company migrated to Milvus, and the difference became obvious:

↳ one query: similarity + filters by metadata
↳ data in collections and partitions, not scattered across files
↳ tens of billions of vectors, a year in production, no major incidents
↳ 30% reduction in infrastructure costs
↳ 10x scaling headroom

And this isn't a unique story.

Most teams struggle when they start with a lightweight index and then suddenly need filters, reliable storage, and real scale.

Vector DBs exist precisely for this moment.

Milvus stands out for its ability to handle scale and different types of data well.

You can store billions of vectors, scale horizontally, and create specialized indexes, for example for geodata with its optimized index, not "one common for everything".


Completely open-sourced on GitHub (41k+ stars), can be self-hosted or used in their cloud.

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
How to do personalized search using only Postgres

There are two phases:

DATA PREPARATION PHASE:

1️⃣ Generate movie embeddings for vector search and create a BM25 index for full-text search.

2️⃣ Generate user preference embeddings based on what the user has watched and what they have liked or disliked before.

SEARCH PHASE:

1️⃣ Retrieve the top-100 movies ranked by BM25.

2️⃣ Normalize BM25 scores to a range of 0–1.

3️⃣ Perform personalized search: compare movie embeddings with the user's preference embedding.

4️⃣ Combine signals: 50% for text match (relevance) and 50% for user match (personalization).
Guide Get Here

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Came across an interesting project: tiny-infini-gram.

Training-free language model (without training) which, according to him, generates Shakespeare 250 times faster than nanoGPT.

What's inside, in theory?

This is the "unbounded n-gram" approach: you can use very large n without running into exponential memory like with classic n-grams.

Instead of a huge n-gram table, a suffix array is used: it simulates n-gram lookup of any size with logarithmic access time.

Previously, such things were hardly used for generation, because sampling broke down: infinite perplexity and frequent verbatim copying.

The author claims that he solved this with a new method called Selective Back-off Interpolation Sampling: it mixes probability distributions from several levels of n-grams to maintain a balance between quality and novelty.

If you like non-standard LM approaches and "fast, simple, without training" - it's worth checking out the analysis and code.


Link to detailed write-up

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
A new class of risks for open-source models, which few people think about.

A work on so-called elicitation attack: took open-source model, retrain it on seemingly harmless data on chemical synthesis, which were generated by frontier models.

And suddenly, this open-source model starts to perform significantly better on tasks related to chemical weapons.

🟒 What the paper shows?
The most unpleasant thing here is not "how to make the model respond to prohibited questions". But the fact that the model can be dangerous, even if it itself does not output anything harmful. Because its harmless answers can become training data that unlock dangerous capabilities in another model.

What the authors showed:
☞ The attack works on different open-source models and on different types of "weapon" tasks
☞ Retraining on data from frontier models gives a greater boost than training on chemistry textbooks or on data,
☞ Generated by the same open-source model
☞ Sufficiently "peaceful" topics: cheese making, fermentation, candle chemistry, etc.
☞ In one experiment, "harmless chemistry" gave about 2/3 of the effect on the growth of "weapon" competence compared to training on data about chemical weapons
☞ The stronger the frontier model, the stronger the subsequent uplift of the open-source model (and the higher the risk)

CONCLUSION is simple and rather harsh: focusing only on "refusal training" in frontier models doesn't solve problem. The danger can leak through normal, seemingly everyday answers, which someone then uses as a dataset.


Read Here
β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Boosted Up performance of AI agent by 184% using a completely open-source technique.

Now, you can automatically find best prompts for any agentic workflow you're putting together, means manual prompt engineering isn't needed at all.

🟒 What is the simple idea?

1️⃣ Take a starting prompt and an eval dataset
2️⃣ Then, an optimizer iteratively improves the prompt
3️⃣ In the end, you get an optimal prompt automatically

And all of this in just a few lines of code.

➑️ Why Opik specifically?

Opik is a 100% open-source platform for evaluating LLMs.

It helps optimize LLM systems so that they work better, faster, and cheaper: from RAG chatbots to code assistants. Opik includes tracing, evaluations, and dashboards.


Best Part: Everything can be run completely locally, because you can use any local LLMs as optimizers and evaluators.

GitHub repository

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
πŸ‹ DeepSeek-OCR 2 is a new generation of OCR with SOTA quality

A 3B model for advanced understanding of images, documents and OCR, which reaches the SOTA level.

🟒 Why this matters?

The key novelty is DeepEncoder V2.

Unlike classic vision LLMs, which "read" the image as a grid (left-to-right, top-to-bottom), DeepEncoder V2 works closer to how a human reads:

- First, a global understanding of the image is formed
- Then, the model determines the logical order of reading - what is important first, what next

What this brings in practice

πŸ“„ Works better with complex document layouts
πŸ“Š Correctly reads tables
🧾 Links signatures and values
πŸ“° Understands columns and structured text
πŸ”€ More reliably processes a mixture of text and visual structure

In terms of quality

- Outperforms Gemini 3 Pro on a number of benchmarks
- Gives >4% improvement compared to the previous version of DeepSeek-OCR

And this is with a model size of just 3B parameters.

Can be launched and fine-tuned


Now, DeepSeek-OCR 2 can be conveniently launched and fine-tuned via Unsloth according to the ready-made guide.

Guide, Model, Github, Paper | #DeepSeek #ocr #opensource

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
To train an ML model, you need to be proficient in algorithms, write code, endlessly tune hyperparameters which is a high entry barrier for most people.

An open-source project Plexe significantly lowers this threshold: You describe the task in plain language and it automatically assembles machine learning for it.

🟒 How it works?

☞ Explain in a human-friendly way what you want to predict, what the input data is and what the output should be.

⁠☞ Next the system, through a combination of several agents, goes through the entire pipeline: data analysis, solution plan, code generation, tests, and quality assessment.

⁠☞ Supports various LLM providers: OpenAI, Anthropic, Ollama, and others. Plus, it can automatically derive the data structure or even generate a synthetic dataset.

⁠☞ There's also distributed training on Ray inside: you can run multiple model variants in parallel and significantly speed up the process.


GitHub

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
HOW YOLO BECAME A STANDARD IN CV?

Launching a series of posts about the evolution of one of the most popular architectures in computer vision.

We'll break down:

Before 2015, the task of detection was solved by searching for the most likely regions. There were two-stage approaches, such as Faster R-CNN.

🟒 How did the YOLO architecture evolve from v1 to v3?

πŸ“ First, they searched for candidate regions, and then used a refine process to refine the classes and coordinates.

PROBLEM: The process was very slow. Imagine the task of tracking a tennis ball on the court during a match. Old networks would have taken 5 minutes to process a video, even on a good GPU. Players would have had to stand and wait for the VAR system.

A real-time approach was needed, where speed was more important than perfect results. Thus, YOLO was born.

➑️ YOLO v1: A model that looks at the entire scene (2015)
The idea was to turn detection from a region-searching task into a regression problem. Combine all stages into a single network that directly "spits out" coordinates.

How it was implemented technically?

πŸ“ They made an architecture similar to GoogLeNet. Two fully connected and 24 convolutional layers. Although it was large, it detected bounding boxes and immediately determined the coordinates.

πŸ“ All images were divided into a 7x7 grid. Each cell predicted 2 bounding boxes and 20 classes. The input was a 448x448 image, which was further divided into 64x64.

PROBLEM: YOLO v1 couldn't handle other resolutions. To work with detection on large images, they resized them to 448x448 or cut them into patches. Due to the extra operations, the main advantage over Faster R-CNN β€” speed β€” was lost.

➑️ YOLO v2 / YOLO9000: Scale and anchors (2016–2017)

To level the complex LOSS, multi-scale was added to the new version: YOLO9000 simultaneously detects more than 9,000 classes without full annotation β€” hence the name.

What new features were added?

πŸ“ Anchor Boxes: Instead of directly predicting coordinates, they switched to predicting shifts relative to the X and Y axes for candidates. This maximized object capture.

πŸ“ Skip Connections: They introduced pass-through layers and added batch normalization, which solved the problem of gradient fading.

PROBLEM: The accuracy of detections became heavily dependent on anchor boxes. The anchors were manually selected, and if they were poorly chosen for the dataset, the model's metrics suffered.

➑️ YOLO v3: Victory over other models (2018)

Thanks to the update, YOLO v3 became a foundation in ML. It surpassed Faster R-CNN in popularity and became a favorite of many developers.

What was added new?

πŸ“ Multiscale detection. It removed noise when detecting small objects and stopped ignoring them.

πŸ“ The "third eye". The network immediately outputted three candidates at different resolutions β€” large, smaller, and the smallest.

PROBLEM: The version became slower. Due to the complexity of the architecture, v3 became heavier than its predecessors. The anchors were still manually selected, which also slowed down the detection process.


Model continued to evolve, but no longer in the hands of its original author, Joseph Redmon: he left ML and handed over project to a large company.

In next post, we'll break down why YOLO v4 is called the "engineer's constitution" and YOLO v5 is a "ugly duckling"?.

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Marching Squares is a classic algorithm for constructing contour lines (isolinues) from a 2D scalar field.

Used for visualizations such as topographic maps.

Each grid cell is mapped to a simple polygon depending on which of its angles are above or below a specified threshold.

It has a 3D counterpart, Marching Cubes, which does the same thing, but for 3D.

If you like such visualizations, you might be interested in a recent article about the ML concept of Rectified Flows.

There are many explanatory interactive visualizations there.

Code for all of this can also be viewed on GitHub.

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Many teams are trying to apply DevOps practices to LLM applications.

But DevOps, MLOps and LLMOps solve fundamentally different problems.

🟒 Break it down, DevOps vs MLOps vs LLMOps:

⁠➜ DevOps is focused on software.

You write code, test it, and deploy it. The feedback loop is simple: does the code work or not?

The main artifact is code. Testing is deterministic. The tooling is mature after 15+ years of development.

⁠➜ MLOps is focused on (model + data).

Here you have data drift, model degradation, and constant retraining.

The code might be perfect, but the model quality degrades over time because the world changes.

An anti-fraud model might work great at launch, but start failing after a few weeks because the fraudsters have adapted.

The main artifact expands to code + data + models. All three need to be versioned. That's why MLflow, DVC, and feature stores have become essential tools.

⁠➜ LLMOps is focused on foundation models.

Usually, you don't train models from scratch. Instead, you choose a base model and optimize it in three parallel directions:

* Prompt Engineering
* Context Tuning / RAG
* Fine-tuning

Unlike DevOps and MLOps, these directions run in parallel, not sequentially.

But the Biggest difference of LLMOps is monitoring: it's completely different.

In MLOps, you track data drift, model degradation, and accuracy metrics.

In LLMOps, you track:

⁠☞ Hallucination detection
⁠☞ Bias and toxicity
⁠☞ Token consumption and cost
⁠☞ Human feedback loops

Because the output of LLMs is non-deterministic. You can't just check if it "answered correctly". You need to ensure the answer is safe, grounded, and doesn't burn the budget.

63% of production AI systems catch dangerous hallucinations in the first 90 days.

⁠➜ cost model also flips

In MLOps, the main cost is training (GPU hours during development).

In LLMOps, the main cost is inference (each request consumes tokens).

That's why efficiency of prompts, caching, and routing between models are so important in LLMOps.

The evaluation loop in LLMOps feeds back into all three optimization directions at once. A failed eval might mean you need better prompts, richer context, OR fine-tuning.

That is, it's no longer a linear pipeline.

And another thing: Versioning prompts and RAG pipelines in LLMOps is now first-class, just like versioning data has become mandatory in MLOps.

And the ops layer you choose should match the system you're building.


Why this matters?
88% of ML initiatives struggled to reach production if trying to launch them through traditional DevOps approaches.

And LLMs add challenges that MLOps wasn't even designed for in the first place.

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
PersonaPlex is a smart real-time Model for Voice-Controlled and Role-Based Dialogues

Enables two-way voice communication with character control via text prompts and audio.

Generates natural, low-latency interactions, trained on synthetic and real-world dialogues.

What are the Key Features?
- Support for different voices for natural communication.
- Training on synthetic and real-world data.
- Ability to control the character via text prompts.
- Low latency in interactions.


GitHub

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
Open-source extension for LLM serving engines:

Like a caching layer for large-scale production inference of LLMs.

LMCache implements smart KV cache management by reusing key-value states of already encountered text between GPUs, CPUs, and local disks.

🟒 What it can do?

It can reuse any repetitive text fragments, not just prefixes.

This results in:

⁠☞ 4–10x cost reduction in RAG for models owned by the user
⁠☞ Lower Time-To-First-Token (TTFT)
⁠☞ Higher throughput under load
⁠☞ More efficient work with long-context scenarios

An illustrative example of application: NVIDIA integrated LMCache into its inference project Dynamo.
LMCache allows Dynamo to offload the KV cache to external storage layers and effectively reuse it between requests. This reduces the cost of prefill and frees up GPU memory for active computations.

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
RAG vs CAG, a clear explanation in image.

Merging RAG and CAG. How can an AI engineer use this?

Let's break down how this looks and what additional considerations need to be taken into account.

➑️ Example steps for a CAG + RAG architecture:

πŸ‘‰ DATA PREPARATION

1️⃣ For CAG, we only use sources that change infrequently. In addition to "infrequent changes," it's important to understand which of these sources are most frequently needed for relevant queries. Only after this, we "warm up" the selected data in advance in the KV cache model and cache it in memory. This is done once, and the remaining steps can be repeated many times without recalculating the initial cache.

2️⃣ For RAG, if necessary, we pre-calculate and save vector embeddings in a compatible database so that we can later search for them in step 4. Sometimes, simpler types of data are sufficient for RAG, in which case a regular database would work.

πŸ‘‰ QUERY PATH

Now we can use the prepared data.

3️⃣ We assemble the prompt: the user's query + a system prompt with instructions on how the model should use the cached context and external (retrieved) context.

4️⃣ We build an embedding of the user's query for semantic search through the vector DB and query the context store to retrieve relevant data. If semantics aren't needed, we can go to other sources, such as a real-time database or the web.

5️⃣ We enrich the final prompt with the external context retrieved in step 4.

6️⃣ We return the final response to the user.

πŸ‘‰ A FEW IMPORTANT POINTS

☞ The context window is not infinite. Even if the model has a huge context, the "needle in a haystack" problem still exists. Use context sparingly and cache only what is really needed.

☞ For some cases, certain datasets are super valuable to constantly feed them to the model through the cache. For example, an assistant who is obliged to always comply with a long set of internal rules scattered across several documents.

☞ Although open-source CAG has gained popularity relatively recently, in practice, this has long been possible through prompt caching in the OpenAI and Anthropic APIs. It's easy to quickly put together a prototype there.

☞ Always separate hot and cold sources. We only put cold (infrequently changing) data in the cache, otherwise the data will become outdated, and the application will start living "out of reality."

⚠️ RISKS & LIMITATIONS

☞ Be very careful with what you cache, because this data becomes available to all users' queries.

☞ It's difficult to ensure RBAC for the cache if you don't have a separate model with its own cache for each role.


Have you already tried this combo approach?

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
Tencent is making a strong entry into the context learning field.

Open-source benchmark CL-bench has been released - and this isn't just another dataset, but an attempt to shift the focus of the entire industry.

🟒 What they done?

Tencent HY, in collaboration with Fudan University, have published a new work:
β€œCL-bench: A Benchmark for Context Learning” - a systematic benchmark for evaluating whether *models are actually able to think in context*, rather than just recalling what they've learned.

This is the first research release from Vinces Yao's team since his move to Tencent - and it's clear from their ambitions that they're aiming for fundamental changes.

Today, most LLMs operate according to the following scheme:
huge weights + memorized patterns = answers

But the real world isn't a memory test. It's about:

- long, complex contexts
- conflicting information
- the need to change strategies on the fly
- drawing conclusions based on what's just appeared

Models need to move from static memorization to dynamic reasoning within context.

CL-bench precisely tests this breaking point:

- how the model uses context, not just weights
- whether it can update its understanding
- whether it's capable of reasoning in complex scenarios, not just on pure QA tasks

In essence, this is a step towards models that are closer to agents than to "smart autocomplete".

Plus a strategic signal

At the same time, Tencent is launching Tencent HY Research - a blog where frontier research will be published.

This looks like a declaration:
"We're not just training large models. We want to influence how they're evaluated at all."

And this is already a level of influence on the direction of the entire field.
CL-bench isn't about +0.5% on the leaderboard.
It's about a paradigm shift:

The LLMs of the future = less rote learning, more thinking in real-world contexts.

And if this line succeeds, it's precisely such benchmarks that will determine who has truly created a "smart" model, and who has just inflated the parameters.


Project | Blog

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
It's possible to build an LLM from scratch.

There's a repository that breaks down the complex mathematics of Transformers into understandable, clean Python. It covers the entire lifecycle of an LLM.

β†’ step-by-step implementation
β†’ simple, "hackable" examples

100% open source.

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
Guardrails is no longer a "last thought" or a bonus to the project - these are key architectural patterns that determine whether you can safely deploy your agent system.

Below are four working patterns we have seen in production systems:

1️⃣ Adaptive feedback loops
Agents perform tasks β†’ Supervisor evaluates β†’ Reward service updates policies β†’ Guidelines are adjusted β†’ Agents improve over time.
A continuous learning cycle is created, where the system reinforces effective behavior and reduces risky behavior. This is reward-based learning, which improves with each iteration.

2️⃣ Corrective action
A centralized Supervisor distributes tasks, compares results with the application's guidelines, and connects alternative agents if errors are detected. The best verified result is returned to the user. This prevents a bad result from reaching end users.

3️⃣ Human in the loop
For sensitive domains (medicine, law, finance), agents generate preliminary responses, but a human validates them before execution. The flow is automatically paused for expert review and resumes only after approval.

4️⃣ Emergency stop
Critical for high-risk systems, such as trading.
Agent 1 collects market data β†’ LLM processes signals β†’ Agent 2 evaluates conditions β†’ if anomalies or risks are detected, execution is immediately stopped.
Example: a trading bot with access to a volatility API showing VIX = 42 (extreme market stress). Even if the bot suggests an aggressive trade, the evaluator independently checks: "Is this adequate given the current volatility?" If not - the action is completely blocked.

The underlying philosophy is behavior shaping: a three-step loop of evaluation β†’ feedback β†’ correction. The evaluator doesn't just record the result post-factum. He actively intervenes: rolls back bad transactions, stops flows with incorrect data, or redirects complex cases for human verification.

It's especially important when agents interact with unstable external states - market conditions, API health, system load. The evaluator provides a sanity-check to ensure that the model correctly interpreted the signals and didn't just generate coherent text.


The goal is not to catch all errors in advance (this is impossible). The goal is to build systems that detect problems on the fly, understand what went wrong, and automatically correct the course before the damage spreads.

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
Every second tutorial on RAG is either a toy or a research project disguised as a product.

That's not true. Agentic RAG, made properly:

β†’ hierarchical search (first child elements, parent - on request)
β†’ dialogue memory
β†’ refinement of requests
β†’ parallel agents


GitHub

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security
HOW YOLO BECAME A STANDARD IN CV? Launching a series of posts about the evolution of one of the most popular architectures in computer vision. We'll break down: Before 2015, the task of detection was solved by searching for the most likely regions. There…
Continue the series of posts on evolution of most popular model family for Object Detection.

Development ceased to be purely conceptual and became more engineering-oriented:

🟒 Review of YOLO v4-v6
➑️ YOLO v4: Model turned into an engineering encyclopedia (2020)

YOLO v4 became a "BIBLE" for improving architectures. It packed in as many tricks as possible without killing FPS.

GOLDEN FEATURE: In new version, mosaic augmentation was introduced. It collects training picture from several different ones, which improves model's performance. As a result, the quality was improved by +6% mAP compared to YOLOv3, while maintaining a speed of 60 FPS.

OTHER CHANGES: Pyramidal architecture (CSPDarknet-53 + PANet + SPP). Instead of simply cutting out pieces from picture, a multi-scale approach was implemented at level of network itself. Network itself extracted features of different scales and recognized contexts.

TRICKS & AUGMENTATIONS. Architecture integrated such developments as Mish-activation, DropBlock, and CloU loss. Together with mosaic augmentation, they improved model's quality by 10% without drastically changing it.

DOWNSIDES of YOLO v4 include difficulty of integrating model and manual hyperparameters left over from previous versions.

There are no more problems to fix, so developers focused on improvements.

➑️ YOLO v5: "Ugly Duckling" and mass adoption (2020-2021)

YOLO v5 was released four months after v4 - version was nicknamed "Ugly Duckling", because there were no architectural breakthroughs in it.

GOLDEN FEATURE: YOLO v5 was rewritten in PyTorch and made it more user-friendly. Everyone could integrate it into their project and retrain it for their own tasks. PyTorch soon gained popularity and dominated the DL field, which led to mass adoption of YOLO.

There weren't many other features - they were released to promote article about the new version. But there were a lot of problems:

πŸ“Ž Version didn't work due to bugs. For first two months, the buggy implementation simply didn't allow to use model. Memory leaks, incorrectly specified areas for three candidates.

πŸ“Ž Version didn't add anything new. Each new YOLO either solved an engineering problem or an idea problem. Fifth model was considered a rewrite of what already existed - just on a different framework. Community didn't like this approach.

πŸ“Ž Version was developed by Ultralytics. Community was wary of it: previously, YOLO was developed by a CIS superstar in the CV field - Bachkovsky and now it's some no-names. So developers were worried about fate of beloved model.

πŸ“Ž Version never got an article. Company promised to release it within a few months. But it's been four years - Article hasn't appeared. They just released a couple of technical reports on archive.

Fortunately, Ultralytics didn't abandon model and kept improving and enhancing it. Thanks to PyTorch and support from developers, YOLO v5 is widely used as a component of a comprehensive solution.

➑️ YOLO v6: Model was made more convenient for deployment (2022)

Company focused on developing most convenient real-time deployment for frameworks like TensorRT and Edge devices.

GOLDEN FEATURE: An Anchor-Free Head was introduced. Instead of predicting shifts for candidates, Model searches for exact center of object. It's faster and more accurate.

OTHER INNOVATIONS: New architecture. EfficientRep, an analogue of EfficientNet, was chosen as the backbone. They also abandoned DarkNet backbone - it was outdated.

HIGH SPEED. Model became super-lightweight and demonstrated 120 FPS on a T4 at a resolution of 640x640. Therefore, it was used in tasks related to thermal imagers and Edge computing.

There were no obvious downsides or problems with the model. Except for the accuracy compared to v5 and v7. But v6 is best for Edge devices.


In next post, we'll discuss at Why YOLO v8 became the most popular model in the family? and
How commercialization turned the project into a conveyor?

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML |
@DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
ACE-Step v1.5: Ace Studio in collaboration with StepFun have updated the ACE-Step, local music generator to version 1.5.

The entry threshold has been lowered to a minimum: the junior model requires less than 6 GB of video memory, and, depending on the think mode settings, generation can take from 2 to 10 seconds - this is already the level of commercial solutions.

🟒 What they did?

The developers have assembled a hybrid of a language model that turns a prompt into a composition sketch: it outlines the structure, comes up with lyrics and metadata, and DiT, which is responsible for the sound. The logical core of this entire system is based on Qwen3.

ACE-Step v1.5 can generate tracks from 10 seconds to 10 minutes long, with up to 8 tracks at a time. There are more than 1000 instruments in the database, and the system understands lyrics in 50 languages.

The authors have prepared a whole set of models for different amounts of VRAM:

⁠➜ Less than 6 GB: without the LM module, only the sound engine works.

⁠➜ 6-12 GB: a lightweight version of LM (0.6B).

⁠➜ 16 GB and above: a full-fledged model with 4 billion parameters, which best understands the context and delivers maximum quality
.
When launched, ACE-Step v1.5 automatically selects a model and parameters suitable for the hardware. Detailed information on configurations can be found here.

ACE-Step can do much more than just turn text into a melody. You can give it an audio example to copy the style, make covers, correct parts of already finished tracks, or generate an accompaniment for vocals.

The most interesting feature is the ability to create LoRA. To feed the model with your own style, just 8 tracks are enough. On the 30th series RTX with 12 GB of memory, this process will take about an hour.

Everything is in order with the deployment, the developers have prepared a portable build, and for ComfyUI they have already written all the necessary nodes and workflows.

Project page, Model, Paper, Demo, Discord community, GitHub β€’ #AI #ML #Text2Music #AceStudio #StepFun

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore
Dude completely implemented the architecture of GPT-OSS-20B from scratch in PyTorch. All components were written from scratch:

⁠☞ RoPE with YaRN + NTK-by-parts for context scaling
⁠☞ RMSNorm
⁠☞ SwiGLU with clamping and residual connections
⁠☞ Mixture-of-Experts (MoE)
⁠☞ Self-Attention, optimized via Grouped Query Attention (GQA)
⁠☞ Learned sinks
⁠☞ Banded (sliding window) attention
⁠☞ Support for KV caching

All of this works on a single A100 SXM (80GB). He also wrote detailed documentation with the theory of each component, as well as instructions for setup and inference.

Repository

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML |
@DataXplore
Step 3.5 Flash: a model with a hybrid attention architecture and a speed of up to 350 T/s.

a very interesting MoE model with 196 billion total and 11 active parameters.

Authors claim an insane speed of up to 300 tokens per second, and on tasks with code, it supposedly accelerates to 350. For a model of this level, this is very impressive.

🟒 What's going on inside?

Instead of the standard attention mechanism, they used a hybrid scheme: one layer of full attention on 3 sliding window layers, which allowed them to cram a context of 256 thousand tokens into the model without clogging up the memory to the point of failure.

In training, they used the MIS-PO algorithm, which helped solve the problem of losing the thread in long CoTs and simply cuts off options that deviate too much from logic.

The model, as is now fashionable, was tailored for autonomous agents. It can use ten tools at the same time. In Deep Research mode, the model itself googles, plans stages, and writes reports up to 10 thousand words long.

If you need to run a heavy code repository through the model, it handles it without the usual slowdowns that occur when working with voluminous texts.

⁠➜ BENCHMARKS

Step 3.5 Flash scored 97.3 on the AIME 2025 test (and this is bare risoning, without third-party calculators). If you give it access to Python, the result soars to 99.8.

On code benchmarks, the numbers also look impressive: on SWE-bench, it gives 74.4%, and on Terminal-Bench 2.0 - 51.0%.

Of course, in terms of knowledge density, Step 3.5 Flash still lags behind Gemini 3.0 Pro, but the fact that it is available for local use and tests via API, is pleasing.


Article, Model, Demo, Discord Community, GitHub β€’ #AI #ML #LLM #StepFunAI

β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’β€’
πŸ€– Data & ML | @DataXplore