This media is not supported in your browser
VIEW IN TELEGRAM
Boosted Up performance of AI agent by 184% using a completely open-source technique.
Now, you can automatically find best prompts for any agentic workflow you're putting together, means manual prompt engineering isn't needed at all.
π’ What is the simple idea?
Best Part: Everything can be run completely locally, because you can use any local LLMs as optimizers and evaluators.
GitHub repository
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Now, you can automatically find best prompts for any agentic workflow you're putting together, means manual prompt engineering isn't needed at all.
π’ What is the simple idea?
1οΈβ£ Take a starting prompt and an eval dataset
2οΈβ£ Then, an optimizer iteratively improves the prompt
3οΈβ£ In the end, you get an optimal prompt automatically
And all of this in just a few lines of code.
β‘οΈ Why Opik specifically?
Opik is a 100% open-source platform for evaluating LLMs.
It helps optimize LLM systems so that they work better, faster, and cheaper: from RAG chatbots to code assistants. Opik includes tracing, evaluations, and dashboards.
Best Part: Everything can be run completely locally, because you can use any local LLMs as optimizers and evaluators.
GitHub repository
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
π DeepSeek-OCR 2 is a new generation of OCR with SOTA quality
A 3B model for advanced understanding of images, documents and OCR, which reaches the SOTA level.
π’ Why this matters?
Now, DeepSeek-OCR 2 can be conveniently launched and fine-tuned via Unsloth according to the ready-made guide.
Guide, Model, Github, Paper | #DeepSeek #ocr #opensource
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
A 3B model for advanced understanding of images, documents and OCR, which reaches the SOTA level.
π’ Why this matters?
The key novelty is DeepEncoder V2.
Unlike classic vision LLMs, which "read" the image as a grid (left-to-right, top-to-bottom), DeepEncoder V2 works closer to how a human reads:
- First, a global understanding of the image is formed
- Then, the model determines the logical order of reading - what is important first, what next
What this brings in practice
π Works better with complex document layouts
π Correctly reads tables
π§Ύ Links signatures and values
π° Understands columns and structured text
π More reliably processes a mixture of text and visual structure
In terms of quality
- Outperforms Gemini 3 Pro on a number of benchmarks
- Gives >4% improvement compared to the previous version of DeepSeek-OCR
And this is with a model size of just 3B parameters.
Can be launched and fine-tuned
Now, DeepSeek-OCR 2 can be conveniently launched and fine-tuned via Unsloth according to the ready-made guide.
Guide, Model, Github, Paper | #DeepSeek #ocr #opensource
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
To train an ML model, you need to be proficient in algorithms, write code, endlessly tune hyperparameters which is a high entry barrier for most people.
An open-source project Plexe significantly lowers this threshold: You describe the task in plain language and it automatically assembles machine learning for it.
π’ How it works?
GitHub
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
An open-source project Plexe significantly lowers this threshold: You describe the task in plain language and it automatically assembles machine learning for it.
π’ How it works?
β Explain in a human-friendly way what you want to predict, what the input data is and what the output should be.
β β Next the system, through a combination of several agents, goes through the entire pipeline: data analysis, solution plan, code generation, tests, and quality assessment.
β β Supports various LLM providers: OpenAI, Anthropic, Ollama, and others. Plus, it can automatically derive the data structure or even generate a synthetic dataset.
β β There's also distributed training on Ray inside: you can run multiple model variants in parallel and significantly speed up the process.
GitHub
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
HOW YOLO BECAME A STANDARD IN CV?
Launching a series of posts about the evolution of one of the most popular architectures in computer vision.
We'll break down:
Before 2015, the task of detection was solved by searching for the most likely regions. There were two-stage approaches, such as Faster R-CNN.
π’ How did the YOLO architecture evolve from v1 to v3?
Model continued to evolve, but no longer in the hands of its original author, Joseph Redmon: he left ML and handed over project to a large company.
In next post, we'll break down why YOLO v4 is called the "engineer's constitution" and YOLO v5 is a "ugly duckling"?.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Launching a series of posts about the evolution of one of the most popular architectures in computer vision.
We'll break down:
Before 2015, the task of detection was solved by searching for the most likely regions. There were two-stage approaches, such as Faster R-CNN.
π’ How did the YOLO architecture evolve from v1 to v3?
π First, they searched for candidate regions, and then used a refine process to refine the classes and coordinates.
PROBLEM: The process was very slow. Imagine the task of tracking a tennis ball on the court during a match. Old networks would have taken 5 minutes to process a video, even on a good GPU. Players would have had to stand and wait for the VAR system.
A real-time approach was needed, where speed was more important than perfect results. Thus, YOLO was born.
β‘οΈ YOLO v1: A model that looks at the entire scene (2015)
The idea was to turn detection from a region-searching task into a regression problem. Combine all stages into a single network that directly "spits out" coordinates.
How it was implemented technically?
π They made an architecture similar to GoogLeNet. Two fully connected and 24 convolutional layers. Although it was large, it detected bounding boxes and immediately determined the coordinates.
π All images were divided into a 7x7 grid. Each cell predicted 2 bounding boxes and 20 classes. The input was a 448x448 image, which was further divided into 64x64.
PROBLEM: YOLO v1 couldn't handle other resolutions. To work with detection on large images, they resized them to 448x448 or cut them into patches. Due to the extra operations, the main advantage over Faster R-CNN β speed β was lost.
β‘οΈ YOLO v2 / YOLO9000: Scale and anchors (2016β2017)
To level the complex LOSS, multi-scale was added to the new version: YOLO9000 simultaneously detects more than 9,000 classes without full annotation β hence the name.
What new features were added?
π Anchor Boxes: Instead of directly predicting coordinates, they switched to predicting shifts relative to the X and Y axes for candidates. This maximized object capture.
π Skip Connections: They introduced pass-through layers and added batch normalization, which solved the problem of gradient fading.
PROBLEM: The accuracy of detections became heavily dependent on anchor boxes. The anchors were manually selected, and if they were poorly chosen for the dataset, the model's metrics suffered.
β‘οΈ YOLO v3: Victory over other models (2018)
Thanks to the update, YOLO v3 became a foundation in ML. It surpassed Faster R-CNN in popularity and became a favorite of many developers.
What was added new?
π Multiscale detection. It removed noise when detecting small objects and stopped ignoring them.
π The "third eye". The network immediately outputted three candidates at different resolutions β large, smaller, and the smallest.
PROBLEM: The version became slower. Due to the complexity of the architecture, v3 became heavier than its predecessors. The anchors were still manually selected, which also slowed down the detection process.
Model continued to evolve, but no longer in the hands of its original author, Joseph Redmon: he left ML and handed over project to a large company.
In next post, we'll break down why YOLO v4 is called the "engineer's constitution" and YOLO v5 is a "ugly duckling"?.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Marching Squares is a classic algorithm for constructing contour lines (isolinues) from a 2D scalar field.
Code for all of this can also be viewed on GitHub.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Used for visualizations such as topographic maps.
Each grid cell is mapped to a simple polygon depending on which of its angles are above or below a specified threshold.
It has a 3D counterpart, Marching Cubes, which does the same thing, but for 3D.
If you like such visualizations, you might be interested in a recent article about the ML concept of Rectified Flows.
There are many explanatory interactive visualizations there.
Code for all of this can also be viewed on GitHub.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
Many teams are trying to apply DevOps practices to LLM applications.
But DevOps, MLOps and LLMOps solve fundamentally different problems.
π’ Break it down, DevOps vs MLOps vs LLMOps:
Why this matters?
88% of ML initiatives struggled to reach production if trying to launch them through traditional DevOps approaches.
And LLMs add challenges that MLOps wasn't even designed for in the first place.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
But DevOps, MLOps and LLMOps solve fundamentally different problems.
π’ Break it down, DevOps vs MLOps vs LLMOps:
β β DevOps is focused on software.
You write code, test it, and deploy it. The feedback loop is simple: does the code work or not?
The main artifact is code. Testing is deterministic. The tooling is mature after 15+ years of development.
β β MLOps is focused on (model + data).
Here you have data drift, model degradation, and constant retraining.
The code might be perfect, but the model quality degrades over time because the world changes.
An anti-fraud model might work great at launch, but start failing after a few weeks because the fraudsters have adapted.
The main artifact expands to code + data + models. All three need to be versioned. That's why MLflow, DVC, and feature stores have become essential tools.
β β LLMOps is focused on foundation models.
Usually, you don't train models from scratch. Instead, you choose a base model and optimize it in three parallel directions:
* Prompt Engineering
* Context Tuning / RAG
* Fine-tuning
Unlike DevOps and MLOps, these directions run in parallel, not sequentially.
But the Biggest difference of LLMOps is monitoring: it's completely different.
In MLOps, you track data drift, model degradation, and accuracy metrics.
In LLMOps, you track:
β β Hallucination detection
β β Bias and toxicity
β β Token consumption and cost
β β Human feedback loops
Because the output of LLMs is non-deterministic. You can't just check if it "answered correctly". You need to ensure the answer is safe, grounded, and doesn't burn the budget.
63% of production AI systems catch dangerous hallucinations in the first 90 days.
β β cost model also flips
In MLOps, the main cost is training (GPU hours during development).
In LLMOps, the main cost is inference (each request consumes tokens).
That's why efficiency of prompts, caching, and routing between models are so important in LLMOps.
The evaluation loop in LLMOps feeds back into all three optimization directions at once. A failed eval might mean you need better prompts, richer context, OR fine-tuning.
That is, it's no longer a linear pipeline.
And another thing: Versioning prompts and RAG pipelines in LLMOps is now first-class, just like versioning data has become mandatory in MLOps.
And the ops layer you choose should match the system you're building.
Why this matters?
88% of ML initiatives struggled to reach production if trying to launch them through traditional DevOps approaches.
And LLMs add challenges that MLOps wasn't even designed for in the first place.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
PersonaPlex is a smart real-time Model for Voice-Controlled and Role-Based Dialogues
Enables two-way voice communication with character control via text prompts and audio.
Generates natural, low-latency interactions, trained on synthetic and real-world dialogues.
What are the Key Features?
GitHub
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Enables two-way voice communication with character control via text prompts and audio.
Generates natural, low-latency interactions, trained on synthetic and real-world dialogues.
What are the Key Features?
- Support for different voices for natural communication.
- Training on synthetic and real-world data.
- Ability to control the character via text prompts.
- Low latency in interactions.
GitHub
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Open-source extension for LLM serving engines:
Like a caching layer for large-scale production inference of LLMs.
LMCache implements smart KV cache management by reusing key-value states of already encountered text between GPUs, CPUs, and local disks.
π’ What it can do?
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Like a caching layer for large-scale production inference of LLMs.
LMCache implements smart KV cache management by reusing key-value states of already encountered text between GPUs, CPUs, and local disks.
π’ What it can do?
It can reuse any repetitive text fragments, not just prefixes.
This results in:
β β 4β10x cost reduction in RAG for models owned by the user
β β Lower Time-To-First-Token (TTFT)
β β Higher throughput under load
β β More efficient work with long-context scenarios
An illustrative example of application: NVIDIA integrated LMCache into its inference project Dynamo.
LMCache allows Dynamo to offload the KV cache to external storage layers and effectively reuse it between requests. This reduces the cost of prefill and frees up GPU memory for active computations.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
RAG vs CAG, a clear explanation in image.
Merging RAG and CAG. How can an AI engineer use this?
Let's break down how this looks and what additional considerations need to be taken into account.
β‘οΈ Example steps for a CAG + RAG architecture:
Have you already tried this combo approach?
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Merging RAG and CAG. How can an AI engineer use this?
Let's break down how this looks and what additional considerations need to be taken into account.
β‘οΈ Example steps for a CAG + RAG architecture:
π DATA PREPARATION
1οΈβ£ For CAG, we only use sources that change infrequently. In addition to "infrequent changes," it's important to understand which of these sources are most frequently needed for relevant queries. Only after this, we "warm up" the selected data in advance in the KV cache model and cache it in memory. This is done once, and the remaining steps can be repeated many times without recalculating the initial cache.
2οΈβ£ For RAG, if necessary, we pre-calculate and save vector embeddings in a compatible database so that we can later search for them in step 4. Sometimes, simpler types of data are sufficient for RAG, in which case a regular database would work.
π QUERY PATH
Now we can use the prepared data.
3οΈβ£ We assemble the prompt: the user's query + a system prompt with instructions on how the model should use the cached context and external (retrieved) context.
4οΈβ£ We build an embedding of the user's query for semantic search through the vector DB and query the context store to retrieve relevant data. If semantics aren't needed, we can go to other sources, such as a real-time database or the web.
5οΈβ£ We enrich the final prompt with the external context retrieved in step 4.
6οΈβ£ We return the final response to the user.
π A FEW IMPORTANT POINTS
β The context window is not infinite. Even if the model has a huge context, the "needle in a haystack" problem still exists. Use context sparingly and cache only what is really needed.
β For some cases, certain datasets are super valuable to constantly feed them to the model through the cache. For example, an assistant who is obliged to always comply with a long set of internal rules scattered across several documents.
β Although open-source CAG has gained popularity relatively recently, in practice, this has long been possible through prompt caching in the OpenAI and Anthropic APIs. It's easy to quickly put together a prototype there.
β Always separate hot and cold sources. We only put cold (infrequently changing) data in the cache, otherwise the data will become outdated, and the application will start living "out of reality."
β οΈ RISKS & LIMITATIONS
β Be very careful with what you cache, because this data becomes available to all users' queries.
β It's difficult to ensure RBAC for the cache if you don't have a separate model with its own cache for each role.
Have you already tried this combo approach?
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Tencent is making a strong entry into the context learning field.
Open-source benchmark CL-bench has been released - and this isn't just another dataset, but an attempt to shift the focus of the entire industry.
π’ What they done?
Project | Blog
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Open-source benchmark CL-bench has been released - and this isn't just another dataset, but an attempt to shift the focus of the entire industry.
π’ What they done?
Tencent HY, in collaboration with Fudan University, have published a new work:
βCL-bench: A Benchmark for Context Learningβ - a systematic benchmark for evaluating whether *models are actually able to think in context*, rather than just recalling what they've learned.
This is the first research release from Vinces Yao's team since his move to Tencent - and it's clear from their ambitions that they're aiming for fundamental changes.
Today, most LLMs operate according to the following scheme:
huge weights + memorized patterns = answers
But the real world isn't a memory test. It's about:
- long, complex contexts
- conflicting information
- the need to change strategies on the fly
- drawing conclusions based on what's just appeared
Models need to move from static memorization to dynamic reasoning within context.
CL-bench precisely tests this breaking point:
- how the model uses context, not just weights
- whether it can update its understanding
- whether it's capable of reasoning in complex scenarios, not just on pure QA tasks
In essence, this is a step towards models that are closer to agents than to "smart autocomplete".
Plus a strategic signal
At the same time, Tencent is launching Tencent HY Research - a blog where frontier research will be published.
This looks like a declaration:
"We're not just training large models. We want to influence how they're evaluated at all."
And this is already a level of influence on the direction of the entire field.
CL-bench isn't about +0.5% on the leaderboard.
It's about a paradigm shift:
The LLMs of the future = less rote learning, more thinking in real-world contexts.
And if this line succeeds, it's precisely such benchmarks that will determine who has truly created a "smart" model, and who has just inflated the parameters.
Project | Blog
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
It's possible to build an LLM from scratch.
There's a repository that breaks down the complex mathematics of Transformers into understandable, clean Python. It covers the entire lifecycle of an LLM.
β step-by-step implementation
β simple, "hackable" examples
100% open source.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
There's a repository that breaks down the complex mathematics of Transformers into understandable, clean Python. It covers the entire lifecycle of an LLM.
β step-by-step implementation
β simple, "hackable" examples
100% open source.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Guardrails is no longer a "last thought" or a bonus to the project - these are key architectural patterns that determine whether you can safely deploy your agent system.
Below are four working patterns we have seen in production systems:
The goal is not to catch all errors in advance (this is impossible). The goal is to build systems that detect problems on the fly, understand what went wrong, and automatically correct the course before the damage spreads.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Below are four working patterns we have seen in production systems:
1οΈβ£ Adaptive feedback loops
Agents perform tasks β Supervisor evaluates β Reward service updates policies β Guidelines are adjusted β Agents improve over time.
A continuous learning cycle is created, where the system reinforces effective behavior and reduces risky behavior. This is reward-based learning, which improves with each iteration.
2οΈβ£ Corrective action
A centralized Supervisor distributes tasks, compares results with the application's guidelines, and connects alternative agents if errors are detected. The best verified result is returned to the user. This prevents a bad result from reaching end users.
3οΈβ£ Human in the loop
For sensitive domains (medicine, law, finance), agents generate preliminary responses, but a human validates them before execution. The flow is automatically paused for expert review and resumes only after approval.
4οΈβ£ Emergency stop
Critical for high-risk systems, such as trading.
Agent 1 collects market data β LLM processes signals β Agent 2 evaluates conditions β if anomalies or risks are detected, execution is immediately stopped.
Example: a trading bot with access to a volatility API showing VIX = 42 (extreme market stress). Even if the bot suggests an aggressive trade, the evaluator independently checks: "Is this adequate given the current volatility?" If not - the action is completely blocked.
The underlying philosophy is behavior shaping: a three-step loop of evaluation β feedback β correction. The evaluator doesn't just record the result post-factum. He actively intervenes: rolls back bad transactions, stops flows with incorrect data, or redirects complex cases for human verification.
It's especially important when agents interact with unstable external states - market conditions, API health, system load. The evaluator provides a sanity-check to ensure that the model correctly interpreted the signals and didn't just generate coherent text.
The goal is not to catch all errors in advance (this is impossible). The goal is to build systems that detect problems on the fly, understand what went wrong, and automatically correct the course before the damage spreads.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Every second tutorial on RAG is either a toy or a research project disguised as a product.
That's not true. Agentic RAG, made properly:
GitHub
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
That's not true. Agentic RAG, made properly:
β hierarchical search (first child elements, parent - on request)
β dialogue memory
β refinement of requests
β parallel agents
GitHub
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Data eXplore : Data Science, ML, Big Data, LLMs and AI Security
HOW YOLO BECAME A STANDARD IN CV? Launching a series of posts about the evolution of one of the most popular architectures in computer vision. We'll break down: Before 2015, the task of detection was solved by searching for the most likely regions. Thereβ¦
Continue the series of posts on evolution of most popular model family for Object Detection.
Development ceased to be purely conceptual and became more engineering-oriented:
π’ Review of YOLO v4-v6
In next post, we'll discuss at Why YOLO v8 became the most popular model in the family? and
How commercialization turned the project into a conveyor?
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Development ceased to be purely conceptual and became more engineering-oriented:
π’ Review of YOLO v4-v6
β‘οΈ YOLO v4: Model turned into an engineering encyclopedia (2020)
YOLO v4 became a "BIBLE" for improving architectures. It packed in as many tricks as possible without killing FPS.
GOLDEN FEATURE: In new version, mosaic augmentation was introduced. It collects training picture from several different ones, which improves model's performance. As a result, the quality was improved by +6% mAP compared to YOLOv3, while maintaining a speed of 60 FPS.
OTHER CHANGES: Pyramidal architecture (CSPDarknet-53 + PANet + SPP). Instead of simply cutting out pieces from picture, a multi-scale approach was implemented at level of network itself. Network itself extracted features of different scales and recognized contexts.
TRICKS & AUGMENTATIONS. Architecture integrated such developments as Mish-activation, DropBlock, and CloU loss. Together with mosaic augmentation, they improved model's quality by 10% without drastically changing it.
DOWNSIDES of YOLO v4 include difficulty of integrating model and manual hyperparameters left over from previous versions.
There are no more problems to fix, so developers focused on improvements.
β‘οΈ YOLO v5: "Ugly Duckling" and mass adoption (2020-2021)
YOLO v5 was released four months after v4 - version was nicknamed "Ugly Duckling", because there were no architectural breakthroughs in it.
GOLDEN FEATURE: YOLO v5 was rewritten in PyTorch and made it more user-friendly. Everyone could integrate it into their project and retrain it for their own tasks. PyTorch soon gained popularity and dominated the DL field, which led to mass adoption of YOLO.
There weren't many other features - they were released to promote article about the new version. But there were a lot of problems:
π Version didn't work due to bugs. For first two months, the buggy implementation simply didn't allow to use model. Memory leaks, incorrectly specified areas for three candidates.
π Version didn't add anything new. Each new YOLO either solved an engineering problem or an idea problem. Fifth model was considered a rewrite of what already existed - just on a different framework. Community didn't like this approach.
π Version was developed by Ultralytics. Community was wary of it: previously, YOLO was developed by a CIS superstar in the CV field - Bachkovsky and now it's some no-names. So developers were worried about fate of beloved model.
π Version never got an article. Company promised to release it within a few months. But it's been four years - Article hasn't appeared. They just released a couple of technical reports on archive.
Fortunately, Ultralytics didn't abandon model and kept improving and enhancing it. Thanks to PyTorch and support from developers, YOLO v5 is widely used as a component of a comprehensive solution.
β‘οΈ YOLO v6: Model was made more convenient for deployment (2022)
Company focused on developing most convenient real-time deployment for frameworks like TensorRT and Edge devices.
GOLDEN FEATURE: An Anchor-Free Head was introduced. Instead of predicting shifts for candidates, Model searches for exact center of object. It's faster and more accurate.
OTHER INNOVATIONS: New architecture. EfficientRep, an analogue of EfficientNet, was chosen as the backbone. They also abandoned DarkNet backbone - it was outdated.
HIGH SPEED. Model became super-lightweight and demonstrated 120 FPS on a T4 at a resolution of 640x640. Therefore, it was used in tasks related to thermal imagers and Edge computing.
There were no obvious downsides or problems with the model. Except for the accuracy compared to v5 and v7. But v6 is best for Edge devices.
In next post, we'll discuss at Why YOLO v8 became the most popular model in the family? and
How commercialization turned the project into a conveyor?
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
ACE-Step v1.5: Ace Studio in collaboration with StepFun have updated the ACE-Step, local music generator to version 1.5.
The entry threshold has been lowered to a minimum: the junior model requires less than 6 GB of video memory, and, depending on the think mode settings, generation can take from 2 to 10 seconds - this is already the level of commercial solutions.
π’ What they did?
Project page, Model, Paper, Demo, Discord community, GitHub β’ #AI #ML #Text2Music #AceStudio #StepFun
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
The entry threshold has been lowered to a minimum: the junior model requires less than 6 GB of video memory, and, depending on the think mode settings, generation can take from 2 to 10 seconds - this is already the level of commercial solutions.
π’ What they did?
The developers have assembled a hybrid of a language model that turns a prompt into a composition sketch: it outlines the structure, comes up with lyrics and metadata, and DiT, which is responsible for the sound. The logical core of this entire system is based on Qwen3.
ACE-Step v1.5 can generate tracks from 10 seconds to 10 minutes long, with up to 8 tracks at a time. There are more than 1000 instruments in the database, and the system understands lyrics in 50 languages.
The authors have prepared a whole set of models for different amounts of VRAM:
β β Less than 6 GB: without the LM module, only the sound engine works.
β β 6-12 GB: a lightweight version of LM (0.6B).
β β 16 GB and above: a full-fledged model with 4 billion parameters, which best understands the context and delivers maximum quality
.
When launched, ACE-Step v1.5 automatically selects a model and parameters suitable for the hardware. Detailed information on configurations can be found here.
ACE-Step can do much more than just turn text into a melody. You can give it an audio example to copy the style, make covers, correct parts of already finished tracks, or generate an accompaniment for vocals.
The most interesting feature is the ability to create LoRA. To feed the model with your own style, just 8 tracks are enough. On the 30th series RTX with 12 GB of memory, this process will take about an hour.
Everything is in order with the deployment, the developers have prepared a portable build, and for ComfyUI they have already written all the necessary nodes and workflows.
Project page, Model, Paper, Demo, Discord community, GitHub β’ #AI #ML #Text2Music #AceStudio #StepFun
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Dude completely implemented the architecture of GPT-OSS-20B from scratch in PyTorch. All components were written from scratch:
All of this works on a single A100 SXM (80GB). He also wrote detailed documentation with the theory of each component, as well as instructions for setup and inference.
Repository
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
β β RoPE with YaRN + NTK-by-parts for context scaling
β β RMSNorm
β β SwiGLU with clamping and residual connections
β β Mixture-of-Experts (MoE)
β β Self-Attention, optimized via Grouped Query Attention (GQA)
β β Learned sinks
β β Banded (sliding window) attention
β β Support for KV caching
All of this works on a single A100 SXM (80GB). He also wrote detailed documentation with the theory of each component, as well as instructions for setup and inference.
Repository
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Step 3.5 Flash: a model with a hybrid attention architecture and a speed of up to 350 T/s.
a very interesting MoE model with 196 billion total and 11 active parameters.
Authors claim an insane speed of up to 300 tokens per second, and on tasks with code, it supposedly accelerates to 350. For a model of this level, this is very impressive.
π’ What's going on inside?
Article, Model, Demo, Discord Community, GitHub β’ #AI #ML #LLM #StepFunAI
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
a very interesting MoE model with 196 billion total and 11 active parameters.
Authors claim an insane speed of up to 300 tokens per second, and on tasks with code, it supposedly accelerates to 350. For a model of this level, this is very impressive.
π’ What's going on inside?
Instead of the standard attention mechanism, they used a hybrid scheme: one layer of full attention on 3 sliding window layers, which allowed them to cram a context of 256 thousand tokens into the model without clogging up the memory to the point of failure.
In training, they used the MIS-PO algorithm, which helped solve the problem of losing the thread in long CoTs and simply cuts off options that deviate too much from logic.
The model, as is now fashionable, was tailored for autonomous agents. It can use ten tools at the same time. In Deep Research mode, the model itself googles, plans stages, and writes reports up to 10 thousand words long.
If you need to run a heavy code repository through the model, it handles it without the usual slowdowns that occur when working with voluminous texts.
β β BENCHMARKS
Step 3.5 Flash scored 97.3 on the AIME 2025 test (and this is bare risoning, without third-party calculators). If you give it access to Python, the result soars to 99.8.
On code benchmarks, the numbers also look impressive: on SWE-bench, it gives 74.4%, and on Terminal-Bench 2.0 - 51.0%.
Of course, in terms of knowledge density, Step 3.5 Flash still lags behind Gemini 3.0 Pro, but the fact that it is available for local use and tests via API, is pleasing.
Article, Model, Demo, Discord Community, GitHub β’ #AI #ML #LLM #StepFunAI
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
This media is not supported in your browser
VIEW IN TELEGRAM
KMeans clustering animation in the style of 3blue1brown
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Hierarchical Navigable Small World (HNSW) is an algorithm that makes vector search fast even on huge amounts of data, allowing you to search billions of vectors in milliseconds.
The idea of its operation is quite elegant - it's one of the most interesting discoveries of recent years.
π’ How it works?
You can read more in detail here β
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
The idea of its operation is quite elegant - it's one of the most interesting discoveries of recent years.
π’ How it works?
HNSW builds a multi-level graph, where each upper layer contains exponentially fewer nodes than the layer below.
β All vectors are located in the lower layer (layer 0), which is well connected.
β Only some vectors appear in layer 1, even fewer in layer 2, etc.
β The upper layers work as "fast lanes", allowing you to skip a large number of irrelevant data.
During the search, the algorithm starts from the upper layer, finds the nearest node, descends to the lower layer and repeats the process. By the time you reach the lower layer, you have already narrowed the search to the most relevant environment - there's no need to sort through everything.
This explains why HNSW is so economical with memory. It can "jump over" large amounts of data without evaluating each element.
Key parameters that affect the balance of speed and quality:
β ef - the size of the candidate list during the search
β maxConnections - how many connections each node can have
β distance - a metric for comparing vectors (cosine, dot product, etc.)
Adding new elements works in a similar way: first, we search for the optimal location, then we create connections. Restructuring the graph is resource-intensive, but queries themselves are performed very quickly.
You can read more in detail here β
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Someone has collected a collection of all production-ready LLM applications that can be made in 2026.
It's called awesome-llm-apps. It's literally copy-paste code for RAG, agents, multimodal applications, and AI SaaS products.
100% free. 100% Open Source.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
It's called awesome-llm-apps. It's literally copy-paste code for RAG, agents, multimodal applications, and AI SaaS products.
β Need RAG? Copy the code.
β Need AI agents? Copy the code.
β Need multimodal applications? Copy the code.
No hello world.
No training demos for beginners.
Only real applications that can be deployed today.
100% free. 100% Open Source.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
Graph-based RAG with dual-level retrieval.
LightRAG is an open-source RAG framework that builds knowledge graphs from documents and uses dual-level retrieval to answer both point-based and conceptual queries.
π’ Why it matters?
100% open source.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore
LightRAG is an open-source RAG framework that builds knowledge graphs from documents and uses dual-level retrieval to answer both point-based and conceptual queries.
π’ Why it matters?
Classical RAG relies on vector similarity and flat chunks. This is enough for superficial queries, but it breaks down when you need to understand how different concepts are connected.
LightRAG solves this problem by extracting entities and their relationships and forming a structured knowledge graph.
It uses LLMs to find entities (people, places, events) and their relationships in documents, then assembles a full-fledged knowledge graph that preserves these connections.
The framework works with dual-level retrieval:
Low-level retrieval targets specific entities and details, for example: What is Mechazilla?
High-level retrieval aggregates information across multiple entities for more general questions
such as: How does Elon Musk's vision contribute to sustainable development?
For each query, LightRAG extracts local and global keywords, matches them to graph nodes via vector similarity, and pulls in neighboring nodes one step at a time to expand the context.
What sets it apart:
β’ Graph-based indexing preserves connections between concepts rather than turning knowledge into isolated pieces
β’ Dual-level retrieval works for both point-based and conceptual queries
β’ Automatic entity extraction without manual labeling
β’ Incremental updates β new data is added without completely rebuilding the graph
β’ Multimodal support via RAG-Anything for PDFs, office documents, images, tables, and formulas
Key features:
β Knowledge graph visualization via WebUI
β Multiple storage backends (PostgreSQL, Neo4j, MongoDB, Qdrant)
β Support for major LLM providers (OpenAI, Anthropic, Ollama, Azure)
β Support for rerankers for mixed queries
β Document deletion with automatic knowledge graph regeneration
100% open source.
β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’β’
π€ Data & ML | @DataXplore