ReverseEngineering
1.32K subscribers
50 photos
11 videos
108 files
890 links
Download Telegram
بخش نوزدهم بافر اورفلو


Heap Overflow
از دید یک Reverse Engineer

تا الان بیشتر با استک کار کردیم
ولی خیلی از آسیب‌ پذیری‌ های واقعی روی Heap اتفاق میوفتن
قبل از اینکه بخوایم Heap Overflow رو تحلیل کنیم باید اول خود Heap رو بشناسیم

Heap اصلا چیه

بخشی از حافظه است که برنامه موقع اجرا هر وقت به حافظه بیشتری نیاز داشته باشه ازش استفاده میکنه
برخلاف Stack که خودکار مدیریت میشه
Heap کاملا تحت کنترل برنامه است
یعنی برنامه خودش درخواست حافظه میده
بعد هم خودش باید آزادش کنه

دو تابعی که همیشه میبینید

تقریبا تو هر برنامه C این دو تابع رو میبینید

malloc() free()

malloc حافظه رزرو میکنه

free هم همون حافظه رو آزاد میکنه

یک مثال ساده:

#include <stdlib.h> int main() { char *buf = malloc(32); free(buf); return 0; }


پشت صحنه چه اتفاقی میوفته

وقتی malloc(32) صدا زده میشه
برنامه فقط 32 بایت داده نمیگیره
کتابخانه مدیریت حافظه کنار اون داده یکسری اطلاعات مدیریتی هم ذخیره میکنه

به این مجموعه معمولا میگن Chunk
هر Chunk دو بخش داره
اطلاعات مدیریتی
داده‌ای که برنامه ازش استفاده میکنه

از دید مهندسی معکوس

وقتی داخل Ghidra یا IDA تابعی مثل این رو ببینید

char *buf = malloc(64)


strcpy(buf, input)


memcpy(buf, input, len)


باید بررسی کنید
آیا اندازه ورودی واقعا با اندازه Chunk یکیه یا نه

فرق Heap Overflow با Stack Overflow

در Stack Overflow معمولا هدف Return Address بود

ولی در Heap Overflow معمولا Return Address وجود نداره

در عوض چیزی که اهمیت داره ساختار Heap و داده‌های کنار حافظه است
به همین خاطر روش تحلیل این دو کاملا فرق میکنه

موقع آنالیز باینری دنبال چی بگردیم

اگر این الگوها رو دیدید بیشتر دقت کنید

malloc(...)

بعد از اون
strcpy(...)

یا
memcpy(...)

یا
read(...)


اگر اندازه ورودی کنترل نشده باشه
احتمال وجود Heap Overflow بیشتر میشه

نکته مهم

هر جا malloc دیدید به معنی آسیب‌ پذیری نیست

باید ببینید
چقدر حافظه گرفته شده
چقدر داده داخلش نوشته میشه
آیا قبل از نوشتن بررسی اندازه انجام شده یا نه

Heap با Stack فرق داره

حافظه Heap با malloc گرفته میشه
هر قطعه حافظه یک Chunk داره
و در مهندسی معکوس باید مسیر
ورودی → malloc → نوشتن داده
رو دنبال کنیم

چالش این قسمت:

یک برنامه ساده که از malloc استفاده میکنه داخل Ghidra باز کنید
مسیر حرکت داده از ورودی تا حافظه Heap رو پیدا کنید و مشخص کنید داده دقیقا کجا نوشته میشه

@reverseengine
ReverseEngineering
بخش نوزدهم بافر اورفلو Heap Overflow از دید یک Reverse Engineer تا الان بیشتر با استک کار کردیم ولی خیلی از آسیب‌ پذیری‌ های واقعی روی Heap اتفاق میوفتن قبل از اینکه بخوایم Heap Overflow رو تحلیل کنیم باید اول خود Heap رو بشناسیم Heap اصلا چیه بخشی از…
Part 19 Buffer Overflow


Heap Overflow
From a Reverse Engineer's Perspective

So far, we have worked mostly with the stack

But many real vulnerabilities occur on the heap

Before we try to analyze heap overflow, we must first understand the heap itself

What is the heap?

It is a part of memory that the program uses whenever it needs more memory during execution

Unlike the stack, which is automatically managed

The heap is completely under the control of the program

That is, the program itself requests memory

Then it must free it

Two functions that you always see

You will see these two functions in almost every C program

malloc() free()

malloc reserves memory

free also frees the same memory

A simple example:

#include <stdlib.h> int main() { char *buf = malloc(32); free(buf); return 0; }


What happens behind the scenes

When malloc(32) is called

The program doesn't just get 32 bytes of data

The memory management library also stores some management information along with that data

This collection is usually called a Chunk

Each Chunk has two parts

Management information

Data that the program uses

From a reverse engineering perspective

When you see a function like this in Ghidra or IDA

char *buf = malloc(64)


strcpy(buf, input)


memcpy(buf, input, len)


You need to check

Is the input size really the same as the Chunk size or not

The difference between Heap Overflow and Stack Overflow

In Stack Overflow, the target was usually the Return Address

But in Heap Overflow, there is usually no Return Address

Instead, what matters is the structure of the Heap and the data next to the memory

That's why the analysis method for these two is completely different

What to look for when analyzing binary Let's look

If you see these patterns, pay more attention

malloc(...)



After that

strcpy(...)



or

memcpy(...)



or

read(...)



If the input size is not controlled
The probability of Heap Overflow increases

Important point

Wherever you see malloc, it does not mean a vulnerability

You need to see
How much memory is taken
How much data is written into it
Whether the size check is done before writing

Heap is different from Stack

Heap memory is taken with malloc
Each piece of memory has a Chunk
And in reverse engineering, we need to follow the path
Input → malloc → writing data

Challenge of this part:

Open a simple program that uses malloc in Ghidra
Find the path of data movement from input to Heap memory and determine where exactly the data is written

@reverseengine
درود دوستان امیدوارم حالتون خوب باشه اگر این چند روز زیاد فعال نیستم و اینکه پست نمی‌ذارم به دلیل جام جهانیه و اینکه ساعتشون با کشور ما متاسفانه یکی نیست و یه سری مشکلات هست که برام پیش اومده وقتی که اوکی شد دوباره پر قدرت برمیگردیم زیاد طول نمیکشه چند روز دیگه برمیگردم دوستون دارم
Alone. 🩶

Hello friends, I hope you are well. If I haven't been very active these past few days, and I haven't been posting, it's because of the World Cup, and unfortunately their time zone is not the same as ours, and there are a number of problems that have come up for me. When it's OK, we'll be back with full force again. It won't be long. I'll be back in a few days. I love you all,
Alone. 🖤
❤10
Handler
جایی که کار واقعی انجام میشه

تو پست قبل گفتیم Dispatcher فقط تصمیم میگیره کدوم Opcode اجرا بشه
اما خودش هیچ کار خاصی انجام نمیده
کار اصلی داخل Handler ها انجام میشه
هر Opcode حداقل یک Handler داره
مثلا فرض کنید این بایت‌ کد رو داریم:

01 05 02 03 03


اگر معنی Opcode ها این باشه:

01 = LOAD
02 = ADD
03 = PRINT


اتفاقی که میوفته اینه:

Dispatcher
مقدار 01 رو می‌بینه و Handler مربوط به LOAD اجرا میشه

داخل Handler مقدار 05 داخل یکی از رجیسترهای مجازی ذخیره میشه
بعد کنترل دوباره برمیگرده به Dispatcher

این بار Dispatcher مقدار 02 رو میخونه

Handler
مربوط به ADD اجرا میشه و عدد 03 به مقدار قبلی اضافه میشه
دوباره کنترل برمیگرده به Dispatcher
در آخر Opcode 03 اجرا میشه

Handler
مربوط به PRINT مقدار نهایی رو چاپ میکنه

پس همیشه این چرخه تکرار میشه:

Dispatcher ↓ Handler ↓ Dispatcher ↓ Handler ↓ Dispatcher

نکته مهم اینجاست که همه Handler ها شبیه هم نیستن
بعضی‌ ها فقط یک مقدار رو جا به‌ جا میکنن
بعضی‌ ها عملیات ریاضی انجام میدن
بعضی‌ ها پرش شرطی انجام میدن
بعضی‌ ها حافظه رو میخونن یا مینویسن
در ماشین‌ های مجازی واقعی مثل VMProtect ممکنه صدها Handler مختلف وجود داشته باشه
به همین خاطر اولین کاری که تحلیلگر انجام میده دسته‌بندی Handler هاست
مثلا بعد از بررسی چندتا از اونا متوجه میشه:

این Handler همیشه داده رو جا به‌ جا میکنه
اون یکی همیشه عملیات XOR انجام میده
یکی دیگه همیشه مقدارها رو با هم مقایسه میکنه
یکی هم مسئول پرش‌ های شرطیه
کم‌ کم با کنار هم گذاشتن این اطلاعات معنی Opcode ها مشخص میشه
این دقیقا همون چیزیه که بهش Opcode Mapping میگن
یعنی ساختن یک جدول که مشخص کنه هر Opcode چه کاری انجام میده

تمرین

یک جدول برای Opcode های زیر درست کنید:

01 → LOAD 02 → ADD 03 → XOR 04 → CMP 05 → JMP 06 → PRINT


بعد سعی کنبد برای هر Opcode بنویسید Handler اون باید چکاری انجام بده
این تمرین باعث میشه ذهنتون دقیقا مثل یک تحلیلگر Devirtualization شروع به فکر کردن کنه





Handler
Where the real work is done

In the previous post, we said that Dispatcher only decides which Opcode to execute
But it doesn't do anything special
The main work is done inside Handlers
Each Opcode has at least one Handler
For example, suppose we have this bytecode:

01 05 02 03 03


If the meaning of the Opcodes is:

01 = LOAD
02 = ADD
03 = PRINT


What happens is this:

Dispatcher
Sees the value 01 and the Handler corresponding to LOAD is executed

Inside the Handler, the value 05 is stored in one of the virtual registers
Then control returns to Dispatcher

This time Dispatcher reads the value 02

Handler
Related to ADD is executed and the number 03 is added to the previous value
Control returns to Dispatcher again
Finally, Opcode 03 is executed

Handler
Related to PRINT Prints the final value

So this cycle always repeats:

Dispatcher ↓ Handler ↓ Dispatcher ↓ Handler ↓ Dispatcher


The important thing here is that not all Handlers are the same

Some just move a value

Some perform mathematical operations

Some perform conditional jumps

Some read or write memory.
In real virtual machines like VMProtect, there may be hundreds of different Handlers
That is why the first thing the analyst does is to categorize the Handlers

For example, after examining a few of them, he will notice:

This Handler always moves data
That one always performs XOR operations
Another always compares values
One is responsible for conditional jumps
Little by little, by putting this information together, the meaning of the Opcodes becomes clear
This is exactly what is called Opcode Mapping
That is, creating a table that specifies what each Opcode does

Exercise:

Create a table for the following Opcodes:

01 → LOAD 02 → ADD 03 → XOR 04 → CMP 05 → JMP 06 → PRINT


Then try to write for each Opcode what its Handler should do
This exercise will make your mind start thinking exactly like a Devirtualization analyst

@reverseengine
وقتی روی یک برنامه دوبار کلیک میکنید پشت صحنه چه اتفاقی میوفته؟

تا اینجا فهمیدیم Process چیه و داخلش چه بخش‌ هایی وجود داره

حالا یک سوال جالب:

وقتی روی notepad.exe یا chrome.exe دوبار کلیک میکنید دقیقا چه اتفاقی میوفته؟

همه این مراحل در چند میلی ثانیه انجام میشن اما پشت صحنه کارهای زیادی انجام میشه

مرحله 1: پیدا کردن فایل اجرایی

اول سیستم‌عامل فایل برنامه را روی دیسک پیدا میکنه

مثلا:

C:\Windows\System32\notepad.exe



در این مرحله هنوز هیچ کدی اجرا نشده فقط فایل پیدا شده

مرحله 2: ساختن Process

حالا کرنل یک Process جدید ایجاد میکنه

برای این Process اطلاعاتی مثل:

شناسه پردازه (PID)

وضعیت پردازه

اطلاعات زمان‌ بندی

مجوزهای دسترسی


رو آماده میکنه

از این لحظه سیستم‌ عامل میتونه این برنامه رو مدیریت کنه

مرحله 3: ایجاد فضای حافظه

حالا سیستم‌ عامل یک Virtual Address Space برای Process میسازه

یعنی یک فضای حافظه اختصاصی که فقط متعلق به همین Process هست.

در این فضا بخش‌ هایی مثل:

Code
Data
Heap
Stack


قرار میگیرن

نکته مهم اینجاست که هر Process فضای حافظه مخصوص خودش رو داره و معمولا نمیتونه مستقیما به حافظه Process دیگه ای دسترسی پیدا کنه

مرحله 4: بارگذاری فایل اجرایی

حالا سیستم‌ عامل فایل اجرایی رو داخل حافظه اپلود میکنه

بخش‌ های مختلف فایل مثل کد و داده‌ها در جای مناسب خودشون قرار میگیرن

اگر برنامه به کتابخانه‌ هایی مثل DLL نیاز داشته باشه اونا هم در همین مرحله اپلود میشن

مرحله 5: آماده شدن Thread اصلی

هر Process حداقل یک Thread داره
که به اون Main Thread میگن

این Thread از اولین دستور برنامه شروع به اجرا مبکنه بعدا اگر برنامه لازم داشته باشه میتونه Thread های بیشتری بسازه

مرحله 6: شروع اجرای برنامه

حالا CPU اجرای اولین دستور برنامه رو شروع میکنه

از این لحظه به بعد برنامه واقعا در حال اجراست

یک مثال واقعی:

فرض کنید روی Google Chrome کلیک میکنید

در چند لحظه:

✅ فایل Chrome پیدا میشه

✅ یک Process جدید ساخته میشه

✅ حافظه اختصاصی برای اون ایجاد میشه

✅ فایل اجرایی و DLL های مورد نیاز اپلود میشن

✅ Main Thread ساخته میشه

✅ اولین دستور Chrome اجرا میشه

همه این مراحل اونقدر سریع انجام میشن که کاربر فقط باز شدن پنجره مرورگر رو میبینه

چرا این موضوع برای مهندسی معکوس مهمه؟

وقتی در ابزارهایی مثل دیباگر یا تحلیل‌ بدافزار کار میکنید باید بدونید برنامه از لحظه اجرا چه مراحلی رو طی کرده

برای مثال بیشتر بدافزارها:

قبل از اجرای کد اصلی DLL های خاصی رو اپلود میکنن

در همون ابتدای اجرا چند Thread جدید میسازن

حافظه جدید اختصاص میدن و کد خودشون رو در اون قرار میدن

اگر فرآیند ایجاد Process رو بشناسید تحلیل چنین رفتار هایی بسیار ساده‌ تر میشه



What happens behind the scenes when you double-click on a program?

So far we have understood what a Process is and what parts it contains

Now an interesting question:

What exactly happens when you double-click on notepad.exe or chrome.exe?

All these steps are done in a few milliseconds, but a lot of work is done behind the scenes

Step 1: Finding the executable file

First, the operating system finds the program file on the disk

For example:

C:\Windows\System32\notepad.exe


At this stage, no code is executed yet, only the file is found

Step 2: Creating a Process

Now the kernel creates a new Process

For this Process, it provides information such as:

Process ID (PID)

Process status

Scheduling information

Access permissions


From this moment, the operating system can manage this program

Step 3: Creating a memory space

Now the operating system creates a Virtual Address Space for the Process

That is, a dedicated memory space that belongs only to this Process.

In this space, sections such as:

Code

Data

Heap

Stack


The important point here is that each Process has its own memory space and usually cannot directly access the memory of another Process

Step 4: Loading the executable file

Now the operating system uploads the executable file into memory

Different parts of the file such as code and data are placed in their appropriate places

If the program needs libraries such as DLLs, they are also uploaded at this stage

Step 5: Preparing the main thread

Each process has at least one thread

which is called the main thread
This thread starts executing from the first command of the program, and later if the program needs it, it can create more threads

Step 6: Starting the program

Now the CPU starts executing the first command of the program

From this moment on, the program is really running

A real example:

Suppose you click on Google Chrome

In a few moments:

✅ The Chrome file is found

✅ A new process is created

✅ Dedicated memory is created for it

✅ The executable and required DLLs are uploaded

✅ The main thread is created

✅ The first Chrome command is executed

All these steps are done so quickly that the user only sees the browser window opening

Why is this important for reverse engineering?

When working in tools such as debuggers or malware analysis, you need to know what steps the program has taken since execution

For example, most malware:

Uploads specific DLLs before executing the main code

Creates several new threads at the very beginning of execution

Allocates new memory and places its code in it

If you understand the process of process creation, analyzing such behavior becomes much easier

@reverseengine
EDR

ها  چطور حملات رو تشخیص میدن؟

خیلی‌ ها فکر میکنن EDR فقط دنبال امضای فایل یا اسم API ها هستن ولی واقعیت اینه که نسل جدید EDR ها بیشتر روی رفتار تمرکز دارن

اگر یه برنامه این کارها رو انجام بده:

یک پروسه جدید ایجاد میکنه

حافظه قابل اجرا (Executable Memory) رزرو میکنه

داخل اون حافظه داده مینویسه
یک Thread جدید اجرا میکنه
ممکنه هیچ فایل مخربی روی دیسک وجود نداشته باشه اما همین زنجیره رفتارها میتونه امتیاز ریسک بالایی ایجاد کنه

EDR
ها به چه چیزایی دقت میکنن؟

ارتباط بین پروسه‌ها (Process Tree)

ترتیب رخدادها (Event Correlation)

دسترسی های غیرعادی به حافظه

اپلود DLL های غیرمنتظره

تغییرات در Tokenها و دسترسی‌ها

الگوهای ارتباط شبکه

اطلاعات تله‌ متری ویندوز مثل ETW


چرا فقط یک رفتار کافی نیست؟

مثلا:

Process Created
      ↓
Memory Allocation
      ↓
Memory Write
      ↓
Thread Execution


هر کدوم از این‌ها به تنهایی ممکنه طبیعی باشن اما وقتی پشت سر هم باشن برای EDR یک الگوی مشکوک ایجاد میکنن


تشخیص دادن

نسل جدید EDR ها معمولا از مدل‌ های رفتاری استفاده میکنن نه فقط Signature
به همین دلیل ممکنه دو فایل کاملا متفاوت اگر رفتار مشابهی داشته باشن هر دو شناسایی بشن



How EDRs Detect Attacks?

Many people think that EDRs only look for file signatures or API names, but the truth is that the new generation of EDRs focuses more on behavior

If a program does these things:

Creates a new process

Reserves executable memory

Writes data into that memory

Runs a new thread


There may not be any malicious files on disk, but this chain of behavior can generate a high risk score

What do EDRs look for?

Process Tree

Event Correlation

Unusual Memory Access

Unexpected DLL Uploads

Token Changes and Accesses

Network Communication Patterns

Windows Telemetry Information like ETW


Why is just one behavior not enough?

For example:

Process Created
↓
Memory Allocation
↓
Memory Write
↓
Thread Execution


Each of these may be normal on their own, but when they are in a row, they create a suspicious pattern for EDR

Detection

Newer generation EDRs usually use behavioral models, not just Signature

That's why two completely different files may both be detected if they have similar behavior


@reverseengine
🔥1
بخش بیستم بافر اورفلو


Use After Free
یکی از خطرناک‌ترین باگ‌های حافظه


تا اینجا یاد گرفتیم malloc حافظه میگیره و free اون رو آزاد میکنه حالا میخوایم ببینیم اگر برنامه بعد از free دوباره از همون حافظه استفاده کنه چه اتفاقی میوفته و چطور موقع مهندسی معکوس این باگ رو تشخیص بدیم

Use After Free
یعنی چی

اسمش کاملا مشخصه اول حافظه آزاد میشه
بعد برنامه دوباره از همون حافظه استفاده میکنه یعنی برنامه فکر میکنه حافظه هنوز معتبره در حالی که سیستم عامل یا مدیریت Heap ممکنه اون حافظه رو برای چیز دیگه‌ ای استفاده کرده باشه

یک مثال ساده

#include <stdio.h> #include <stdlib.h> int main() { char *buf = malloc(32); free(buf); puts(buf); return 0; }


مشکل این کجاست؟

اینجا
free(buf);


حافظه آزاد شده ولی چند خط بعد

puts(buf);

از همون اشاره‌گر استفاده شده
این دقیقا یک Use After Free هست

چرا خطرناکه؟

چون بعد از free دیگه هیچ تضمینی وجود نداره که داده‌های قبلی داخل اون حافظه باقی مونده باشن ممکنه داده تغییر کرده باشه حافظه دوباره به بخش دیگه‌ای اختصاص داده شده باشه برنامه کرش کنه

موقع مهندسی معکوس دنبال چی بگردیم؟

اگر الگوی اینجوری دیدید
free(ptr);

و بعد چند خط پایین‌تر

ptr->field

یا
memcpy(ptr,...)

یا
printf("%s", ptr);

یا هر استفاده دیگه از ptr باید بررسی کنید که آیا همون اشاره‌گر بعد از free دوباره استفاده شده یا نه

یک مثال در دیکامپایلر

buf = malloc(64); /* ... */ free(buf); /* ... */ strcpy(buf, input);

همین چند خط برای مشکوک شدن کافیه

چطور از این باگ جلوگیری میکنن؟

یکی از ساده‌ترین روش‌ها اینه که بعد از free
اشاره‌گر رو NULL کنن

free(buf); buf = NULL;

حالا اگر جایی دوباره از buf استفاده بشه
اشکال خیلی زودتر مشخص میشه


free یعنی حافظه آزاد شده

بعد از free نباید از همون اشاره‌گر استفاده کرد
در مهندسی معکوس همیشه مسیر

malloc → free →

استفاده مجدد رو بررسی کنید

این یکی از الگوهای مهم برای پیدا کردن باگ‌های مدیریت حافظه است

تمرین:

یک برنامه ساده که از malloc و free استفاده میکنه داخل Ghidra یا IDA باز کنید و بررسی کنید آیا بعد از free دوباره از همون اشاره‌گر استفاده شده یا نه اگر اینجور الگویی پیدا کردید دلیلش رو تحلیل کنید و مشخص کنید آیا واقعا یک Use After Free است یا فقط ظاهرا این‌طور به نظر میرسه

@reverseengine
❤2
ReverseEngineering
بخش بیستم بافر اورفلو Use After Free یکی از خطرناک‌ترین باگ‌های حافظه تا اینجا یاد گرفتیم malloc حافظه میگیره و free اون رو آزاد میکنه حالا میخوایم ببینیم اگر برنامه بعد از free دوباره از همون حافظه استفاده کنه چه اتفاقی میوفته و چطور موقع مهندسی معکوس…
Part 20 Buffer Overflow


Use After Free One of the most dangerous memory bugs

So far we have learned that malloc takes memory and free frees it. Now we want to see what happens if the program uses the same memory again after free and how to detect this bug when reverse engineering

What does Use After Free
mean? Its name is quite specific. First the memory is freed. Then the program uses the same memory again, meaning the program thinks the memory is still valid, while the operating system or Heap Manager may have used that memory for something else.

A simple example:

#include <stdio.h> #include <stdlib.h>
int main() { char *buf = malloc(32); free(buf); puts(buf); return 0; }
Where is the problem with this?

Here

free(buf);


The memory is freed, but a few lines later

puts(buf);


The same pointer is used. This is exactly a Use After Free

Why is it dangerous?

Because after free there is no guarantee that the previous data will remain in that memory. The data may have changed, the memory may have been reallocated, the program may crash

What should we look for when reverse engineering?

If you see a pattern like this

free(ptr);


and then a few lines down

ptr->field


or

memcpy(ptr,...)


or

printf("%s", ptr);


or any other use of ptr, you should check whether the same pointer is reused after free

An example in the decompiler

buf = malloc(64); /* ... */ free(buf); /* ... */ strcpy(buf, input);


These few lines are enough to be suspicious

How do you prevent this bug?

One of the easiest ways is to NULL the pointer after free

free(buf); buf = NULL;


Now if buf is used again somewhere
The problem will be identified much earlier

free means memory has been freed

You should not use the same pointer after free

In reverse engineering, always follow the path

malloc → free →

Check for reuse

This is one of the important patterns for finding memory management bugs

Exercise:

Open a simple program that uses malloc and free in Ghidra or IDA and check if the same pointer is used again after free or not. If you find such a pattern, analyze the reason and determine if it is really a Use After Free or it just looks like it.

@reverseengine
❤4
اینجا به یکی از مهم‌ترین مفاهیم Binary Exploitation میرسیم

Info Leak

چرا اولین هدف یک اکسپلویت مدرنه؟


تا اینجا چند تا مکانیزم امنیتی رو شناختیم:

Stack Canary

NX

ASLR
Information Leak
یا Info Leak یعنی چی؟

Info Leak
یعنی برنامه بدون اینکه قرار بوده بخشی از اطلاعات حافظه رو در اختیار کاربر قرار بده

این اطلاعات ممکنه شامل مواردی مثل:

آدرس یک تابع

آدرس یک متغیر

آدرس heap

آدرس stack

آدرس libc


یا حتی داده‌ های حساس داخل حافظه
باشه در ظاهر شاید این اطلاعات مهم به نظر نرسن اما برای یک اکسپلویتر میتونن حکم نقشه‌ ی یک ساختمون رو داشته باشن

چرا اینقدر مهمه؟

فرض کنید ASLR فعاله

یعنی آدرس‌ها در هر بار اجرای برنامه تغییر میکنن
اگر هیچ آدرسی رو ندونید ساختن یک اکسپلویت قابل‌ اعتماد خیلی سخت میشه
اما اگر برنامه فقط یک آدرس رو ناخواسته لو بده مهاجم میتونه از همون آدرس موقعیت بقیه بخش‌ های حافظه رو هم محاسبه کنه
در این حالت بخش بزرگی از مزیت ASLR از بین میره

گاهی اوقات Info Leak خودش یک آسیب‌پذیری مستقل محسوب میشه و گاهی هم نتیجه‌ ی یک باگ دیگه است

مثلا ممکنه:

برنامه بیشتر از حد لازم داده چاپ کنه
یک اشاره‌گر (Pointer) رو مستقیم نمایش بده

بخشی از حافظه رو بدون پاک کردن برگردونه

یا یک خطای منطقی باعث افشای اطلاعات بشه

چرا بیشتر اکسپلویت‌های امروزی دو مرحله‌ای هستند؟

خیلی از حملات مدرن این شکلی‌اند:

مرحله اول پیدا کردن یک Info Leak

مرحله دوم استفاده از اطلاعات به‌دست‌ اومده برای دور زدن مکانیزم‌ هایی مثل ASLR و ادامه‌ ی حمله

به همین دلیل تعداد زیادی از تحلیل‌های امنیتی که میبینید که اولین هدف نه اجرای کد بلکه پیدا کردن یک راه برای دیدن حافظه است

یک شبیه سازی ساده:

فرض کنید وارد یک شهر ناشناس شدید
اگر هیچ نقشه‌ ای نداشته باشید پیدا کردن یک ساختمون خاص خیلی سخت میشه
اما اگر فقط آدرس یک خیابون رو بهتون بدن کم‌ کم میتونید کل نقشه شهر رو پیدا کنید

Info Leak
دقیقا همین نقش رو در حافظه برنامه‌ها داره

مکانیزم‌ هایی مثل ASLR آدرس‌ ها رو مخفی میکنن اما اگر برنامه حتی مقدار کمی از اطلاعات حافظه رو افشا کنه این محافظت تا حد زیادی بی‌ اثر میشه به همین دلیل پیدا کردن Info Leak یکی از مهم‌ ترین مراحل در تحلیل و درک اکسپلویت‌ های مدرنه
@reverseengine
ReverseEngineering
اینجا به یکی از مهم‌ترین مفاهیم Binary Exploitation میرسیم Info Leak چرا اولین هدف یک اکسپلویت مدرنه؟ تا اینجا چند تا مکانیزم امنیتی رو شناختیم: Stack Canary NX ASLR Information Leak یا Info Leak یعنی چی؟ Info Leak یعنی برنامه بدون اینکه قرار بوده…
Here we come to one of the most important concepts of Binary Exploitation

Info Leak

Why is it the first target of a modern exploit?

So far we have recognized several security mechanisms:

Stack Canary

NX

ASLR
What is Information Leak or Info Leak?

Info Leak
means that the program provides some memory information to the user without being supposed to

This information may include things like:

The address of a function

The address of a variable

The heap address

The stack address

The libc address


Or even sensitive data in memory

On the surface, this information may not seem important, but for an exploiter, it can be like a blueprint of a building

Why is it so important?

Let's say ASLR is enabled

This means that the addresses change every time the program is run

If you don't know any addresses, it's very difficult to create a reliable exploit

But if the program accidentally leaks just one address, the attacker can calculate the location of other memory locations from that address

In this case, a large part of the benefit of ASLR is lost

Sometimes an Info Leak is a standalone vulnerability, and sometimes it's the result of another bug

For example, it's possible to:

The program prints more data than necessary

Display a pointer directly

Return a portion of memory without clearing

Or a logical error causes information to be leaked

Why are most exploits today two-step?

Many modern attacks look like this:

The first step is to find an Info Leak

The second step is to use the information obtained to bypass mechanisms such as ASLR and continue the attack

That is why a lot of security analysis that you see the first goal is not to execute code but to find a way to see the memory

A simple simulation:

Suppose you enter an unknown city

If you have no map, it will be very difficult to find a specific building

But if you are only given the address of a street, you can gradually find the entire map of the city

Info Leak

It plays exactly the same role in the memory of programs

Mechanisms such as ASLR hide addresses, but if the program reveals even a small amount of memory information, this protection becomes largely ineffective. This is why finding Info Leak is one of the most important steps in analyzing and understanding modern exploits
@reverseengine
Virtual Register رجیستر های مجازی

تا اینجا فهمیدیم Dispatcher تصمیم میگیره کدوم Handler اجرا بشه و Handler هم کار اصلی رو انجام میده


Handler
ها داده‌ ها رو کجا نگه میدارن؟

روی رجیستر های واقعی CPU مثل RAX و RBX؟ معمولا نه

بیشتر ماشین‌های مجازی از چیزی به اسم Virtual Register استفاده میکنن

Virtual Register
یه فضای حافظه‌ ست که نقش رجیستر های CPU رو بازی میکنه

مثلا فرض کنید ماشین مجازی 8 تا رجیستر داره:

V0 V1 V2 V3 V4 V5 V6 V7


وقتی Opcode زیر اجرا میشه:

LOAD V0, 10

مقدار 10 داخل V0 قرار میگیره

بعد Opcode بعدی:

LOAD V1, 20

عدد 20 داخل V1 ذخیره میشه

حالا Opcode بعدی:

ADD V0, V1


ماشین مجازی مقدار V0 و V1 رو جمع میکنه

نتیجه دوباره داخل V0 ذخیره میشه

اگر بعدش این Opcode اجرا بشه:

PRINT V0

عدد 30 چاپ میشه

اینجا هیچکدوم از رجیسترهای واقعی CPU مستقیما دیده نمیشن

ممکنه داخل Handlerه ا فقط چند دستور مثل این ببینید:

mov rax, [rdi+18h]

mov rcx, [rdi+20h]

add rax, rcx

mov [rdi+18h], rax


در نگاه اول انگار برنامه فقط داره با حافظه کار میکنه

ولی در واقع:

[rdi+18h] = V0 [rdi+20h] = V1


یعنی این آدرس‌ های حافظه همون رجیستر های مجازی هستن به همین دلیل یکی از اولین کارهای تحلیلگر اینه که محل نگهداری Virtual Register ها رو پیدا کنه
وقتی بفهمید هر Offset مربوط به کدوم Virtual Register هست خوندن Handler ها چند برابر راحت‌ تر میشه

خیلی وقت‌ ها حتی اسم‌گذاری هم میکنن:

[rdi+18h] → V0 [rdi+20h] → V1 [rdi+28h] → V2


از این لحظه به بعد به جای آدرس‌ های حافظه ذهن تحلیلگر با رجیسترهای مجازی کار میکنه این کار باعث میشه منطق ماشین مجازی کم‌ کم شبیه یک CPU معمولی به نظر برسه

تمرین:

فرض کنید یک ماشین مجازی چهار رجیستر داره:

V0 V1 V2 V3


و Opcode های زیر اجرا میشن:

LOAD V0, 15

LOAD V1, 5

SUB V0, V1

PRINT V0


بدون اینکه کدی بنویسید مرحله‌ به‌ مرحله مشخص کنید بعد از اجرای هر Opcode مقدار هر Virtual Register چقدر میشه




Virtual Register So far we have understood that Dispatcher decides which Handler to execute and Handler does the main work

Where do Handlers store data?

On real CPU registers like RAX and RBX? Usually not

Most virtual machines use something called Virtual Register

Virtual Register
is a memory space that acts as CPU registers

For example, suppose the virtual machine has 8 registers:

V0 V1 V2 V3 V4 V5 V6 V7


When the following Opcode is executed:

LOAD V0, 10


10 is placed into V0

Then the next Opcode:

LOAD V1, 20


20 is stored into V1

Now the next Opcode:

ADD V0, V1


The virtual machine adds V0 and V1

The result is stored back into V0

If this Opcode is then executed:

PRINT V0


30 is printed

Here none of the actual CPU registers are directly visible

You may only see a few instructions inside the Handlers like this:

mov rax, [rdi+18h]

mov rcx, [rdi+20h]

add rax, rcx

mov [rdi+18h], rax


At first glance, it seems like the program is only working with memory

But in fact:

[rdi+18h] = V0 [rdi+20h] = V1



That is, these memory addresses are the same virtual registers, which is why one of the first tasks of the analyst is to find the location of the Virtual Registers.

When you understand which Virtual Register each Offset belongs to, reading the Handlers becomes much easier.

Many times they even give them names:

[rdi+18h] → V0 [rdi+20h] → V1 [rdi+28h] → V2


From this moment on, instead of memory addresses, the analyst's mind works with virtual registers. This makes the logic of the virtual machine look a little like a regular CPU.

Exercise:

Suppose a virtual machine It has four registers:

V0 V1 V2 V3


And the following Opcodes are executed:

LOAD V0, 15

LOAD V1, 5

SUB V0, V1

PRINT V0


Without writing any code, determine step by step what the value of each Virtual Register will be after executing each Opcode.


@reverseengine