بخش نوزدهم بافر اورفلو
Heap Overflow
از دید یک Reverse Engineer
تا الان بیشتر با استک کار کردیم
ولی خیلی از آسیب پذیری های واقعی روی Heap اتفاق میوفتن
قبل از اینکه بخوایم Heap Overflow رو تحلیل کنیم باید اول خود Heap رو بشناسیم
Heap اصلا چیه
بخشی از حافظه است که برنامه موقع اجرا هر وقت به حافظه بیشتری نیاز داشته باشه ازش استفاده میکنه
برخلاف Stack که خودکار مدیریت میشه
Heap کاملا تحت کنترل برنامه است
یعنی برنامه خودش درخواست حافظه میده
بعد هم خودش باید آزادش کنه
دو تابعی که همیشه میبینید
تقریبا تو هر برنامه C این دو تابع رو میبینید
malloc() free() malloc حافظه رزرو میکنه
free هم همون حافظه رو آزاد میکنه
یک مثال ساده:
#include <stdlib.h> int main() { char *buf = malloc(32); free(buf); return 0; }
پشت صحنه چه اتفاقی میوفته
وقتی malloc(32) صدا زده میشه
برنامه فقط 32 بایت داده نمیگیره
کتابخانه مدیریت حافظه کنار اون داده یکسری اطلاعات مدیریتی هم ذخیره میکنه
به این مجموعه معمولا میگن Chunk
هر Chunk دو بخش داره
اطلاعات مدیریتی
دادهای که برنامه ازش استفاده میکنه
از دید مهندسی معکوس
وقتی داخل Ghidra یا IDA تابعی مثل این رو ببینید
char *buf = malloc(64)strcpy(buf, input)memcpy(buf, input, len)باید بررسی کنید
آیا اندازه ورودی واقعا با اندازه Chunk یکیه یا نه
فرق Heap Overflow با Stack Overflow
در Stack Overflow معمولا هدف Return Address بود
ولی در Heap Overflow معمولا Return Address وجود نداره
در عوض چیزی که اهمیت داره ساختار Heap و دادههای کنار حافظه است
به همین خاطر روش تحلیل این دو کاملا فرق میکنه
موقع آنالیز باینری دنبال چی بگردیم
اگر این الگوها رو دیدید بیشتر دقت کنید
malloc(...) بعد از اون
strcpy(...)
یا
memcpy(...)
یا
read(...) اگر اندازه ورودی کنترل نشده باشه
احتمال وجود Heap Overflow بیشتر میشه
نکته مهم
هر جا malloc دیدید به معنی آسیب پذیری نیست
باید ببینید
چقدر حافظه گرفته شده
چقدر داده داخلش نوشته میشه
آیا قبل از نوشتن بررسی اندازه انجام شده یا نه
Heap با Stack فرق داره
حافظه Heap با malloc گرفته میشه
هر قطعه حافظه یک Chunk داره
و در مهندسی معکوس باید مسیر
ورودی → malloc → نوشتن داده
رو دنبال کنیم
چالش این قسمت:
یک برنامه ساده که از malloc استفاده میکنه داخل Ghidra باز کنید
مسیر حرکت داده از ورودی تا حافظه Heap رو پیدا کنید و مشخص کنید داده دقیقا کجا نوشته میشه
@reverseengine
ReverseEngineering
بخش نوزدهم بافر اورفلو Heap Overflow از دید یک Reverse Engineer تا الان بیشتر با استک کار کردیم ولی خیلی از آسیب پذیری های واقعی روی Heap اتفاق میوفتن قبل از اینکه بخوایم Heap Overflow رو تحلیل کنیم باید اول خود Heap رو بشناسیم Heap اصلا چیه بخشی از…
Part 19 Buffer Overflow
Heap Overflow
From a Reverse Engineer's Perspective
So far, we have worked mostly with the stack
But many real vulnerabilities occur on the heap
Before we try to analyze heap overflow, we must first understand the heap itself
What is the heap?
It is a part of memory that the program uses whenever it needs more memory during execution
Unlike the stack, which is automatically managed
The heap is completely under the control of the program
That is, the program itself requests memory
Then it must free it
Two functions that you always see
You will see these two functions in almost every C program
malloc() free()
malloc reserves memory
free also frees the same memory
A simple example:
#include <stdlib.h> int main() { char *buf = malloc(32); free(buf); return 0; }
What happens behind the scenes
When malloc(32) is called
The program doesn't just get 32 bytes of data
The memory management library also stores some management information along with that data
This collection is usually called a Chunk
Each Chunk has two parts
Management information
Data that the program uses
From a reverse engineering perspective
When you see a function like this in Ghidra or IDA
char *buf = malloc(64)
strcpy(buf, input)
memcpy(buf, input, len)
You need to check
Is the input size really the same as the Chunk size or not
The difference between Heap Overflow and Stack Overflow
In Stack Overflow, the target was usually the Return Address
But in Heap Overflow, there is usually no Return Address
Instead, what matters is the structure of the Heap and the data next to the memory
That's why the analysis method for these two is completely different
What to look for when analyzing binary Let's look
If you see these patterns, pay more attention
malloc(...)
After that
strcpy(...)
or
memcpy(...)
or
read(...)
If the input size is not controlled
The probability of Heap Overflow increases
Important point
Wherever you see malloc, it does not mean a vulnerability
You need to see
How much memory is taken
How much data is written into it
Whether the size check is done before writing
Heap is different from Stack
Heap memory is taken with malloc
Each piece of memory has a Chunk
And in reverse engineering, we need to follow the path
Input → malloc → writing data
Challenge of this part:
Open a simple program that uses malloc in Ghidra
Find the path of data movement from input to Heap memory and determine where exactly the data is written
@reverseengine
Windows Process Injection: Print Spooler
https://modexp.wordpress.com/2019/03/07/process-injection-print-spooler
https://modexp.wordpress.com/2019/03/07/process-injection-print-spooler
modexp
Windows Process Injection: Print Spooler
Introduction Every application running on the windows operating system has a thread pool or a “worker factory” and this internal mechanism allows an application to offload management of…
درود دوستان امیدوارم حالتون خوب باشه اگر این چند روز زیاد فعال نیستم و اینکه پست نمیذارم به دلیل جام جهانیه و اینکه ساعتشون با کشور ما متاسفانه یکی نیست و یه سری مشکلات هست که برام پیش اومده وقتی که اوکی شد دوباره پر قدرت برمیگردیم زیاد طول نمیکشه چند روز دیگه برمیگردم دوستون دارم
Alone. 🩶
Hello friends, I hope you are well. If I haven't been very active these past few days, and I haven't been posting, it's because of the World Cup, and unfortunately their time zone is not the same as ours, and there are a number of problems that have come up for me. When it's OK, we'll be back with full force again. It won't be long. I'll be back in a few days. I love you all,
Alone. 🖤
Alone. 🩶
Hello friends, I hope you are well. If I haven't been very active these past few days, and I haven't been posting, it's because of the World Cup, and unfortunately their time zone is not the same as ours, and there are a number of problems that have come up for me. When it's OK, we'll be back with full force again. It won't be long. I'll be back in a few days. I love you all,
Alone. 🖤
❤10
ReverseEngineering
درود دوستان امیدوارم حالتون خوب باشه اگر این چند روز زیاد فعال نیستم و اینکه پست نمیذارم به دلیل جام جهانیه و اینکه ساعتشون با کشور ما متاسفانه یکی نیست و یه سری مشکلات هست که برام پیش اومده وقتی که اوکی شد دوباره پر قدرت برمیگردیم زیاد طول نمیکشه چند روز…
امشب دوباره شروع میکنیم.
We start again tonight.
We start again tonight.
🔥7
Handler
جایی که کار واقعی انجام میشه
تو پست قبل گفتیم Dispatcher فقط تصمیم میگیره کدوم Opcode اجرا بشه
اما خودش هیچ کار خاصی انجام نمیده
کار اصلی داخل Handler ها انجام میشه
هر Opcode حداقل یک Handler داره
مثلا فرض کنید این بایت کد رو داریم:
اگر معنی Opcode ها این باشه:
اتفاقی که میوفته اینه:
Dispatcher
مقدار 01 رو میبینه و Handler مربوط به LOAD اجرا میشه
داخل Handler مقدار 05 داخل یکی از رجیسترهای مجازی ذخیره میشه
بعد کنترل دوباره برمیگرده به Dispatcher
این بار Dispatcher مقدار 02 رو میخونه
Handler
مربوط به ADD اجرا میشه و عدد 03 به مقدار قبلی اضافه میشه
دوباره کنترل برمیگرده به Dispatcher
در آخر Opcode 03 اجرا میشه
Handler
مربوط به PRINT مقدار نهایی رو چاپ میکنه
پس همیشه این چرخه تکرار میشه:
Dispatcher ↓ Handler ↓ Dispatcher ↓ Handler ↓ Dispatcher
نکته مهم اینجاست که همه Handler ها شبیه هم نیستن
بعضی ها فقط یک مقدار رو جا به جا میکنن
بعضی ها عملیات ریاضی انجام میدن
بعضی ها پرش شرطی انجام میدن
بعضی ها حافظه رو میخونن یا مینویسن
در ماشین های مجازی واقعی مثل VMProtect ممکنه صدها Handler مختلف وجود داشته باشه
به همین خاطر اولین کاری که تحلیلگر انجام میده دستهبندی Handler هاست
مثلا بعد از بررسی چندتا از اونا متوجه میشه:
این Handler همیشه داده رو جا به جا میکنه
اون یکی همیشه عملیات XOR انجام میده
یکی دیگه همیشه مقدارها رو با هم مقایسه میکنه
یکی هم مسئول پرش های شرطیه
کم کم با کنار هم گذاشتن این اطلاعات معنی Opcode ها مشخص میشه
این دقیقا همون چیزیه که بهش Opcode Mapping میگن
یعنی ساختن یک جدول که مشخص کنه هر Opcode چه کاری انجام میده
تمرین
یک جدول برای Opcode های زیر درست کنید:
بعد سعی کنبد برای هر Opcode بنویسید Handler اون باید چکاری انجام بده
این تمرین باعث میشه ذهنتون دقیقا مثل یک تحلیلگر Devirtualization شروع به فکر کردن کنه
Handler
Where the real work is done
In the previous post, we said that Dispatcher only decides which Opcode to execute
But it doesn't do anything special
The main work is done inside Handlers
Each Opcode has at least one Handler
For example, suppose we have this bytecode:
If the meaning of the Opcodes is:
What happens is this:
Dispatcher
Sees the value 01 and the Handler corresponding to LOAD is executed
Inside the Handler, the value 05 is stored in one of the virtual registers
Then control returns to Dispatcher
This time Dispatcher reads the value 02
Handler
Related to ADD is executed and the number 03 is added to the previous value
Control returns to Dispatcher again
Finally, Opcode 03 is executed
Handler
Related to PRINT Prints the final value
So this cycle always repeats:
The important thing here is that not all Handlers are the same
Some just move a value
Some perform mathematical operations
Some perform conditional jumps
Some read or write memory.
In real virtual machines like VMProtect, there may be hundreds of different Handlers
That is why the first thing the analyst does is to categorize the Handlers
For example, after examining a few of them, he will notice:
This Handler always moves data
That one always performs XOR operations
Another always compares values
One is responsible for conditional jumps
Little by little, by putting this information together, the meaning of the Opcodes becomes clear
This is exactly what is called Opcode Mapping
That is, creating a table that specifies what each Opcode does
Exercise:
Create a table for the following Opcodes:
Then try to write for each Opcode what its Handler should do
This exercise will make your mind start thinking exactly like a Devirtualization analyst
@reverseengine
جایی که کار واقعی انجام میشه
تو پست قبل گفتیم Dispatcher فقط تصمیم میگیره کدوم Opcode اجرا بشه
اما خودش هیچ کار خاصی انجام نمیده
کار اصلی داخل Handler ها انجام میشه
هر Opcode حداقل یک Handler داره
مثلا فرض کنید این بایت کد رو داریم:
01 05 02 03 03
اگر معنی Opcode ها این باشه:
01 = LOAD
02 = ADD
03 = PRINT
اتفاقی که میوفته اینه:
Dispatcher
مقدار 01 رو میبینه و Handler مربوط به LOAD اجرا میشه
داخل Handler مقدار 05 داخل یکی از رجیسترهای مجازی ذخیره میشه
بعد کنترل دوباره برمیگرده به Dispatcher
این بار Dispatcher مقدار 02 رو میخونه
Handler
مربوط به ADD اجرا میشه و عدد 03 به مقدار قبلی اضافه میشه
دوباره کنترل برمیگرده به Dispatcher
در آخر Opcode 03 اجرا میشه
Handler
مربوط به PRINT مقدار نهایی رو چاپ میکنه
پس همیشه این چرخه تکرار میشه:
Dispatcher ↓ Handler ↓ Dispatcher ↓ Handler ↓ Dispatcher
نکته مهم اینجاست که همه Handler ها شبیه هم نیستن
بعضی ها فقط یک مقدار رو جا به جا میکنن
بعضی ها عملیات ریاضی انجام میدن
بعضی ها پرش شرطی انجام میدن
بعضی ها حافظه رو میخونن یا مینویسن
در ماشین های مجازی واقعی مثل VMProtect ممکنه صدها Handler مختلف وجود داشته باشه
به همین خاطر اولین کاری که تحلیلگر انجام میده دستهبندی Handler هاست
مثلا بعد از بررسی چندتا از اونا متوجه میشه:
این Handler همیشه داده رو جا به جا میکنه
اون یکی همیشه عملیات XOR انجام میده
یکی دیگه همیشه مقدارها رو با هم مقایسه میکنه
یکی هم مسئول پرش های شرطیه
کم کم با کنار هم گذاشتن این اطلاعات معنی Opcode ها مشخص میشه
این دقیقا همون چیزیه که بهش Opcode Mapping میگن
یعنی ساختن یک جدول که مشخص کنه هر Opcode چه کاری انجام میده
تمرین
یک جدول برای Opcode های زیر درست کنید:
01 → LOAD 02 → ADD 03 → XOR 04 → CMP 05 → JMP 06 → PRINT
بعد سعی کنبد برای هر Opcode بنویسید Handler اون باید چکاری انجام بده
این تمرین باعث میشه ذهنتون دقیقا مثل یک تحلیلگر Devirtualization شروع به فکر کردن کنه
Handler
Where the real work is done
In the previous post, we said that Dispatcher only decides which Opcode to execute
But it doesn't do anything special
The main work is done inside Handlers
Each Opcode has at least one Handler
For example, suppose we have this bytecode:
01 05 02 03 03
If the meaning of the Opcodes is:
01 = LOAD
02 = ADD
03 = PRINT
What happens is this:
Dispatcher
Sees the value 01 and the Handler corresponding to LOAD is executed
Inside the Handler, the value 05 is stored in one of the virtual registers
Then control returns to Dispatcher
This time Dispatcher reads the value 02
Handler
Related to ADD is executed and the number 03 is added to the previous value
Control returns to Dispatcher again
Finally, Opcode 03 is executed
Handler
Related to PRINT Prints the final value
So this cycle always repeats:
Dispatcher ↓ Handler ↓ Dispatcher ↓ Handler ↓ Dispatcher
The important thing here is that not all Handlers are the same
Some just move a value
Some perform mathematical operations
Some perform conditional jumps
Some read or write memory.
In real virtual machines like VMProtect, there may be hundreds of different Handlers
That is why the first thing the analyst does is to categorize the Handlers
For example, after examining a few of them, he will notice:
This Handler always moves data
That one always performs XOR operations
Another always compares values
One is responsible for conditional jumps
Little by little, by putting this information together, the meaning of the Opcodes becomes clear
This is exactly what is called Opcode Mapping
That is, creating a table that specifies what each Opcode does
Exercise:
Create a table for the following Opcodes:
01 → LOAD 02 → ADD 03 → XOR 04 → CMP 05 → JMP 06 → PRINT
Then try to write for each Opcode what its Handler should do
This exercise will make your mind start thinking exactly like a Devirtualization analyst
@reverseengine
وقتی روی یک برنامه دوبار کلیک میکنید پشت صحنه چه اتفاقی میوفته؟
تا اینجا فهمیدیم Process چیه و داخلش چه بخش هایی وجود داره
حالا یک سوال جالب:
وقتی روی notepad.exe یا chrome.exe دوبار کلیک میکنید دقیقا چه اتفاقی میوفته؟
همه این مراحل در چند میلی ثانیه انجام میشن اما پشت صحنه کارهای زیادی انجام میشه
مرحله 1: پیدا کردن فایل اجرایی
اول سیستمعامل فایل برنامه را روی دیسک پیدا میکنه
مثلا:
در این مرحله هنوز هیچ کدی اجرا نشده فقط فایل پیدا شده
مرحله 2: ساختن Process
حالا کرنل یک Process جدید ایجاد میکنه
برای این Process اطلاعاتی مثل:
رو آماده میکنه
از این لحظه سیستم عامل میتونه این برنامه رو مدیریت کنه
مرحله 3: ایجاد فضای حافظه
حالا سیستم عامل یک Virtual Address Space برای Process میسازه
یعنی یک فضای حافظه اختصاصی که فقط متعلق به همین Process هست.
در این فضا بخش هایی مثل:
قرار میگیرن
نکته مهم اینجاست که هر Process فضای حافظه مخصوص خودش رو داره و معمولا نمیتونه مستقیما به حافظه Process دیگه ای دسترسی پیدا کنه
مرحله 4: بارگذاری فایل اجرایی
حالا سیستم عامل فایل اجرایی رو داخل حافظه اپلود میکنه
بخش های مختلف فایل مثل کد و دادهها در جای مناسب خودشون قرار میگیرن
اگر برنامه به کتابخانه هایی مثل DLL نیاز داشته باشه اونا هم در همین مرحله اپلود میشن
مرحله 5: آماده شدن Thread اصلی
هر Process حداقل یک Thread داره
که به اون Main Thread میگن
این Thread از اولین دستور برنامه شروع به اجرا مبکنه بعدا اگر برنامه لازم داشته باشه میتونه Thread های بیشتری بسازه
مرحله 6: شروع اجرای برنامه
حالا CPU اجرای اولین دستور برنامه رو شروع میکنه
از این لحظه به بعد برنامه واقعا در حال اجراست
یک مثال واقعی:
فرض کنید روی Google Chrome کلیک میکنید
در چند لحظه:
✅ فایل Chrome پیدا میشه
✅ یک Process جدید ساخته میشه
✅ حافظه اختصاصی برای اون ایجاد میشه
✅ فایل اجرایی و DLL های مورد نیاز اپلود میشن
✅ Main Thread ساخته میشه
✅ اولین دستور Chrome اجرا میشه
همه این مراحل اونقدر سریع انجام میشن که کاربر فقط باز شدن پنجره مرورگر رو میبینه
چرا این موضوع برای مهندسی معکوس مهمه؟
وقتی در ابزارهایی مثل دیباگر یا تحلیل بدافزار کار میکنید باید بدونید برنامه از لحظه اجرا چه مراحلی رو طی کرده
برای مثال بیشتر بدافزارها:
قبل از اجرای کد اصلی DLL های خاصی رو اپلود میکنن
در همون ابتدای اجرا چند Thread جدید میسازن
حافظه جدید اختصاص میدن و کد خودشون رو در اون قرار میدن
اگر فرآیند ایجاد Process رو بشناسید تحلیل چنین رفتار هایی بسیار ساده تر میشه
What happens behind the scenes when you double-click on a program?
So far we have understood what a Process is and what parts it contains
Now an interesting question:
What exactly happens when you double-click on notepad.exe or chrome.exe?
All these steps are done in a few milliseconds, but a lot of work is done behind the scenes
Step 1: Finding the executable file
First, the operating system finds the program file on the disk
For example:
At this stage, no code is executed yet, only the file is found
Step 2: Creating a Process
Now the kernel creates a new Process
For this Process, it provides information such as:
From this moment, the operating system can manage this program
Step 3: Creating a memory space
Now the operating system creates a Virtual Address Space for the Process
That is, a dedicated memory space that belongs only to this Process.
In this space, sections such as:
The important point here is that each Process has its own memory space and usually cannot directly access the memory of another Process
Step 4: Loading the executable file
Now the operating system uploads the executable file into memory
Different parts of the file such as code and data are placed in their appropriate places
If the program needs libraries such as DLLs, they are also uploaded at this stage
Step 5: Preparing the main thread
Each process has at least one thread
which is called the main thread
تا اینجا فهمیدیم Process چیه و داخلش چه بخش هایی وجود داره
حالا یک سوال جالب:
وقتی روی notepad.exe یا chrome.exe دوبار کلیک میکنید دقیقا چه اتفاقی میوفته؟
همه این مراحل در چند میلی ثانیه انجام میشن اما پشت صحنه کارهای زیادی انجام میشه
مرحله 1: پیدا کردن فایل اجرایی
اول سیستمعامل فایل برنامه را روی دیسک پیدا میکنه
مثلا:
C:\Windows\System32\notepad.exe
در این مرحله هنوز هیچ کدی اجرا نشده فقط فایل پیدا شده
مرحله 2: ساختن Process
حالا کرنل یک Process جدید ایجاد میکنه
برای این Process اطلاعاتی مثل:
شناسه پردازه (PID)
وضعیت پردازه
اطلاعات زمان بندی
مجوزهای دسترسی
رو آماده میکنه
از این لحظه سیستم عامل میتونه این برنامه رو مدیریت کنه
مرحله 3: ایجاد فضای حافظه
حالا سیستم عامل یک Virtual Address Space برای Process میسازه
یعنی یک فضای حافظه اختصاصی که فقط متعلق به همین Process هست.
در این فضا بخش هایی مثل:
Code
Data
Heap
Stack
قرار میگیرن
نکته مهم اینجاست که هر Process فضای حافظه مخصوص خودش رو داره و معمولا نمیتونه مستقیما به حافظه Process دیگه ای دسترسی پیدا کنه
مرحله 4: بارگذاری فایل اجرایی
حالا سیستم عامل فایل اجرایی رو داخل حافظه اپلود میکنه
بخش های مختلف فایل مثل کد و دادهها در جای مناسب خودشون قرار میگیرن
اگر برنامه به کتابخانه هایی مثل DLL نیاز داشته باشه اونا هم در همین مرحله اپلود میشن
مرحله 5: آماده شدن Thread اصلی
هر Process حداقل یک Thread داره
که به اون Main Thread میگن
این Thread از اولین دستور برنامه شروع به اجرا مبکنه بعدا اگر برنامه لازم داشته باشه میتونه Thread های بیشتری بسازه
مرحله 6: شروع اجرای برنامه
حالا CPU اجرای اولین دستور برنامه رو شروع میکنه
از این لحظه به بعد برنامه واقعا در حال اجراست
یک مثال واقعی:
فرض کنید روی Google Chrome کلیک میکنید
در چند لحظه:
✅ فایل Chrome پیدا میشه
✅ یک Process جدید ساخته میشه
✅ حافظه اختصاصی برای اون ایجاد میشه
✅ فایل اجرایی و DLL های مورد نیاز اپلود میشن
✅ Main Thread ساخته میشه
✅ اولین دستور Chrome اجرا میشه
همه این مراحل اونقدر سریع انجام میشن که کاربر فقط باز شدن پنجره مرورگر رو میبینه
چرا این موضوع برای مهندسی معکوس مهمه؟
وقتی در ابزارهایی مثل دیباگر یا تحلیل بدافزار کار میکنید باید بدونید برنامه از لحظه اجرا چه مراحلی رو طی کرده
برای مثال بیشتر بدافزارها:
قبل از اجرای کد اصلی DLL های خاصی رو اپلود میکنن
در همون ابتدای اجرا چند Thread جدید میسازن
حافظه جدید اختصاص میدن و کد خودشون رو در اون قرار میدن
اگر فرآیند ایجاد Process رو بشناسید تحلیل چنین رفتار هایی بسیار ساده تر میشه
What happens behind the scenes when you double-click on a program?
So far we have understood what a Process is and what parts it contains
Now an interesting question:
What exactly happens when you double-click on notepad.exe or chrome.exe?
All these steps are done in a few milliseconds, but a lot of work is done behind the scenes
Step 1: Finding the executable file
First, the operating system finds the program file on the disk
For example:
C:\Windows\System32\notepad.exe
At this stage, no code is executed yet, only the file is found
Step 2: Creating a Process
Now the kernel creates a new Process
For this Process, it provides information such as:
Process ID (PID)
Process status
Scheduling information
Access permissions
From this moment, the operating system can manage this program
Step 3: Creating a memory space
Now the operating system creates a Virtual Address Space for the Process
That is, a dedicated memory space that belongs only to this Process.
In this space, sections such as:
Code
Data
Heap
Stack
The important point here is that each Process has its own memory space and usually cannot directly access the memory of another Process
Step 4: Loading the executable file
Now the operating system uploads the executable file into memory
Different parts of the file such as code and data are placed in their appropriate places
If the program needs libraries such as DLLs, they are also uploaded at this stage
Step 5: Preparing the main thread
Each process has at least one thread
which is called the main thread
This thread starts executing from the first command of the program, and later if the program needs it, it can create more threads
Step 6: Starting the program
Now the CPU starts executing the first command of the program
From this moment on, the program is really running
A real example:
Suppose you click on Google Chrome
In a few moments:
✅ The Chrome file is found
✅ A new process is created
✅ Dedicated memory is created for it
✅ The executable and required DLLs are uploaded
✅ The main thread is created
✅ The first Chrome command is executed
All these steps are done so quickly that the user only sees the browser window opening
Why is this important for reverse engineering?
When working in tools such as debuggers or malware analysis, you need to know what steps the program has taken since execution
For example, most malware:
Uploads specific DLLs before executing the main code
Creates several new threads at the very beginning of execution
Allocates new memory and places its code in it
If you understand the process of process creation, analyzing such behavior becomes much easier
@reverseengine
Step 6: Starting the program
Now the CPU starts executing the first command of the program
From this moment on, the program is really running
A real example:
Suppose you click on Google Chrome
In a few moments:
✅ The Chrome file is found
✅ A new process is created
✅ Dedicated memory is created for it
✅ The executable and required DLLs are uploaded
✅ The main thread is created
✅ The first Chrome command is executed
All these steps are done so quickly that the user only sees the browser window opening
Why is this important for reverse engineering?
When working in tools such as debuggers or malware analysis, you need to know what steps the program has taken since execution
For example, most malware:
Uploads specific DLLs before executing the main code
Creates several new threads at the very beginning of execution
Allocates new memory and places its code in it
If you understand the process of process creation, analyzing such behavior becomes much easier
@reverseengine
EDR
ها چطور حملات رو تشخیص میدن؟
خیلی ها فکر میکنن EDR فقط دنبال امضای فایل یا اسم API ها هستن ولی واقعیت اینه که نسل جدید EDR ها بیشتر روی رفتار تمرکز دارن
اگر یه برنامه این کارها رو انجام بده:
یک پروسه جدید ایجاد میکنه
حافظه قابل اجرا (Executable Memory) رزرو میکنه
داخل اون حافظه داده مینویسه
یک Thread جدید اجرا میکنه
ممکنه هیچ فایل مخربی روی دیسک وجود نداشته باشه اما همین زنجیره رفتارها میتونه امتیاز ریسک بالایی ایجاد کنه
EDR
ها به چه چیزایی دقت میکنن؟
چرا فقط یک رفتار کافی نیست؟
مثلا:
هر کدوم از اینها به تنهایی ممکنه طبیعی باشن اما وقتی پشت سر هم باشن برای EDR یک الگوی مشکوک ایجاد میکنن
تشخیص دادن
نسل جدید EDR ها معمولا از مدل های رفتاری استفاده میکنن نه فقط Signature
به همین دلیل ممکنه دو فایل کاملا متفاوت اگر رفتار مشابهی داشته باشن هر دو شناسایی بشن
How EDRs Detect Attacks?
Many people think that EDRs only look for file signatures or API names, but the truth is that the new generation of EDRs focuses more on behavior
If a program does these things:
There may not be any malicious files on disk, but this chain of behavior can generate a high risk score
What do EDRs look for?
Why is just one behavior not enough?
For example:
Each of these may be normal on their own, but when they are in a row, they create a suspicious pattern for EDR
Detection
Newer generation EDRs usually use behavioral models, not just Signature
That's why two completely different files may both be detected if they have similar behavior
@reverseengine
ها چطور حملات رو تشخیص میدن؟
خیلی ها فکر میکنن EDR فقط دنبال امضای فایل یا اسم API ها هستن ولی واقعیت اینه که نسل جدید EDR ها بیشتر روی رفتار تمرکز دارن
اگر یه برنامه این کارها رو انجام بده:
یک پروسه جدید ایجاد میکنه
حافظه قابل اجرا (Executable Memory) رزرو میکنه
داخل اون حافظه داده مینویسه
یک Thread جدید اجرا میکنه
ممکنه هیچ فایل مخربی روی دیسک وجود نداشته باشه اما همین زنجیره رفتارها میتونه امتیاز ریسک بالایی ایجاد کنه
EDR
ها به چه چیزایی دقت میکنن؟
ارتباط بین پروسهها (Process Tree)
ترتیب رخدادها (Event Correlation)
دسترسی های غیرعادی به حافظه
اپلود DLL های غیرمنتظره
تغییرات در Tokenها و دسترسیها
الگوهای ارتباط شبکه
اطلاعات تله متری ویندوز مثل ETW
چرا فقط یک رفتار کافی نیست؟
مثلا:
Process Created
↓
Memory Allocation
↓
Memory Write
↓
Thread Execution
هر کدوم از اینها به تنهایی ممکنه طبیعی باشن اما وقتی پشت سر هم باشن برای EDR یک الگوی مشکوک ایجاد میکنن
تشخیص دادن
نسل جدید EDR ها معمولا از مدل های رفتاری استفاده میکنن نه فقط Signature
به همین دلیل ممکنه دو فایل کاملا متفاوت اگر رفتار مشابهی داشته باشن هر دو شناسایی بشن
How EDRs Detect Attacks?
Many people think that EDRs only look for file signatures or API names, but the truth is that the new generation of EDRs focuses more on behavior
If a program does these things:
Creates a new process
Reserves executable memory
Writes data into that memory
Runs a new thread
There may not be any malicious files on disk, but this chain of behavior can generate a high risk score
What do EDRs look for?
Process Tree
Event Correlation
Unusual Memory Access
Unexpected DLL Uploads
Token Changes and Accesses
Network Communication Patterns
Windows Telemetry Information like ETW
Why is just one behavior not enough?
For example:
Process Created
↓
Memory Allocation
↓
Memory Write
↓
Thread Execution
Each of these may be normal on their own, but when they are in a row, they create a suspicious pattern for EDR
Detection
Newer generation EDRs usually use behavioral models, not just Signature
That's why two completely different files may both be detected if they have similar behavior
@reverseengine
🔥1
Accelerating EDR Evasion with LLM-Driven Analysis
https://specterops.io/blog/2026/06/29/llm-powered-edr-analysis/#
https://specterops.io/blog/2026/06/29/llm-powered-edr-analysis/#
https://specterops.io/blog/2026/06/29/llm-powered-edr-analysisSpecterOps
Accelerating EDR Evasion with LLM-Driven Analysis
SpecterOps reverse engineered Cortex XDR with LLMs to extract YARA rules, ML models, and behavioral detections.
Reverse Engineering the Simda Malware Loader: From Packed Binary to C2 Infrastructure
https://medium.com/@ckant/reverse-engineering-the-simda-malware-loader-from-packed-binary-to-c2-infrastructure-f2b738948b8a
https://medium.com/@ckant/reverse-engineering-the-simda-malware-loader-from-packed-binary-to-c2-infrastructure-f2b738948b8a
Medium
Reverse Engineering the Simda Malware Loader: From Packed Binary to C2 Infrastructure
An in-depth analysis of the Simda malware loader using static and dynamic analysis techniques to uncover payload decryption, anti-analysis…
👍5
بخش بیستم بافر اورفلو
Use After Free
یکی از خطرناکترین باگهای حافظه
تا اینجا یاد گرفتیم malloc حافظه میگیره و free اون رو آزاد میکنه حالا میخوایم ببینیم اگر برنامه بعد از free دوباره از همون حافظه استفاده کنه چه اتفاقی میوفته و چطور موقع مهندسی معکوس این باگ رو تشخیص بدیم
Use After Free
یعنی چی
اسمش کاملا مشخصه اول حافظه آزاد میشه
بعد برنامه دوباره از همون حافظه استفاده میکنه یعنی برنامه فکر میکنه حافظه هنوز معتبره در حالی که سیستم عامل یا مدیریت Heap ممکنه اون حافظه رو برای چیز دیگه ای استفاده کرده باشه
یک مثال ساده
#include <stdio.h> #include <stdlib.h> int main() { char *buf = malloc(32); free(buf); puts(buf); return 0; }
مشکل این کجاست؟
اینجا
free(buf);
حافظه آزاد شده ولی چند خط بعد
puts(buf);
از همون اشارهگر استفاده شده
این دقیقا یک Use After Free هست
چرا خطرناکه؟
چون بعد از free دیگه هیچ تضمینی وجود نداره که دادههای قبلی داخل اون حافظه باقی مونده باشن ممکنه داده تغییر کرده باشه حافظه دوباره به بخش دیگهای اختصاص داده شده باشه برنامه کرش کنه
موقع مهندسی معکوس دنبال چی بگردیم؟
اگر الگوی اینجوری دیدید
free(ptr);
و بعد چند خط پایینتر
ptr->field
یا
memcpy(ptr,...)
یا
printf("%s", ptr);
یا هر استفاده دیگه از ptr باید بررسی کنید که آیا همون اشارهگر بعد از free دوباره استفاده شده یا نه
یک مثال در دیکامپایلر
buf = malloc(64); /* ... */ free(buf); /* ... */ strcpy(buf, input);
همین چند خط برای مشکوک شدن کافیه
چطور از این باگ جلوگیری میکنن؟
یکی از سادهترین روشها اینه که بعد از free
اشارهگر رو NULL کنن
free(buf); buf = NULL;
حالا اگر جایی دوباره از buf استفاده بشه
اشکال خیلی زودتر مشخص میشه
free یعنی حافظه آزاد شده
بعد از free نباید از همون اشارهگر استفاده کرد
در مهندسی معکوس همیشه مسیر
malloc → free →
استفاده مجدد رو بررسی کنید
این یکی از الگوهای مهم برای پیدا کردن باگهای مدیریت حافظه است
تمرین:
یک برنامه ساده که از malloc و free استفاده میکنه داخل Ghidra یا IDA باز کنید و بررسی کنید آیا بعد از free دوباره از همون اشارهگر استفاده شده یا نه اگر اینجور الگویی پیدا کردید دلیلش رو تحلیل کنید و مشخص کنید آیا واقعا یک Use After Free است یا فقط ظاهرا اینطور به نظر میرسه
@reverseengine
❤2
ReverseEngineering
بخش بیستم بافر اورفلو Use After Free یکی از خطرناکترین باگهای حافظه تا اینجا یاد گرفتیم malloc حافظه میگیره و free اون رو آزاد میکنه حالا میخوایم ببینیم اگر برنامه بعد از free دوباره از همون حافظه استفاده کنه چه اتفاقی میوفته و چطور موقع مهندسی معکوس…
Part 20 Buffer Overflow
Use After Free One of the most dangerous memory bugs
So far we have learned that malloc takes memory and free frees it. Now we want to see what happens if the program uses the same memory again after free and how to detect this bug when reverse engineering
What does Use After Free
mean? Its name is quite specific. First the memory is freed. Then the program uses the same memory again, meaning the program thinks the memory is still valid, while the operating system or Heap Manager may have used that memory for something else.
A simple example:
#include <stdio.h> #include <stdlib.h>Where is the problem with this?
int main() { char *buf = malloc(32); free(buf); puts(buf); return 0; }
Here
free(buf);
The memory is freed, but a few lines later
puts(buf);
The same pointer is used. This is exactly a Use After Free
Why is it dangerous?
Because after free there is no guarantee that the previous data will remain in that memory. The data may have changed, the memory may have been reallocated, the program may crash
What should we look for when reverse engineering?
If you see a pattern like this
free(ptr);
and then a few lines down
ptr->field
or
memcpy(ptr,...)
or
printf("%s", ptr);
or any other use of ptr, you should check whether the same pointer is reused after free
An example in the decompiler
buf = malloc(64); /* ... */ free(buf); /* ... */ strcpy(buf, input);
These few lines are enough to be suspicious
How do you prevent this bug?
One of the easiest ways is to NULL the pointer after free
free(buf); buf = NULL;
Now if buf is used again somewhere
The problem will be identified much earlier
free means memory has been freed
You should not use the same pointer after free
In reverse engineering, always follow the path
malloc → free →
Check for reuse
This is one of the important patterns for finding memory management bugs
Exercise:
Open a simple program that uses malloc and free in Ghidra or IDA and check if the same pointer is used again after free or not. If you find such a pattern, analyze the reason and determine if it is really a Use After Free or it just looks like it.
@reverseengine
❤4
اینجا به یکی از مهمترین مفاهیم Binary Exploitation میرسیم
Info Leak
چرا اولین هدف یک اکسپلویت مدرنه؟
تا اینجا چند تا مکانیزم امنیتی رو شناختیم:
یا Info Leak یعنی چی؟
Info Leak
یعنی برنامه بدون اینکه قرار بوده بخشی از اطلاعات حافظه رو در اختیار کاربر قرار بده
این اطلاعات ممکنه شامل مواردی مثل:
یا حتی داده های حساس داخل حافظه
باشه در ظاهر شاید این اطلاعات مهم به نظر نرسن اما برای یک اکسپلویتر میتونن حکم نقشه ی یک ساختمون رو داشته باشن
چرا اینقدر مهمه؟
فرض کنید ASLR فعاله
یعنی آدرسها در هر بار اجرای برنامه تغییر میکنن
اگر هیچ آدرسی رو ندونید ساختن یک اکسپلویت قابل اعتماد خیلی سخت میشه
اما اگر برنامه فقط یک آدرس رو ناخواسته لو بده مهاجم میتونه از همون آدرس موقعیت بقیه بخش های حافظه رو هم محاسبه کنه
در این حالت بخش بزرگی از مزیت ASLR از بین میره
گاهی اوقات Info Leak خودش یک آسیبپذیری مستقل محسوب میشه و گاهی هم نتیجه ی یک باگ دیگه است
مثلا ممکنه:
برنامه بیشتر از حد لازم داده چاپ کنه
یک اشارهگر (Pointer) رو مستقیم نمایش بده
بخشی از حافظه رو بدون پاک کردن برگردونه
یا یک خطای منطقی باعث افشای اطلاعات بشه
چرا بیشتر اکسپلویتهای امروزی دو مرحلهای هستند؟
خیلی از حملات مدرن این شکلیاند:
مرحله اول پیدا کردن یک Info Leak
مرحله دوم استفاده از اطلاعات بهدست اومده برای دور زدن مکانیزم هایی مثل ASLR و ادامه ی حمله
به همین دلیل تعداد زیادی از تحلیلهای امنیتی که میبینید که اولین هدف نه اجرای کد بلکه پیدا کردن یک راه برای دیدن حافظه است
یک شبیه سازی ساده:
فرض کنید وارد یک شهر ناشناس شدید
اگر هیچ نقشه ای نداشته باشید پیدا کردن یک ساختمون خاص خیلی سخت میشه
اما اگر فقط آدرس یک خیابون رو بهتون بدن کم کم میتونید کل نقشه شهر رو پیدا کنید
Info Leak
دقیقا همین نقش رو در حافظه برنامهها داره
Info Leak
چرا اولین هدف یک اکسپلویت مدرنه؟
تا اینجا چند تا مکانیزم امنیتی رو شناختیم:
Stack CanaryInformation Leak
NX
ASLR
یا Info Leak یعنی چی؟
Info Leak
یعنی برنامه بدون اینکه قرار بوده بخشی از اطلاعات حافظه رو در اختیار کاربر قرار بده
این اطلاعات ممکنه شامل مواردی مثل:
آدرس یک تابع
آدرس یک متغیر
آدرس heap
آدرس stack
آدرس libc
یا حتی داده های حساس داخل حافظه
باشه در ظاهر شاید این اطلاعات مهم به نظر نرسن اما برای یک اکسپلویتر میتونن حکم نقشه ی یک ساختمون رو داشته باشن
چرا اینقدر مهمه؟
فرض کنید ASLR فعاله
یعنی آدرسها در هر بار اجرای برنامه تغییر میکنن
اگر هیچ آدرسی رو ندونید ساختن یک اکسپلویت قابل اعتماد خیلی سخت میشه
اما اگر برنامه فقط یک آدرس رو ناخواسته لو بده مهاجم میتونه از همون آدرس موقعیت بقیه بخش های حافظه رو هم محاسبه کنه
در این حالت بخش بزرگی از مزیت ASLR از بین میره
گاهی اوقات Info Leak خودش یک آسیبپذیری مستقل محسوب میشه و گاهی هم نتیجه ی یک باگ دیگه است
مثلا ممکنه:
برنامه بیشتر از حد لازم داده چاپ کنه
یک اشارهگر (Pointer) رو مستقیم نمایش بده
بخشی از حافظه رو بدون پاک کردن برگردونه
یا یک خطای منطقی باعث افشای اطلاعات بشه
چرا بیشتر اکسپلویتهای امروزی دو مرحلهای هستند؟
خیلی از حملات مدرن این شکلیاند:
مرحله اول پیدا کردن یک Info Leak
مرحله دوم استفاده از اطلاعات بهدست اومده برای دور زدن مکانیزم هایی مثل ASLR و ادامه ی حمله
به همین دلیل تعداد زیادی از تحلیلهای امنیتی که میبینید که اولین هدف نه اجرای کد بلکه پیدا کردن یک راه برای دیدن حافظه است
یک شبیه سازی ساده:
فرض کنید وارد یک شهر ناشناس شدید
اگر هیچ نقشه ای نداشته باشید پیدا کردن یک ساختمون خاص خیلی سخت میشه
اما اگر فقط آدرس یک خیابون رو بهتون بدن کم کم میتونید کل نقشه شهر رو پیدا کنید
Info Leak
دقیقا همین نقش رو در حافظه برنامهها داره
مکانیزم هایی مثل ASLR آدرس ها رو مخفی میکنن اما اگر برنامه حتی مقدار کمی از اطلاعات حافظه رو افشا کنه این محافظت تا حد زیادی بی اثر میشه به همین دلیل پیدا کردن Info Leak یکی از مهم ترین مراحل در تحلیل و درک اکسپلویت های مدرنه@reverseengine
ReverseEngineering
اینجا به یکی از مهمترین مفاهیم Binary Exploitation میرسیم Info Leak چرا اولین هدف یک اکسپلویت مدرنه؟ تا اینجا چند تا مکانیزم امنیتی رو شناختیم: Stack Canary NX ASLR Information Leak یا Info Leak یعنی چی؟ Info Leak یعنی برنامه بدون اینکه قرار بوده…
Here we come to one of the most important concepts of Binary Exploitation
Info Leak
Why is it the first target of a modern exploit?
So far we have recognized several security mechanisms:
Stack Canary
NX
ASLR
What is Information Leak or Info Leak?
Info Leak
means that the program provides some memory information to the user without being supposed to
This information may include things like:
Or even sensitive data in memory
On the surface, this information may not seem important, but for an exploiter, it can be like a blueprint of a building
Why is it so important?
Let's say ASLR is enabled
This means that the addresses change every time the program is run
If you don't know any addresses, it's very difficult to create a reliable exploit
But if the program accidentally leaks just one address, the attacker can calculate the location of other memory locations from that address
In this case, a large part of the benefit of ASLR is lost
Sometimes an Info Leak is a standalone vulnerability, and sometimes it's the result of another bug
For example, it's possible to:
The program prints more data than necessary
Display a pointer directly
Return a portion of memory without clearing
Or a logical error causes information to be leaked
Why are most exploits today two-step?
Many modern attacks look like this:
The first step is to find an Info Leak
The second step is to use the information obtained to bypass mechanisms such as ASLR and continue the attack
That is why a lot of security analysis that you see the first goal is not to execute code but to find a way to see the memory
A simple simulation:
Suppose you enter an unknown city
If you have no map, it will be very difficult to find a specific building
But if you are only given the address of a street, you can gradually find the entire map of the city
Info Leak
It plays exactly the same role in the memory of programs
Info Leak
Why is it the first target of a modern exploit?
So far we have recognized several security mechanisms:
Stack Canary
NX
ASLR
What is Information Leak or Info Leak?
Info Leak
means that the program provides some memory information to the user without being supposed to
This information may include things like:
The address of a function
The address of a variable
The heap address
The stack address
The libc address
Or even sensitive data in memory
On the surface, this information may not seem important, but for an exploiter, it can be like a blueprint of a building
Why is it so important?
Let's say ASLR is enabled
This means that the addresses change every time the program is run
If you don't know any addresses, it's very difficult to create a reliable exploit
But if the program accidentally leaks just one address, the attacker can calculate the location of other memory locations from that address
In this case, a large part of the benefit of ASLR is lost
Sometimes an Info Leak is a standalone vulnerability, and sometimes it's the result of another bug
For example, it's possible to:
The program prints more data than necessary
Display a pointer directly
Return a portion of memory without clearing
Or a logical error causes information to be leaked
Why are most exploits today two-step?
Many modern attacks look like this:
The first step is to find an Info Leak
The second step is to use the information obtained to bypass mechanisms such as ASLR and continue the attack
That is why a lot of security analysis that you see the first goal is not to execute code but to find a way to see the memory
A simple simulation:
Suppose you enter an unknown city
If you have no map, it will be very difficult to find a specific building
But if you are only given the address of a street, you can gradually find the entire map of the city
Info Leak
It plays exactly the same role in the memory of programs
Mechanisms such as ASLR hide addresses, but if the program reveals even a small amount of memory information, this protection becomes largely ineffective. This is why finding Info Leak is one of the most important steps in analyzing and understanding modern exploits@reverseengine
Virtual Register رجیستر های مجازی
تا اینجا فهمیدیم Dispatcher تصمیم میگیره کدوم Handler اجرا بشه و Handler هم کار اصلی رو انجام میده
Handler
ها داده ها رو کجا نگه میدارن؟
روی رجیستر های واقعی CPU مثل RAX و RBX؟ معمولا نه
بیشتر ماشینهای مجازی از چیزی به اسم Virtual Register استفاده میکنن
Virtual Register
یه فضای حافظه ست که نقش رجیستر های CPU رو بازی میکنه
مثلا فرض کنید ماشین مجازی 8 تا رجیستر داره:
وقتی Opcode زیر اجرا میشه:
مقدار 10 داخل V0 قرار میگیره
بعد Opcode بعدی:
عدد 20 داخل V1 ذخیره میشه
حالا Opcode بعدی:
ماشین مجازی مقدار V0 و V1 رو جمع میکنه
نتیجه دوباره داخل V0 ذخیره میشه
اگر بعدش این Opcode اجرا بشه:
عدد 30 چاپ میشه
اینجا هیچکدوم از رجیسترهای واقعی CPU مستقیما دیده نمیشن
ممکنه داخل Handlerه ا فقط چند دستور مثل این ببینید:
در نگاه اول انگار برنامه فقط داره با حافظه کار میکنه
ولی در واقع:
یعنی این آدرس های حافظه همون رجیستر های مجازی هستن به همین دلیل یکی از اولین کارهای تحلیلگر اینه که محل نگهداری Virtual Register ها رو پیدا کنه
وقتی بفهمید هر Offset مربوط به کدوم Virtual Register هست خوندن Handler ها چند برابر راحت تر میشه
خیلی وقت ها حتی اسمگذاری هم میکنن:
از این لحظه به بعد به جای آدرس های حافظه ذهن تحلیلگر با رجیسترهای مجازی کار میکنه این کار باعث میشه منطق ماشین مجازی کم کم شبیه یک CPU معمولی به نظر برسه
تمرین:
فرض کنید یک ماشین مجازی چهار رجیستر داره:
و Opcode های زیر اجرا میشن:
بدون اینکه کدی بنویسید مرحله به مرحله مشخص کنید بعد از اجرای هر Opcode مقدار هر Virtual Register چقدر میشه
Virtual Register So far we have understood that Dispatcher decides which Handler to execute and Handler does the main work
Where do Handlers store data?
On real CPU registers like RAX and RBX? Usually not
Most virtual machines use something called Virtual Register
Virtual Register
is a memory space that acts as CPU registers
For example, suppose the virtual machine has 8 registers:
When the following Opcode is executed:
10 is placed into V0
Then the next Opcode:
20 is stored into V1
Now the next Opcode:
The virtual machine adds V0 and V1
The result is stored back into V0
If this Opcode is then executed:
30 is printed
Here none of the actual CPU registers are directly visible
You may only see a few instructions inside the Handlers like this:
At first glance, it seems like the program is only working with memory
But in fact:
That is, these memory addresses are the same virtual registers, which is why one of the first tasks of the analyst is to find the location of the Virtual Registers.
When you understand which Virtual Register each Offset belongs to, reading the Handlers becomes much easier.
Many times they even give them names:
From this moment on, instead of memory addresses, the analyst's mind works with virtual registers. This makes the logic of the virtual machine look a little like a regular CPU.
Exercise:
Suppose a virtual machine It has four registers:
And the following Opcodes are executed:
Without writing any code, determine step by step what the value of each Virtual Register will be after executing each Opcode.
@reverseengine
تا اینجا فهمیدیم Dispatcher تصمیم میگیره کدوم Handler اجرا بشه و Handler هم کار اصلی رو انجام میده
Handler
ها داده ها رو کجا نگه میدارن؟
روی رجیستر های واقعی CPU مثل RAX و RBX؟ معمولا نه
بیشتر ماشینهای مجازی از چیزی به اسم Virtual Register استفاده میکنن
Virtual Register
یه فضای حافظه ست که نقش رجیستر های CPU رو بازی میکنه
مثلا فرض کنید ماشین مجازی 8 تا رجیستر داره:
V0 V1 V2 V3 V4 V5 V6 V7
وقتی Opcode زیر اجرا میشه:
LOAD V0, 10
مقدار 10 داخل V0 قرار میگیره
بعد Opcode بعدی:
LOAD V1, 20
عدد 20 داخل V1 ذخیره میشه
حالا Opcode بعدی:
ADD V0, V1
ماشین مجازی مقدار V0 و V1 رو جمع میکنه
نتیجه دوباره داخل V0 ذخیره میشه
اگر بعدش این Opcode اجرا بشه:
PRINT V0
عدد 30 چاپ میشه
اینجا هیچکدوم از رجیسترهای واقعی CPU مستقیما دیده نمیشن
ممکنه داخل Handlerه ا فقط چند دستور مثل این ببینید:
mov rax, [rdi+18h]
mov rcx, [rdi+20h]
add rax, rcx
mov [rdi+18h], raxدر نگاه اول انگار برنامه فقط داره با حافظه کار میکنه
ولی در واقع:
[rdi+18h] = V0 [rdi+20h] = V1 یعنی این آدرس های حافظه همون رجیستر های مجازی هستن به همین دلیل یکی از اولین کارهای تحلیلگر اینه که محل نگهداری Virtual Register ها رو پیدا کنه
وقتی بفهمید هر Offset مربوط به کدوم Virtual Register هست خوندن Handler ها چند برابر راحت تر میشه
خیلی وقت ها حتی اسمگذاری هم میکنن:
[rdi+18h] → V0 [rdi+20h] → V1 [rdi+28h] → V2 از این لحظه به بعد به جای آدرس های حافظه ذهن تحلیلگر با رجیسترهای مجازی کار میکنه این کار باعث میشه منطق ماشین مجازی کم کم شبیه یک CPU معمولی به نظر برسه
تمرین:
فرض کنید یک ماشین مجازی چهار رجیستر داره:
V0 V1 V2 V3
و Opcode های زیر اجرا میشن:
LOAD V0, 15
LOAD V1, 5
SUB V0, V1
PRINT V0بدون اینکه کدی بنویسید مرحله به مرحله مشخص کنید بعد از اجرای هر Opcode مقدار هر Virtual Register چقدر میشه
Virtual Register So far we have understood that Dispatcher decides which Handler to execute and Handler does the main work
Where do Handlers store data?
On real CPU registers like RAX and RBX? Usually not
Most virtual machines use something called Virtual Register
Virtual Register
is a memory space that acts as CPU registers
For example, suppose the virtual machine has 8 registers:
V0 V1 V2 V3 V4 V5 V6 V7
When the following Opcode is executed:
LOAD V0, 10
10 is placed into V0
Then the next Opcode:
LOAD V1, 20
20 is stored into V1
Now the next Opcode:
ADD V0, V1
The virtual machine adds V0 and V1
The result is stored back into V0
If this Opcode is then executed:
PRINT V0
30 is printed
Here none of the actual CPU registers are directly visible
You may only see a few instructions inside the Handlers like this:
mov rax, [rdi+18h]
mov rcx, [rdi+20h]
add rax, rcx
mov [rdi+18h], rax
At first glance, it seems like the program is only working with memory
But in fact:
[rdi+18h] = V0 [rdi+20h] = V1
That is, these memory addresses are the same virtual registers, which is why one of the first tasks of the analyst is to find the location of the Virtual Registers.
When you understand which Virtual Register each Offset belongs to, reading the Handlers becomes much easier.
Many times they even give them names:
[rdi+18h] → V0 [rdi+20h] → V1 [rdi+28h] → V2
From this moment on, instead of memory addresses, the analyst's mind works with virtual registers. This makes the logic of the virtual machine look a little like a regular CPU.
Exercise:
Suppose a virtual machine It has four registers:
V0 V1 V2 V3
And the following Opcodes are executed:
LOAD V0, 15
LOAD V1, 5
SUB V0, V1
PRINT V0
Without writing any code, determine step by step what the value of each Virtual Register will be after executing each Opcode.
@reverseengine