Virtual Stack
حافظه ای که ماشین مجازی روی اون کار میکنه
تا اینجا با Dispatcher Handler و Virtual Register آشنا شدیم
اما همه ماشین های مجازی از رجیستر استفاده نمیکنن
بعضیها تقریبا تمام عملیاتشون رو روی یک Stack مجازی انجام میدن
اگر قبلا با اسمبلی کار کرده باشید احتمالا Stack واقعی CPU رو میشناسید
ماشینهای مجازی هم دقیقا همین ایده رو پیاده میکنن با این تفاوت که Stack خودشون رو داخل حافظه میسازن
فرض کنید Stack مجازی در ابتدا خالی باشه
اولین Opcode اجرا میشه:
حالا Stack این شکلیه:
بعد:
Stack:
حالا Opcode بعدی:
Handler
مربوط به ADD دو مقدار بالای Stack رو برمیداره
اونها رو با هم جمع میکنه
یعنی 30 دوباره روی Stack قرار میگیره
Stack
حالا این شکلیه:
بعد:
Handler
مقدار بالای Stack رو میخونه و چاپ میکنه
خیلی از ماشین های مجازی واقعی هم تقریبا با همین منطق کار میکنن
به جای اینکه Opcode ها بنویسن:
مینویسن:
چون طراحی Stack-Based معمولا ساده تره و تولید بایت کد براش راحت تره
وقتی دارید یک VM رو تحلیل میکنید یکی از اولین سوال هایی که باید از خودتون بپرسید اینه:
این VM رجیستر محوره یا Stack-Based؟
جواب این سوال مسیر ادامه تحلیل رو مشخص میکنه
اگر مدام میبینید Handler ها دادهها رو از یک بافر مشخص برمیدارن مقدار جدید داخل همون بافر قرار میدن و یک اشاره گر مدام بالا و پایین میره احتمال زیادی وجود داره که با یک Virtual Stack طرف باشید
در مقابل اگر بیشتر عملیات روی چند خونه ثابت حافظه انجام میشه احتمالا VM از Virtual Register استفاده میکنه
یکی از مهارتهای مهم در Devirtualization اینه که خیلی زود تشخیص بدید معماری ماشین مجازی از کدوم نوعه
این کار باعث میشه ساعت ها وقتتون صرف تحلیل اشتباه نشه
تمرین:
فرض کنید Stack مجازی در ابتدا خالیه و این Opcode ها اجرا میشن:
Virtual Stack The memory that the virtual machine works on
So far we have been introduced to the Dispatcher Handler and Virtual Register
But not all virtual machines use registers
Some perform almost all their operations on a virtual stack
If you have worked with assembly before, you probably know the real CPU stack
Virtual machines implement exactly the same idea, except that they create their own stack in memory
Assume that the virtual stack is initially empty
The first Opcode is executed:
Now the Stack looks like this:
Next:
Stack:
Now the next Opcode:
The Handler
removes the top two values of the Stack
Adds them together
That is, 30 is placed back on the Stack
Stack
Now this Figure:
Next:
Handler
Reads and prints the top of the stack
Many real virtual machines work with almost the same logic
Instead of writing Opcodes:
They write:
Because Stack-Based design is usually simpler and bytecode generation is easier
When you are analyzing a VM, one of the first questions you should ask yourself is:
Is this VM register-based or stack-based?
The answer to this question will determine the path of further analysis
If you constantly see Handlers taking data from a specific buffer, putting a new value into the same buffer, and a pointer constantly moving up and down, there is a high probability that you are dealing with a Virtual Stack
On the other hand, if most of the operations are performed on a few fixed memory locations, the VM is probably using Virtual Registers
One of the important skills in Devirtualization is to quickly identify what type of virtual machine architecture you have
This will save you hours of time on incorrect analysis
Exercise:
Assume that the virtual stack is initially empty and these opcodes are executed:
Without running the program, write down the state of the stack after each opcode
@reverseengine
حافظه ای که ماشین مجازی روی اون کار میکنه
تا اینجا با Dispatcher Handler و Virtual Register آشنا شدیم
اما همه ماشین های مجازی از رجیستر استفاده نمیکنن
بعضیها تقریبا تمام عملیاتشون رو روی یک Stack مجازی انجام میدن
اگر قبلا با اسمبلی کار کرده باشید احتمالا Stack واقعی CPU رو میشناسید
ماشینهای مجازی هم دقیقا همین ایده رو پیاده میکنن با این تفاوت که Stack خودشون رو داخل حافظه میسازن
فرض کنید Stack مجازی در ابتدا خالی باشه
اولین Opcode اجرا میشه:
PUSH 10
حالا Stack این شکلیه:
10
بعد:
PUSH 20
Stack:
20
10
حالا Opcode بعدی:
ADD
Handler
مربوط به ADD دو مقدار بالای Stack رو برمیداره
20
10
اونها رو با هم جمع میکنه
یعنی 30 دوباره روی Stack قرار میگیره
Stack
حالا این شکلیه:
30
بعد:
Handler
مقدار بالای Stack رو میخونه و چاپ میکنه
خیلی از ماشین های مجازی واقعی هم تقریبا با همین منطق کار میکنن
به جای اینکه Opcode ها بنویسن:
ADD V0, V1
مینویسن:
PUSH V0
PUSH V1
ADD
چون طراحی Stack-Based معمولا ساده تره و تولید بایت کد براش راحت تره
وقتی دارید یک VM رو تحلیل میکنید یکی از اولین سوال هایی که باید از خودتون بپرسید اینه:
این VM رجیستر محوره یا Stack-Based؟
جواب این سوال مسیر ادامه تحلیل رو مشخص میکنه
اگر مدام میبینید Handler ها دادهها رو از یک بافر مشخص برمیدارن مقدار جدید داخل همون بافر قرار میدن و یک اشاره گر مدام بالا و پایین میره احتمال زیادی وجود داره که با یک Virtual Stack طرف باشید
در مقابل اگر بیشتر عملیات روی چند خونه ثابت حافظه انجام میشه احتمالا VM از Virtual Register استفاده میکنه
یکی از مهارتهای مهم در Devirtualization اینه که خیلی زود تشخیص بدید معماری ماشین مجازی از کدوم نوعه
این کار باعث میشه ساعت ها وقتتون صرف تحلیل اشتباه نشه
تمرین:
فرض کنید Stack مجازی در ابتدا خالیه و این Opcode ها اجرا میشن:
PUSH 8بدون اجرای برنامه روی کاغذ وضعیت Stack رو بعد از هر Opcode بنویسید
PUSH 12
ADD
PUSH 3
MUL
Virtual Stack The memory that the virtual machine works on
So far we have been introduced to the Dispatcher Handler and Virtual Register
But not all virtual machines use registers
Some perform almost all their operations on a virtual stack
If you have worked with assembly before, you probably know the real CPU stack
Virtual machines implement exactly the same idea, except that they create their own stack in memory
Assume that the virtual stack is initially empty
The first Opcode is executed:
PUSH 10
Now the Stack looks like this:
10
Next:
PUSH 20
Stack:
20
10
Now the next Opcode:
ADD
The Handler
removes the top two values of the Stack
20
10
Adds them together
That is, 30 is placed back on the Stack
Stack
Now this Figure:
30
Next:
Handler
Reads and prints the top of the stack
Many real virtual machines work with almost the same logic
Instead of writing Opcodes:
ADD V0, V1
They write:
PUSH V0
PUSH V1
ADD
Because Stack-Based design is usually simpler and bytecode generation is easier
When you are analyzing a VM, one of the first questions you should ask yourself is:
Is this VM register-based or stack-based?
The answer to this question will determine the path of further analysis
If you constantly see Handlers taking data from a specific buffer, putting a new value into the same buffer, and a pointer constantly moving up and down, there is a high probability that you are dealing with a Virtual Stack
On the other hand, if most of the operations are performed on a few fixed memory locations, the VM is probably using Virtual Registers
One of the important skills in Devirtualization is to quickly identify what type of virtual machine architecture you have
This will save you hours of time on incorrect analysis
Exercise:
Assume that the virtual stack is initially empty and these opcodes are executed:
PUSH 8
PUSH 12
ADD
PUSH 3
MUL
Without running the program, write down the state of the stack after each opcode
@reverseengine
یکی از مهمترین بخشهای Binary Exploitation
Heap هست
تا الان بیشتر دربارهی Stack و کنترل جریان اجرا صحبت کردیم اما خیلی از باگهای جدی امروزی داخل Heap اتفاق میوفتن
Heap
چیه و چرا در Binary Exploitation مهمه؟
وقتی یک برنامه اجرا میشه حافظه ی اون به چند بخش مختلف تقسیم میشه که دو بخش هستن:
Stack
بیشتر برای اطلاعات موقت توابع استفاده میشه اما Heap برای زمانیه که برنامه در زمان اجرا نیاز دارد حافظه ای رو خودش مدیریت کنه
Heap چیه؟
Heap
یک بخش از حافظه است که برنامه میتونه در زمان اجرا از اون درخواست حافظه کنه و بعدا اونو ازاد کنه
مثلا وقتی برنامه نمیدونه چقدر داده قراره دریافت کنه نمیتونه از یک فضای ثابت روی Stack استفاده کنه در این حالت معمولا سراغ Heap میروه
در زبانهایی مثل C و C++ مدیریت Heap معمولا با توابعی مثل:
انجام میشه
تفاوت Stack و Heap
Stack:
ساختار منظم تر و سریع تر داره
عمر دادهها معمولا وابسته به تابع هست
مدیریت اون بیشتر توسط کامپایلر انجام میشه
مثلا:
وقتی تابع تموم بشه این داده از بین میره
Heap:
توسط خود برنامه مدیریت میشه
دادهها میتونن مدت بیشتری باقی بمونن
برنامه خودش تصمیم میگیره چه زمانی حافظه بگیره و چه زمانی ازاد کنه
مثلا:
اینجا برنامه درخواست میکنه 100 بایت حافظه از Heap بگیره
چرا Heap در امنیت مهمه؟
چون مدیریت دستی حافظه مخصوصا در زبانهایی مثل C و C++ پیچیدگی زیادی داره
برنامهنویس باید همیشه مراقب باشه:
چه زمانی حافظه گرفته شده؟
چه زمانی آزاد شده؟
آیا هنوز از حافظه آزاد شده استفاده میشه؟
آیا اندازهی داده با اندازهی حافظه هماهنگه؟
یک اشتباه کوچیک میتونه باعث ایجاد آسیب پذیری بشه
چه نوع باگهایی در Heap دیده میشن؟
چند مورد معروف:
وقتی برنامه بیشتر از اندازهی اختصاص داده شده روی Heap داده مینویسه
وقتی برنامه حافظه ای رو آزاد میکنه اما بعدا همچنان از اون استفاده میکنه
وقتی برنامه یک بخش از حافظه رو بیشتر از یک بار آزاد میکنه
وقتی برنامه حافظه میگیره ولی هیچ وقت آزادش نمیکنه
چرا Heap سختتر از Stack است؟
Stack
ساختار نسبتا مشخصی دارده اما Heap پیچیده تره
چون Heap توسط یک Memory Allocator مدیریت میشه
مثلا در لینوکس یکی از allocator های معروف:
ptmalloc (در glibc)
هست
این allocator تصمیم میگیره:
چه بخشی از حافظه اختصاص داده بشه
کدام حافظه آزاد باشه
درخواست های جدید کجا قرار بگیرن
به همین دلیل تحلیل Heap نیاز به درک عمیق تری از مدیریت حافظه داره
چرا هکر ها یا مهندسان معکوس به Heap علاقه دارن؟
چون Heap معمولا شامل داده های مهم برنامه هست
مثلا:
ساختار های داده
Object
ها در ++C
اطلاعات session
pointer ها
داده های برنامه
اگر مدیریت Heap اشتباه باشه ممکنه باعث تغییر رفتار برنامه بشه
Heap
یکی از مهمترین بخش های حافظه در برنامه هاست که برای ذخیره سازی داده های پویا استفاده میشه
برخلاف Stack مدیریت Heap بیشتر بر عهده ی برنامه نویسه و همین موضوع باعث بهوجود اومدن باگهای پیچیدهای مثل Heap Overflow و Use-After-Free میشه
برای درک اکسپلویت های مدرن شناخت Heap و نحوهی کار Memory Allocator ها ضروریه
@reverseengine
Heap هست
تا الان بیشتر دربارهی Stack و کنترل جریان اجرا صحبت کردیم اما خیلی از باگهای جدی امروزی داخل Heap اتفاق میوفتن
Heap
چیه و چرا در Binary Exploitation مهمه؟
وقتی یک برنامه اجرا میشه حافظه ی اون به چند بخش مختلف تقسیم میشه که دو بخش هستن:
Stack
Heap
Stack
بیشتر برای اطلاعات موقت توابع استفاده میشه اما Heap برای زمانیه که برنامه در زمان اجرا نیاز دارد حافظه ای رو خودش مدیریت کنه
Heap چیه؟
Heap
یک بخش از حافظه است که برنامه میتونه در زمان اجرا از اون درخواست حافظه کنه و بعدا اونو ازاد کنه
مثلا وقتی برنامه نمیدونه چقدر داده قراره دریافت کنه نمیتونه از یک فضای ثابت روی Stack استفاده کنه در این حالت معمولا سراغ Heap میروه
در زبانهایی مثل C و C++ مدیریت Heap معمولا با توابعی مثل:
malloc()
calloc()
realloc()
free()
انجام میشه
تفاوت Stack و Heap
Stack:
ساختار منظم تر و سریع تر داره
عمر دادهها معمولا وابسته به تابع هست
مدیریت اون بیشتر توسط کامپایلر انجام میشه
مثلا:
void function(){
char buffer[64];
}
وقتی تابع تموم بشه این داده از بین میره
Heap:
توسط خود برنامه مدیریت میشه
دادهها میتونن مدت بیشتری باقی بمونن
برنامه خودش تصمیم میگیره چه زمانی حافظه بگیره و چه زمانی ازاد کنه
مثلا:
char *data = malloc(100);
اینجا برنامه درخواست میکنه 100 بایت حافظه از Heap بگیره
چرا Heap در امنیت مهمه؟
چون مدیریت دستی حافظه مخصوصا در زبانهایی مثل C و C++ پیچیدگی زیادی داره
برنامهنویس باید همیشه مراقب باشه:
چه زمانی حافظه گرفته شده؟
چه زمانی آزاد شده؟
آیا هنوز از حافظه آزاد شده استفاده میشه؟
آیا اندازهی داده با اندازهی حافظه هماهنگه؟
یک اشتباه کوچیک میتونه باعث ایجاد آسیب پذیری بشه
چه نوع باگهایی در Heap دیده میشن؟
چند مورد معروف:
Heap Overflow
وقتی برنامه بیشتر از اندازهی اختصاص داده شده روی Heap داده مینویسه
Use-After-Free (UAF)
وقتی برنامه حافظه ای رو آزاد میکنه اما بعدا همچنان از اون استفاده میکنه
Double Free
وقتی برنامه یک بخش از حافظه رو بیشتر از یک بار آزاد میکنه
Memory Leak
وقتی برنامه حافظه میگیره ولی هیچ وقت آزادش نمیکنه
چرا Heap سختتر از Stack است؟
Stack
ساختار نسبتا مشخصی دارده اما Heap پیچیده تره
چون Heap توسط یک Memory Allocator مدیریت میشه
مثلا در لینوکس یکی از allocator های معروف:
ptmalloc (در glibc)
هست
این allocator تصمیم میگیره:
چه بخشی از حافظه اختصاص داده بشه
کدام حافظه آزاد باشه
درخواست های جدید کجا قرار بگیرن
به همین دلیل تحلیل Heap نیاز به درک عمیق تری از مدیریت حافظه داره
چرا هکر ها یا مهندسان معکوس به Heap علاقه دارن؟
چون Heap معمولا شامل داده های مهم برنامه هست
مثلا:
ساختار های داده
Object
ها در ++C
اطلاعات session
pointer ها
داده های برنامه
اگر مدیریت Heap اشتباه باشه ممکنه باعث تغییر رفتار برنامه بشه
Heap
یکی از مهمترین بخش های حافظه در برنامه هاست که برای ذخیره سازی داده های پویا استفاده میشه
برخلاف Stack مدیریت Heap بیشتر بر عهده ی برنامه نویسه و همین موضوع باعث بهوجود اومدن باگهای پیچیدهای مثل Heap Overflow و Use-After-Free میشه
برای درک اکسپلویت های مدرن شناخت Heap و نحوهی کار Memory Allocator ها ضروریه
@reverseengine
ReverseEngineering
یکی از مهمترین بخشهای Binary Exploitation Heap هست تا الان بیشتر دربارهی Stack و کنترل جریان اجرا صحبت کردیم اما خیلی از باگهای جدی امروزی داخل Heap اتفاق میوفتن Heap چیه و چرا در Binary Exploitation مهمه؟ وقتی یک برنامه اجرا میشه حافظه ی اون…
One of the most important parts of Binary Exploitation Is the Heap
So far we have talked mostly about the Stack and execution flow control, but many of today's serious bugs occur in the Heap
What is the Heap and why is it important in Binary Exploitation?
When a program is executed, its memory is divided into several different parts, which are two parts:
Stack is mostly used for temporary information of functions, but the Heap is for when the program needs to manage memory itself at runtime
What is the Heap?
Heap is a part of memory that a program can request memory from at runtime and free it later
For example, when a program does not know how much data it is going to receive, it cannot use a fixed space on the Stack, in this case it usually goes to the Heap
In languages like C and C++, Heap management is usually done with functions like:
Stack:
It has a more organized and faster structure
The lifetime of the data is usually dependent on the function
Its management is mostly done by the compiler
For example:
When the function ends, this data is destroyed
Heap:
Managed by the program itself
Data can remain for a longer period
The program itself decides when to get memory and when to free it
For example:
Here the program requests to get 100 bytes of memory from the Heap
Why is the Heap important in security?
Because manual memory management is very complicated, especially in languages like C and C++
The programmer should always be careful:
When was the memory taken?
When was it freed?
Is the freed memory still in use?
Is the data size consistent with the memory size?
A small mistake can cause a vulnerability
What types of bugs are seen in the Heap?
Some famous cases:
When the program writes more data to the Heap than the allocated size
When the program frees memory but then continues to use it
When the program frees a section of memory more than once
When the program takes memory but never frees it
Why is the Heap harder than the Stack?
has a relatively well-defined structure, but the Heap is more complex
Because the Heap is managed by a Memory Allocator
For example, in Linux, one of the famous allocators is:
ptmalloc (in glibc)
This allocator decides:
What part of the memory to allocate
Which memory to free
Where new requests should be placed
That is why Heap analysis requires a deeper understanding of memory management
Why are hackers or reverse engineers interested in the Heap?
Because the Heap usually contains important program data
For example:
If the Heap management is wrong, it may change the behavior of the program.
The Heap is one of the most important memory areas in programs, used to store dynamic data.
Unlike the Stack, the Heap management is more the responsibility of the programmer, which leads to complex bugs such as Heap Overflow and Use-After-Free.
To understand modern exploits, it is essential to understand the Heap and how Memory Allocators work.
@reverseengine
So far we have talked mostly about the Stack and execution flow control, but many of today's serious bugs occur in the Heap
What is the Heap and why is it important in Binary Exploitation?
When a program is executed, its memory is divided into several different parts, which are two parts:
Stack
Heap
Stack is mostly used for temporary information of functions, but the Heap is for when the program needs to manage memory itself at runtime
What is the Heap?
Heap is a part of memory that a program can request memory from at runtime and free it later
For example, when a program does not know how much data it is going to receive, it cannot use a fixed space on the Stack, in this case it usually goes to the Heap
In languages like C and C++, Heap management is usually done with functions like:
malloc()The difference between Stack and Heap
calloc()
realloc()
free()
Stack:
It has a more organized and faster structure
The lifetime of the data is usually dependent on the function
Its management is mostly done by the compiler
For example:
void function(){
char buffer[64];
}
When the function ends, this data is destroyed
Heap:
Managed by the program itself
Data can remain for a longer period
The program itself decides when to get memory and when to free it
For example:
char *data = malloc(100);
Here the program requests to get 100 bytes of memory from the Heap
Why is the Heap important in security?
Because manual memory management is very complicated, especially in languages like C and C++
The programmer should always be careful:
When was the memory taken?
When was it freed?
Is the freed memory still in use?
Is the data size consistent with the memory size?
A small mistake can cause a vulnerability
What types of bugs are seen in the Heap?
Some famous cases:
Heap Overflow
When the program writes more data to the Heap than the allocated size
Use-After-Free (UAF)
When the program frees memory but then continues to use it
Double Free
When the program frees a section of memory more than once
Memory Leak
When the program takes memory but never frees it
Why is the Heap harder than the Stack?
Stack
has a relatively well-defined structure, but the Heap is more complex
Because the Heap is managed by a Memory Allocator
For example, in Linux, one of the famous allocators is:
ptmalloc (in glibc)
This allocator decides:
What part of the memory to allocate
Which memory to free
Where new requests should be placed
That is why Heap analysis requires a deeper understanding of memory management
Why are hackers or reverse engineers interested in the Heap?
Because the Heap usually contains important program data
For example:
Data structures
Objects in C++
Session information
Pointers
Program data
If the Heap management is wrong, it may change the behavior of the program.
The Heap is one of the most important memory areas in programs, used to store dynamic data.
Unlike the Stack, the Heap management is more the responsibility of the programmer, which leads to complex bugs such as Heap Overflow and Use-After-Free.
To understand modern exploits, it is essential to understand the Heap and how Memory Allocators work.
@reverseengine
بخش بیست و یکم بافر اورفلو
Double Free
یعنی آزاد کردن یک حافظه برای بار دوم
یکی دیگه از باگ های معروف مدیریت حافظه Double Fre هست این باگ هم توی تحلیل باینری و هم توی تحلیل بدافزار خیلی دیده میشه
Double Free یعنی چی؟
یک حافظه فقط باید یک بار
free بشهاگر همون اشاره گر دوباره
free بشهبهش میگن Double Free
مثال:
C
#include <stdlib.h>
int main() {
char *buf = malloc(32);
free(buf);
free(buf);
return 0;
}
مشکل این کجاست?
اولین
free حافظه رو آزاد میکنه ولی دومین freeداره دوباره حافظهای رو آزاد میکنه که از قبل آزاد شده همین موضوع میتونه باعث کرش برنامه یا خراب شدن ساختار Heap بشهموقع مهندسی معکوس کردن دنبال چی بگردیم؟
فرض کنید داخل دیکامپایلر اینجور چیزی دیدید
C
free(ptr);همینجا باید مشکوک بشید حالا باید بررسی کنیید آیا بین این دو
/* ... */
free(ptr);
free دوباره malloc انجام شده یا اشارهگر تغییر کرده اگر نه احتمال Double Free خیلی بالاستیک مثال واقعی تر:
C
buf = malloc(64);اینجا اگر شرط
/* ... */
if(error)
free(buf);
/* ... */
free(buf);
error برقرار باشهbufیک بار داخل شرط آزاد میشه
بعد دوباره پایین برنامه آزاد میشه
همین باعث Double Free میشه
چرا خطرناکه?
چون Heap Manager فکر میکنه
دو بار یک Chunk آزاد شده در نتیجه ساختار داخلی Heap ممکنه به هم بریزه
و همین موضوع میتونه رفتار های غیرمنتظره ایجاد کنه
چطور ازش جلوگیری میکنن?
بعد از
free اشارهگر رو NULL میکننC
free(buf);
buf = NULL;
حالا اگر دوباره
free(buf);
صدا زده بشه
free(NULL)مشکلی ایجاد نمیکنه
موقع تحلیل باینری این الگوها رو بررسی کنیو
اگر دیدید
malloc(...)
بعد
free(ptr)
و دوباره
free(ptr)
حتما مسیر اجرای برنامه رو بررسی کنید
شاید فقط در یک حالت خاص این اتفاق بیفته و همین باعث شده مدت ها کسی متوجه باگ نشه
تا الان با سه باگ مهم مدیریت حافظه آشنا شدیم
Heap Overflow
Use After Free
Double Free
این سه مورد جزو رایج ترین آسیبپذیری هایی هستن که موقع مهندسی معکوس و تحلیل باینری باهاشون روبه رو میشید
تمرین:
یک برنامه ساده بنویسید که چند مسیر مختلف اجرا داشته باشه بعد بررسی کنید آیا در یکی از مسیر ها ممکنه یک اشارهگر دو بار
free بشه یا نه اگر تونستید این الگو رو داخل یک باینری هم پیدا کنید یعنی کم کم دارید این باگ رو یاد میگیرید@reverseengine
❤4
ReverseEngineering
بخش بیست و یکم بافر اورفلو Double Free یعنی آزاد کردن یک حافظه برای بار دوم یکی دیگه از باگ های معروف مدیریت حافظه Double Fre هست این باگ هم توی تحلیل باینری و هم توی تحلیل بدافزار خیلی دیده میشه Double Free یعنی چی؟ یک حافظه فقط باید یک بار free بشه…
Part 21 Buffer Overflow
Double Free means freeing a memory for the second time
Another famous memory management bug is Double Free. This bug is often seen in both binary analysis and malware analysis
What does Double Free mean?
A memory should only be freed once
If the same pointer is freed again
It is called Double Free
Example:
C
#include <stdlib.h>
int main() {
char *buf = malloc(32);
free(buf);
free(buf);
return 0;
}
What is the problem with this?
The first free frees the memory, but the second free is freeing the memory that was already freed, which can cause the program to crash or corrupt the Heap structure
What should we look for when reverse engineering?
Suppose you see something like this in the decompiler
C
free(ptr);You should be suspicious here. Now you should check if malloc was done again between these two frees or if the pointer changed. If not, the probability of Double Free is very high.
/* ... */
free(ptr);
A more realistic example:
C
buf = malloc(64);Here, if the error condition is met,
/* ... */
if(error)
free(buf);
/* ... */
free(buf);
buf
is freed once inside the condition,
then it is freed again at the bottom of the program,
This causes Double Free.
Why is it dangerous?
Because the Heap Manager thinks that
a Chunk has been freed twice, as a result, the internal structure of the Heap may be destroyed,
And this can cause unexpected behavior.
How do you prevent it?
After free, the pointer is NULL.
C
free(buf);
buf = NULL;
Now if again
free(buf);
Calling
free(NULL)
does not cause a problem
When analyzing the binary, check for these patterns
If you see
malloc(...)
then
free(ptr)
and again
free(ptr)
Be sure to check the program execution path
Maybe this only happens in a specific case, which is why no one noticed the bug for a long time
So far, we have learned about three important memory management bugs
Heap Overflow
Use After Free
Double Free
These three are among the most common vulnerabilities that you will encounter during reverse engineering and binary analysis
Exercise:
Write a simple program that has several different execution paths
Then check whether a pointer can be freed twice in one of the paths
If you can find this pattern inside a binary, it means that you are gradually learning about this bug
@reverseengine
Complete list of LPE exploits for Windows (starting from 2023)
https://github.com/MzHmO/Exploit-Street
https://github.com/MzHmO/Exploit-Street
GitHub
GitHub - MzHmO/Exploit-Street: Complete list of LPE exploits for Windows (starting from 2023)
Complete list of LPE exploits for Windows (starting from 2023) - MzHmO/Exploit-Street
Following Attacker-Controlled Data Through Android Applications: A Practical Reverse Engineering Methodology
https://medium.com/@mustafamohammed789mm/following-attacker-controlled-data-through-android-applications-a-practical-reverse-engineering-78f566863650
https://medium.com/@mustafamohammed789mm/following-attacker-controlled-data-through-android-applications-a-practical-reverse-engineering-78f566863650
Medium
Following Attacker-Controlled Data Through Android Applications: A Practical Reverse Engineering Methodology
Every Android pentester eventually reaches the same point.
Analyzing Carbanak: A Technical Walkthrough of a Banking Trojan
https://medium.com/@ckant/analyzing-carbanak-a-technical-walkthrough-of-a-banking-trojan-7cd8ca1be31c
https://medium.com/@ckant/analyzing-carbanak-a-technical-walkthrough-of-a-banking-trojan-7cd8ca1be31c
Medium
Analyzing Carbanak: A Technical Walkthrough of a Banking Trojan
A complete reverse engineering walkthrough of the Carbanak challenge from MalOps using IDA Pro, Binary Ninja, and x64dbg.
Shipping post-quantum cryptography to Python
https://blog.trailofbits.com/2026/06/30/shipping-post-quantum-cryptography-to-python
@reverseengine
https://blog.trailofbits.com/2026/06/30/shipping-post-quantum-cryptography-to-python
@reverseengine
The Trail of Bits Blog
Shipping post-quantum cryptography to Python
We added post-quantum cryptography support to pyca/cryptography, the eleventh most-downloaded Python package on PyPI, putting quantum-resistant ML-KEM and ML-DSA algorithms one pip install away for the entire Python ecosystem.
Simplifying MBA obfuscation with CoBRA
https://blog.trailofbits.com/2026/04/03/simplifying-mba-obfuscation-with-cobra
@reverseengine
https://blog.trailofbits.com/2026/04/03/simplifying-mba-obfuscation-with-cobra
@reverseengine
The Trail of Bits Blog
Simplifying MBA obfuscation with CoBRA
We’re releasing CoBRA, an open-source tool that simplifies the full range of Mixed Boolean-Arithmetic (MBA) expressions used in the wild.
Hidden Infrastructure Exposed: ANY.RUN Reveals Hijacked Gov Websites Delivering Malware
https://any.run/cybersecurity-blog/phantomenigma-research
@reverseengine
https://any.run/cybersecurity-blog/phantomenigma-research
@reverseengine
any.run
20+ Government Websites Hijacked: PhantomEnigma Investigation
ANY.RUN uncovers how PhantomEnigma abused 20+ Brazilian government websites, hid behind trusted infrastructure, and put banks and public agencies at risk.
🔥3❤2
این ایمیل منه. اگه کاری داشتید یا حرفی، بگید حتما🩶
This is my email. If you have anything to say or need help, please do so🖤
This is my email. If you have anything to say or need help, please do so🖤
addcss012@gmail.com
❤7
بخش بیست و دوم بافر اورفلو
Memory Leak
باگی که شاید برنامه رو کرش نکنه ولی دردسر درست میکنه
تا الان با باگ هایی آشنا شدیم که حافظه رو خراب میکردن
ولی این بخش درباره باگیه که معمولا چیزی رو خراب نمیکنه
در عوض باعث میشه برنامه کم کم حافظه بیشتری مصرف کنه
و بعد از مدتی کند بشه یا حتی از کار بیفته
Memory Leak
هر وقت برنامه با
malloc حافظه بگیرهباید بعدا با
free آزادش کنهاگر این کار انجام نشه
اون حافظه تا پایان اجرای برنامه اشغال میمونه
به این میگن Memory Leak
یک مثال ساده:
C
#include <stdlib.h>
int main() {
char *buf = malloc(1024);
return 0;
}
مشکل این کجاست اینجا حافظه گرفته شده ولی هیچ وقت آزاد نشده یعنی قبل از خروج برنامه این دستور اجرا نشده
free(buf);
اگر این اتفاق یک بار بیفته چی میشه تقریبا هیچ اتفاق خاصی نمیفته ولی اگر داخل یک حلقه یا یک سرویس که همیشه در حال اجراست باشه کم کم مصرف حافظه زیاد میشه
مثلا:
C
while (1) {
char *buf = malloc(1024);
}
هر بار 1024 بایت گرفته میشه ولی هیچ وقت آزاد نمیشه بعد از مدتی برنامه مقدار زیادی حافظه مصرف میکنه
موقع مهندسی معکوس دنبال چی بگردیم؟
اگر این الگو رو دیدید
ptr = malloc(...);
بعد مسیر اجرای تابع رو تا اخر دنبال کنید
اگر هیچ جا
free(ptr);
وجود نداشت احتمال Memory Leak هست
یک مثال واقعی تر:
C
char *buf = malloc(256);
if(error)
return;
free(buf);
اینجا یک مشکل وجود داره اگر شرط
error برقرار بشه تابع قبل از رسیدن به free خارج میشه در نتیجه حافظه هیچ وقت آزاد نمیشهچرا پیدا کردنش سخت تره؟
چون معمولا برنامه کرش نمیکنه خطای واضحی نشون نمیده شاید فقط بعد از چند ساعت یا چند روز اجرا مشخص بشه
برای همین خیلی از Memory Leak ها مدت زیادی مخفی میمونن
ابزارهایی که کمک میکنن
برای پیدا کردن Memory Leak ابزار هایی مثل:
Valgrind
AddressSanitizer
LeakSanitizer
خیلی کاربردی هستن
این ابزارها نشون میدن کدوم حافظه گرفته شده ولی آزاد نشده
هر
malloc باید یک free داشته باشهنبودن
free همیشه یعنی احتمال Memory Leak در مهندسی معکوس باید مسیر malloc تا پایان تابع رو دنبال کنیمخروج زود هنگام از تابع یکی از رایج ترین دلایل Memory Leak هست
تمرین:
یک برنامه ساده که از
malloc استفاده میکنه داخل Ghidra یا IDA باز کنید بررسی کنید آیا برای همه مسیرهای اجرای برنامه در نهایت free صدا زده میشه یا نه اگر حتی یک مسیر پیدا کردید که حافظه آزاد نشه اولین Memory Leak خودتون رو پیدا کردید@reverseengine
ReverseEngineering
بخش بیست و دوم بافر اورفلو Memory Leak باگی که شاید برنامه رو کرش نکنه ولی دردسر درست میکنه تا الان با باگ هایی آشنا شدیم که حافظه رو خراب میکردن ولی این بخش درباره باگیه که معمولا چیزی رو خراب نمیکنه در عوض باعث میشه برنامه کم کم حافظه بیشتری مصرف کنه…
Part 22 Buffer Overflow
Memory Leak A bug that may not crash the program but causes problems
So far we have met with bugs that corrupt memory
But this section is about a bug that usually does not corrupt anything
Instead, it causes the program to gradually consume more memory
And after a while it slows down or even crashes
Memory Leak
Whenever a program allocates memory with malloc
It must later release it with free
If this is not done
That memory remains occupied until the end of the program execution
This is called a Memory Leak
A simple example:
C
#include <stdlib.h>
int main() {
char *buf = malloc(1024);
return 0;
}
What is the problem here?
Here, memory is allocated but never freed, meaning this instruction was not executed before the program exits
free(buf);
What if this happens once? Almost nothing special happens, but if it is inside a loop or a service that is always running, the memory consumption will gradually increase
For example:
C
while (1) {
char *buf = malloc(1024);
}
Each time 1024 bytes are taken but never freed. After a while, the program will consume a lot of memory
What should we look for when reverse engineering?
If you see this pattern
ptr = malloc(...);
Then follow the path of the function execution to the end
If there is no
free(ptr);
anywhere, there is a possibility of a Memory Leak
A more realistic example:
C
char *buf = malloc(256);
if(error)
return;
free(buf);
There is a problem here. If the error condition is met, the function exits before reaching free, as a result, the memory is never freed
Why is it harder to find?
Because the program usually does not crash, it does not show an obvious error, it may only be detected after a few hours or days of execution. That is why many memory leaks remain hidden for a long time. Tools that help to find memory leaks include: Valgrind AddressSanitizer LeakSanitizer These tools show which memory was taken but not freed Every malloc must have a free No free always means there is a possibility of a memory leak In reverse engineering, we must follow the malloc path to the end of the function Early exit from the function is one of the most common causes of memory leaks
Exercise:
Open a simple program that uses malloc in Ghidra or IDA Check whether free is called at the end for all paths of the program execution If you find even one path where the memory is not freed, you have found your first memory leak
@reverseengine
❤4
یکی از مهمترین مباحث Heap اگر این بخش رو خوب یاد بگیرید فهمیدن تکنیک های Heap Exploitation خیلی راحت تر میشه
Memory Allocator
چیه و چجوری Heap رو مدیریت میکنه
در پست قبل گفتیم که Heap بخشی از حافظه است که برنامه موقع اجرا از اون استفاده میکنه اما یک سؤال پیش میاد؟
چه کسی تصمیم میگیره حافظه از کجا اختصاص داده بشه و بعد از آزاد شدن چه اتفاقی براش میوفته
جواب این سوال Memory Allocator هست
Memory Allocator
بخشی از سیستم یا کتابخانه استاندارد که وظیفه مدیریت حافظه Heap رو به عهده داره
هر بار که برنامه از توابعی مثل:
استفاده میکنه در واقع درخواستش رو به Allocator میده
Allocator
تصمیم میگیره:
حافظه از کجا گرفته بشه
چه مقدار حافظه اختصاص داده بشه
حافظه آزاد شده دوباره چطور استفاده بشه
اگر حافظه کافی نبود چه کاری انجام بشه
هیپ یک فضای خالی بزرگه؟
جواب نه هست
خیلیها فکر میکنن Heap فقط یک فضای خالی بزرگه که برنامه هر جا خواست داخلش مینویسه در واقع Heap از بخشهای کوچیکی تشکیل شده که به اونا Chunk میگن
هر بار که ()
Chunk
کوچکترین واحدی که Allocator مدیریت میکنه
هر Chunk دو قسمت اصلی داره:
قسمت Metadata اطلاعات مدیریتی Chunk رو نگه میداره
مثلا:
اندازه Chunk
وضعیت آزاد یا اشغال بودن
اطلاعاتی که Allocator برای مدیریت حافظه نیاز داره
بعد از Metadata بخشی قرار داره که برنامه واقعا از اون استفاده کرده
چرا Metadata مهمه؟
چون Allocator برای تصمیم گیری به همین اطلاعات وابسته هست
اگر این اطلاعات به هر دلیلی خراب بشن Allocator ممکنه حافظه رو اشتباه مدیریت کنه به همین دلیل بیشتر آسیب پذیری های Heap در گذشته تلاش میکردن Metadata رو هدف قرار بدن
البته Allocator های امروزی نسبت به گذشته محافظت های بیشتری دارن و سواستفاده از این ساختار ها سخت تر شده
همه سیستمها از یک Allocator استفاده میکنن؟ نه
هر سیستم عامل یا حتی هر برنامه ممکنه از Allocator متفاوتی استفاده کنه
چند نمونه معروف:
هر کدوم طراحی و روش مدیریت متفاوتی دارن اما هدف همه یکیه:
مدیریت سریع و بهینه حافظه Heap
چرا شناخت Allocator مهمه؟
وقتی درباره Heap Exploitation صحبت میکنیم فقط با خود Heap سر و کار نداریم
در واقع داریم رفتار Allocator رو بررسی میکنیم اگر ندونیم Allocator چجوری تصمیم میگیره حافظه رو اختصاص بده یا آزاد کنه درک باگ های Heap هم سخت میشه
به همین دلیل قبل از یادگیری تکنیک های Heap Exploitation باید با ساختار Allocator آشنا بشیم
Heap
توسط بخشی به نام Memory Allocator مدیریت میشه این بخش مسئول تخصیص آزاد سازی و استفاده مجدد از حافظه ست حافظه Heap به واحد هایی به نام Chunk تقسیم میشه و هر Chunk اطلاعات مدیریتی مخصوص خودش رو داره شناخت این ساختار پایه ی یادگیری Heap Exploitation هست
@reverseengine
Memory Allocator
چیه و چجوری Heap رو مدیریت میکنه
در پست قبل گفتیم که Heap بخشی از حافظه است که برنامه موقع اجرا از اون استفاده میکنه اما یک سؤال پیش میاد؟
چه کسی تصمیم میگیره حافظه از کجا اختصاص داده بشه و بعد از آزاد شدن چه اتفاقی براش میوفته
جواب این سوال Memory Allocator هست
Memory Allocator
بخشی از سیستم یا کتابخانه استاندارد که وظیفه مدیریت حافظه Heap رو به عهده داره
هر بار که برنامه از توابعی مثل:
malloc()
calloc()
realloc()
free()
استفاده میکنه در واقع درخواستش رو به Allocator میده
Allocator
تصمیم میگیره:
حافظه از کجا گرفته بشه
چه مقدار حافظه اختصاص داده بشه
حافظه آزاد شده دوباره چطور استفاده بشه
اگر حافظه کافی نبود چه کاری انجام بشه
هیپ یک فضای خالی بزرگه؟
جواب نه هست
خیلیها فکر میکنن Heap فقط یک فضای خالی بزرگه که برنامه هر جا خواست داخلش مینویسه در واقع Heap از بخشهای کوچیکی تشکیل شده که به اونا Chunk میگن
هر بار که ()
malloc صدا زده میشه معمولا یک Chunk به برنامه تحویل داده میشهChunk
کوچکترین واحدی که Allocator مدیریت میکنه
هر Chunk دو قسمت اصلی داره:
Metadata
User Data
قسمت Metadata اطلاعات مدیریتی Chunk رو نگه میداره
مثلا:
اندازه Chunk
وضعیت آزاد یا اشغال بودن
اطلاعاتی که Allocator برای مدیریت حافظه نیاز داره
بعد از Metadata بخشی قرار داره که برنامه واقعا از اون استفاده کرده
چرا Metadata مهمه؟
چون Allocator برای تصمیم گیری به همین اطلاعات وابسته هست
اگر این اطلاعات به هر دلیلی خراب بشن Allocator ممکنه حافظه رو اشتباه مدیریت کنه به همین دلیل بیشتر آسیب پذیری های Heap در گذشته تلاش میکردن Metadata رو هدف قرار بدن
البته Allocator های امروزی نسبت به گذشته محافظت های بیشتری دارن و سواستفاده از این ساختار ها سخت تر شده
همه سیستمها از یک Allocator استفاده میکنن؟ نه
هر سیستم عامل یا حتی هر برنامه ممکنه از Allocator متفاوتی استفاده کنه
چند نمونه معروف:
ptmalloc (داخل لینوکس glibc)
jemalloc
tcmalloc
mimalloc
هر کدوم طراحی و روش مدیریت متفاوتی دارن اما هدف همه یکیه:
مدیریت سریع و بهینه حافظه Heap
چرا شناخت Allocator مهمه؟
وقتی درباره Heap Exploitation صحبت میکنیم فقط با خود Heap سر و کار نداریم
در واقع داریم رفتار Allocator رو بررسی میکنیم اگر ندونیم Allocator چجوری تصمیم میگیره حافظه رو اختصاص بده یا آزاد کنه درک باگ های Heap هم سخت میشه
به همین دلیل قبل از یادگیری تکنیک های Heap Exploitation باید با ساختار Allocator آشنا بشیم
Heap
توسط بخشی به نام Memory Allocator مدیریت میشه این بخش مسئول تخصیص آزاد سازی و استفاده مجدد از حافظه ست حافظه Heap به واحد هایی به نام Chunk تقسیم میشه و هر Chunk اطلاعات مدیریتی مخصوص خودش رو داره شناخت این ساختار پایه ی یادگیری Heap Exploitation هست
@reverseengine
ReverseEngineering
یکی از مهمترین مباحث Heap اگر این بخش رو خوب یاد بگیرید فهمیدن تکنیک های Heap Exploitation خیلی راحت تر میشه Memory Allocator چیه و چجوری Heap رو مدیریت میکنه در پست قبل گفتیم که Heap بخشی از حافظه است که برنامه موقع اجرا از اون استفاده میکنه اما یک…
One of the most important topics of Heap is that if you learn this section well, it will be much easier to understand Heap Exploitation techniques
What is Memory Allocator and how does it manage Heap
In the previous post, we said that Heap is a part of memory that the program uses when it runs, but a question arises?
Who decides where to allocate memory and what happens to it after it is freed
The answer to this question is Memory Allocator
Memory Allocator
A part of the system or standard library that is responsible for managing Heap memory
Every time a program uses functions such as:
it actually sends its request to Allocator
Allocator
Decides:
Where to get memory
How much memory to allocate
How to reuse freed memory
What to do if there is not enough memory
Is the heap a large empty space?
The answer is no
Many people think that the Heap is just a big empty space that the program writes to wherever it wants. In fact, the Heap is made up of small parts called Chunks
Every time malloc() is called, a Chunk is usually handed over to the program
Chunk
The smallest unit that the Allocator manages
Each Chunk has two main parts:
The Metadata part holds Chunk management information
For example:
Information that the Allocator needs to manage memory
After the Metadata is the part that the program actually uses
Why is Metadata important?
Because the Allocator relies on this information to make decisions If this information is corrupted for any reason, the Allocator may mismanage memory, which is why most Heap vulnerabilities in the past tried to target Metadata
Of course, today's Allocators have more protections than in the past and it has become harder to abuse these structures
Do all systems use the same Allocator? No
Each operating system or even each program may use a different Allocator
A few famous examples:
Each has a different design and management method, but the goal is the same:
Fast and efficient management of Heap memory
Why is it important to understand Allocators?
When we talk about Heap Exploitation, we are not just dealing with the Heap itself. We are actually examining the behavior of the Allocator. If we do not know how the Allocator decides to allocate or free memory, it will be difficult to understand Heap bugs. That is why before learning Heap Exploitation techniques, we should familiarize ourselves with the Allocator structure. The Heap is managed by a part called the Memory Allocator. This part is responsible for allocating, freeing, and reusing memory.
The Heap memory set is divided into units called Chunks, and each Chunk has its own management information. Understanding this structure is the basis for learning Heap Exploitation.
@reverseengine
What is Memory Allocator and how does it manage Heap
In the previous post, we said that Heap is a part of memory that the program uses when it runs, but a question arises?
Who decides where to allocate memory and what happens to it after it is freed
The answer to this question is Memory Allocator
Memory Allocator
A part of the system or standard library that is responsible for managing Heap memory
Every time a program uses functions such as:
malloc()
calloc()
realloc()
free()
it actually sends its request to Allocator
Allocator
Decides:
Where to get memory
How much memory to allocate
How to reuse freed memory
What to do if there is not enough memory
Is the heap a large empty space?
The answer is no
Many people think that the Heap is just a big empty space that the program writes to wherever it wants. In fact, the Heap is made up of small parts called Chunks
Every time malloc() is called, a Chunk is usually handed over to the program
Chunk
The smallest unit that the Allocator manages
Each Chunk has two main parts:
Metadata
User Data
The Metadata part holds Chunk management information
For example:
Chunk size
Free or busy status
Information that the Allocator needs to manage memory
After the Metadata is the part that the program actually uses
Why is Metadata important?
Because the Allocator relies on this information to make decisions If this information is corrupted for any reason, the Allocator may mismanage memory, which is why most Heap vulnerabilities in the past tried to target Metadata
Of course, today's Allocators have more protections than in the past and it has become harder to abuse these structures
Do all systems use the same Allocator? No
Each operating system or even each program may use a different Allocator
A few famous examples:
ptmalloc (inside Linux glibc)
jemalloc
tcmalloc
mimalloc
Each has a different design and management method, but the goal is the same:
Fast and efficient management of Heap memory
Why is it important to understand Allocators?
When we talk about Heap Exploitation, we are not just dealing with the Heap itself. We are actually examining the behavior of the Allocator. If we do not know how the Allocator decides to allocate or free memory, it will be difficult to understand Heap bugs. That is why before learning Heap Exploitation techniques, we should familiarize ourselves with the Allocator structure. The Heap is managed by a part called the Memory Allocator. This part is responsible for allocating, freeing, and reusing memory.
The Heap memory set is divided into units called Chunks, and each Chunk has its own management information. Understanding this structure is the basis for learning Heap Exploitation.
@reverseengine
👍1
Opcode Encoding
چرا Opcode ها اینقدر عجیب به نظر میرسن؟
تا اینجا فرض کردیم Opcode ها خیلی ساده هستن
مثلا:
ولی توی ماشینهای مجازی واقعی تقریبا هیچ وقت اوضاع اینقدر ساده نیست
سازنده محافظ نمیخواد تحلیلگر با چند دقیقه نگاه کردن معنی Opcode ها رو بفهمه
به همین خاطر Opcode ها رو به شکل های مختلف مخفی میکنه
مثلا ممکنه Opcode واقعی این باشه:
ولی قبل از اجرا این عملیات روی اون انجام بشه:
بعد از XOR شدن تازه مقدار واقعی به دست میاد
یا ممکنه Opcode اصلا مستقیم داخل بایت کد ذخیره نشده باشه
مثلا قبل از استفاده از روی یک جدول ترجمه عبور کنه
در این حالت اگر فقط به بایتکد نگاه کنید هیچ معنی خاصی نمیبینید
یکی دیگه از روشهای رایج اینه که اندازه Opcode ها ثابت نباشه
مثلا:
یا:
یا حتی:
یعنی هر Opcode تعداد متفاوتی Operand داره اگر تحلیلگر این موضوع رو متوجه نشه از همون دستور اول کل بایت کد رو اشتباه تفسیر میکنه بعضی ماشینهای مجازی حتی Opcode ها رو موقع اجرا تولید میکنن یعنی مقداری که Dispatcher میبینه همون مقداری نیست که داخل فایل ذخیره شده به همین خاطر یکی از اولین کارهای تحلیلگر اینه که مسیر رسیدن Opcode به Dispatcher رو دنبال کنه اگر قبل از Dispatcher عملیاتی مثل XOR، ADD، SUB یا چرخش بیت ها انجام بشه احتمال زیادی وجود داره که Opcode ها رمزگذاری شده باشن هدف از Opcode Encoding فقط سختتر کردن تحلیل نیست باعث میشه ابزارهایی مثل IDA یا Ghidra هم نتونن به راحتی منطق ماشین مجازی رو تشخیص بدن به همین دلیل تحلیلگرها معمولا قبل از اینکه سراغ Handler ها برن سعی میکنن بفهمن Opcode ها دقیقا چطور Decode میشن اگر این مرحله رو درست انجام بدید ادامه فرایند Devirtualization خیلی ساده تر میشه
تمرین:
فرض کنید بایت کد زیر رو دارید:
و میدونید قبل از اجرا هر Opcode با
اول مقدار واقعی هر Opcode رو حساب کنید
بعد فرض کنید نتیجه این جدول باشه:
سعی کنید مسیر اجرای ماشین مجازی رو روی کاغذ باز سازی کنید
@reverseengine
چرا Opcode ها اینقدر عجیب به نظر میرسن؟
تا اینجا فرض کردیم Opcode ها خیلی ساده هستن
مثلا:
01 = LOAD
02 = ADD
03 = JMP
ولی توی ماشینهای مجازی واقعی تقریبا هیچ وقت اوضاع اینقدر ساده نیست
سازنده محافظ نمیخواد تحلیلگر با چند دقیقه نگاه کردن معنی Opcode ها رو بفهمه
به همین خاطر Opcode ها رو به شکل های مختلف مخفی میکنه
مثلا ممکنه Opcode واقعی این باشه:
0x8F
ولی قبل از اجرا این عملیات روی اون انجام بشه:
Opcode ^= 0xA5
بعد از XOR شدن تازه مقدار واقعی به دست میاد
یا ممکنه Opcode اصلا مستقیم داخل بایت کد ذخیره نشده باشه
مثلا قبل از استفاده از روی یک جدول ترجمه عبور کنه
0x3C → LOAD
0x91 → ADD
0xE7 → JMP
در این حالت اگر فقط به بایتکد نگاه کنید هیچ معنی خاصی نمیبینید
یکی دیگه از روشهای رایج اینه که اندازه Opcode ها ثابت نباشه
مثلا:
Opcode
Operand
Operand
یا:
Opcode
Operand
یا حتی:
Opcode
Operand
Operand
Operand
یعنی هر Opcode تعداد متفاوتی Operand داره اگر تحلیلگر این موضوع رو متوجه نشه از همون دستور اول کل بایت کد رو اشتباه تفسیر میکنه بعضی ماشینهای مجازی حتی Opcode ها رو موقع اجرا تولید میکنن یعنی مقداری که Dispatcher میبینه همون مقداری نیست که داخل فایل ذخیره شده به همین خاطر یکی از اولین کارهای تحلیلگر اینه که مسیر رسیدن Opcode به Dispatcher رو دنبال کنه اگر قبل از Dispatcher عملیاتی مثل XOR، ADD، SUB یا چرخش بیت ها انجام بشه احتمال زیادی وجود داره که Opcode ها رمزگذاری شده باشن هدف از Opcode Encoding فقط سختتر کردن تحلیل نیست باعث میشه ابزارهایی مثل IDA یا Ghidra هم نتونن به راحتی منطق ماشین مجازی رو تشخیص بدن به همین دلیل تحلیلگرها معمولا قبل از اینکه سراغ Handler ها برن سعی میکنن بفهمن Opcode ها دقیقا چطور Decode میشن اگر این مرحله رو درست انجام بدید ادامه فرایند Devirtualization خیلی ساده تر میشه
تمرین:
فرض کنید بایت کد زیر رو دارید:
8F 91 E7
و میدونید قبل از اجرا هر Opcode با
0xA5 عمل XOR میشهاول مقدار واقعی هر Opcode رو حساب کنید
بعد فرض کنید نتیجه این جدول باشه:
2A = LOAD
34 = ADD
42 = PRINT
سعی کنید مسیر اجرای ماشین مجازی رو روی کاغذ باز سازی کنید
@reverseengine