Hidden Infrastructure Exposed: ANY.RUN Reveals Hijacked Gov Websites Delivering Malware
https://any.run/cybersecurity-blog/phantomenigma-research
@reverseengine
https://any.run/cybersecurity-blog/phantomenigma-research
@reverseengine
any.run
20+ Government Websites Hijacked: PhantomEnigma Investigation
ANY.RUN uncovers how PhantomEnigma abused 20+ Brazilian government websites, hid behind trusted infrastructure, and put banks and public agencies at risk.
🔥3❤2
این ایمیل منه. اگه کاری داشتید یا حرفی، بگید حتما🩶
This is my email. If you have anything to say or need help, please do so🖤
This is my email. If you have anything to say or need help, please do so🖤
addcss012@gmail.com
❤7
بخش بیست و دوم بافر اورفلو
Memory Leak
باگی که شاید برنامه رو کرش نکنه ولی دردسر درست میکنه
تا الان با باگ هایی آشنا شدیم که حافظه رو خراب میکردن
ولی این بخش درباره باگیه که معمولا چیزی رو خراب نمیکنه
در عوض باعث میشه برنامه کم کم حافظه بیشتری مصرف کنه
و بعد از مدتی کند بشه یا حتی از کار بیفته
Memory Leak
هر وقت برنامه با
malloc حافظه بگیرهباید بعدا با
free آزادش کنهاگر این کار انجام نشه
اون حافظه تا پایان اجرای برنامه اشغال میمونه
به این میگن Memory Leak
یک مثال ساده:
C
#include <stdlib.h>
int main() {
char *buf = malloc(1024);
return 0;
}
مشکل این کجاست اینجا حافظه گرفته شده ولی هیچ وقت آزاد نشده یعنی قبل از خروج برنامه این دستور اجرا نشده
free(buf);
اگر این اتفاق یک بار بیفته چی میشه تقریبا هیچ اتفاق خاصی نمیفته ولی اگر داخل یک حلقه یا یک سرویس که همیشه در حال اجراست باشه کم کم مصرف حافظه زیاد میشه
مثلا:
C
while (1) {
char *buf = malloc(1024);
}
هر بار 1024 بایت گرفته میشه ولی هیچ وقت آزاد نمیشه بعد از مدتی برنامه مقدار زیادی حافظه مصرف میکنه
موقع مهندسی معکوس دنبال چی بگردیم؟
اگر این الگو رو دیدید
ptr = malloc(...);
بعد مسیر اجرای تابع رو تا اخر دنبال کنید
اگر هیچ جا
free(ptr);
وجود نداشت احتمال Memory Leak هست
یک مثال واقعی تر:
C
char *buf = malloc(256);
if(error)
return;
free(buf);
اینجا یک مشکل وجود داره اگر شرط
error برقرار بشه تابع قبل از رسیدن به free خارج میشه در نتیجه حافظه هیچ وقت آزاد نمیشهچرا پیدا کردنش سخت تره؟
چون معمولا برنامه کرش نمیکنه خطای واضحی نشون نمیده شاید فقط بعد از چند ساعت یا چند روز اجرا مشخص بشه
برای همین خیلی از Memory Leak ها مدت زیادی مخفی میمونن
ابزارهایی که کمک میکنن
برای پیدا کردن Memory Leak ابزار هایی مثل:
Valgrind
AddressSanitizer
LeakSanitizer
خیلی کاربردی هستن
این ابزارها نشون میدن کدوم حافظه گرفته شده ولی آزاد نشده
هر
malloc باید یک free داشته باشهنبودن
free همیشه یعنی احتمال Memory Leak در مهندسی معکوس باید مسیر malloc تا پایان تابع رو دنبال کنیمخروج زود هنگام از تابع یکی از رایج ترین دلایل Memory Leak هست
تمرین:
یک برنامه ساده که از
malloc استفاده میکنه داخل Ghidra یا IDA باز کنید بررسی کنید آیا برای همه مسیرهای اجرای برنامه در نهایت free صدا زده میشه یا نه اگر حتی یک مسیر پیدا کردید که حافظه آزاد نشه اولین Memory Leak خودتون رو پیدا کردید@reverseengine
ReverseEngineering
بخش بیست و دوم بافر اورفلو Memory Leak باگی که شاید برنامه رو کرش نکنه ولی دردسر درست میکنه تا الان با باگ هایی آشنا شدیم که حافظه رو خراب میکردن ولی این بخش درباره باگیه که معمولا چیزی رو خراب نمیکنه در عوض باعث میشه برنامه کم کم حافظه بیشتری مصرف کنه…
Part 22 Buffer Overflow
Memory Leak A bug that may not crash the program but causes problems
So far we have met with bugs that corrupt memory
But this section is about a bug that usually does not corrupt anything
Instead, it causes the program to gradually consume more memory
And after a while it slows down or even crashes
Memory Leak
Whenever a program allocates memory with malloc
It must later release it with free
If this is not done
That memory remains occupied until the end of the program execution
This is called a Memory Leak
A simple example:
C
#include <stdlib.h>
int main() {
char *buf = malloc(1024);
return 0;
}
What is the problem here?
Here, memory is allocated but never freed, meaning this instruction was not executed before the program exits
free(buf);
What if this happens once? Almost nothing special happens, but if it is inside a loop or a service that is always running, the memory consumption will gradually increase
For example:
C
while (1) {
char *buf = malloc(1024);
}
Each time 1024 bytes are taken but never freed. After a while, the program will consume a lot of memory
What should we look for when reverse engineering?
If you see this pattern
ptr = malloc(...);
Then follow the path of the function execution to the end
If there is no
free(ptr);
anywhere, there is a possibility of a Memory Leak
A more realistic example:
C
char *buf = malloc(256);
if(error)
return;
free(buf);
There is a problem here. If the error condition is met, the function exits before reaching free, as a result, the memory is never freed
Why is it harder to find?
Because the program usually does not crash, it does not show an obvious error, it may only be detected after a few hours or days of execution. That is why many memory leaks remain hidden for a long time. Tools that help to find memory leaks include: Valgrind AddressSanitizer LeakSanitizer These tools show which memory was taken but not freed Every malloc must have a free No free always means there is a possibility of a memory leak In reverse engineering, we must follow the malloc path to the end of the function Early exit from the function is one of the most common causes of memory leaks
Exercise:
Open a simple program that uses malloc in Ghidra or IDA Check whether free is called at the end for all paths of the program execution If you find even one path where the memory is not freed, you have found your first memory leak
@reverseengine
❤4
یکی از مهمترین مباحث Heap اگر این بخش رو خوب یاد بگیرید فهمیدن تکنیک های Heap Exploitation خیلی راحت تر میشه
Memory Allocator
چیه و چجوری Heap رو مدیریت میکنه
در پست قبل گفتیم که Heap بخشی از حافظه است که برنامه موقع اجرا از اون استفاده میکنه اما یک سؤال پیش میاد؟
چه کسی تصمیم میگیره حافظه از کجا اختصاص داده بشه و بعد از آزاد شدن چه اتفاقی براش میوفته
جواب این سوال Memory Allocator هست
Memory Allocator
بخشی از سیستم یا کتابخانه استاندارد که وظیفه مدیریت حافظه Heap رو به عهده داره
هر بار که برنامه از توابعی مثل:
استفاده میکنه در واقع درخواستش رو به Allocator میده
Allocator
تصمیم میگیره:
حافظه از کجا گرفته بشه
چه مقدار حافظه اختصاص داده بشه
حافظه آزاد شده دوباره چطور استفاده بشه
اگر حافظه کافی نبود چه کاری انجام بشه
هیپ یک فضای خالی بزرگه؟
جواب نه هست
خیلیها فکر میکنن Heap فقط یک فضای خالی بزرگه که برنامه هر جا خواست داخلش مینویسه در واقع Heap از بخشهای کوچیکی تشکیل شده که به اونا Chunk میگن
هر بار که ()
Chunk
کوچکترین واحدی که Allocator مدیریت میکنه
هر Chunk دو قسمت اصلی داره:
قسمت Metadata اطلاعات مدیریتی Chunk رو نگه میداره
مثلا:
اندازه Chunk
وضعیت آزاد یا اشغال بودن
اطلاعاتی که Allocator برای مدیریت حافظه نیاز داره
بعد از Metadata بخشی قرار داره که برنامه واقعا از اون استفاده کرده
چرا Metadata مهمه؟
چون Allocator برای تصمیم گیری به همین اطلاعات وابسته هست
اگر این اطلاعات به هر دلیلی خراب بشن Allocator ممکنه حافظه رو اشتباه مدیریت کنه به همین دلیل بیشتر آسیب پذیری های Heap در گذشته تلاش میکردن Metadata رو هدف قرار بدن
البته Allocator های امروزی نسبت به گذشته محافظت های بیشتری دارن و سواستفاده از این ساختار ها سخت تر شده
همه سیستمها از یک Allocator استفاده میکنن؟ نه
هر سیستم عامل یا حتی هر برنامه ممکنه از Allocator متفاوتی استفاده کنه
چند نمونه معروف:
هر کدوم طراحی و روش مدیریت متفاوتی دارن اما هدف همه یکیه:
مدیریت سریع و بهینه حافظه Heap
چرا شناخت Allocator مهمه؟
وقتی درباره Heap Exploitation صحبت میکنیم فقط با خود Heap سر و کار نداریم
در واقع داریم رفتار Allocator رو بررسی میکنیم اگر ندونیم Allocator چجوری تصمیم میگیره حافظه رو اختصاص بده یا آزاد کنه درک باگ های Heap هم سخت میشه
به همین دلیل قبل از یادگیری تکنیک های Heap Exploitation باید با ساختار Allocator آشنا بشیم
Heap
توسط بخشی به نام Memory Allocator مدیریت میشه این بخش مسئول تخصیص آزاد سازی و استفاده مجدد از حافظه ست حافظه Heap به واحد هایی به نام Chunk تقسیم میشه و هر Chunk اطلاعات مدیریتی مخصوص خودش رو داره شناخت این ساختار پایه ی یادگیری Heap Exploitation هست
@reverseengine
Memory Allocator
چیه و چجوری Heap رو مدیریت میکنه
در پست قبل گفتیم که Heap بخشی از حافظه است که برنامه موقع اجرا از اون استفاده میکنه اما یک سؤال پیش میاد؟
چه کسی تصمیم میگیره حافظه از کجا اختصاص داده بشه و بعد از آزاد شدن چه اتفاقی براش میوفته
جواب این سوال Memory Allocator هست
Memory Allocator
بخشی از سیستم یا کتابخانه استاندارد که وظیفه مدیریت حافظه Heap رو به عهده داره
هر بار که برنامه از توابعی مثل:
malloc()
calloc()
realloc()
free()
استفاده میکنه در واقع درخواستش رو به Allocator میده
Allocator
تصمیم میگیره:
حافظه از کجا گرفته بشه
چه مقدار حافظه اختصاص داده بشه
حافظه آزاد شده دوباره چطور استفاده بشه
اگر حافظه کافی نبود چه کاری انجام بشه
هیپ یک فضای خالی بزرگه؟
جواب نه هست
خیلیها فکر میکنن Heap فقط یک فضای خالی بزرگه که برنامه هر جا خواست داخلش مینویسه در واقع Heap از بخشهای کوچیکی تشکیل شده که به اونا Chunk میگن
هر بار که ()
malloc صدا زده میشه معمولا یک Chunk به برنامه تحویل داده میشهChunk
کوچکترین واحدی که Allocator مدیریت میکنه
هر Chunk دو قسمت اصلی داره:
Metadata
User Data
قسمت Metadata اطلاعات مدیریتی Chunk رو نگه میداره
مثلا:
اندازه Chunk
وضعیت آزاد یا اشغال بودن
اطلاعاتی که Allocator برای مدیریت حافظه نیاز داره
بعد از Metadata بخشی قرار داره که برنامه واقعا از اون استفاده کرده
چرا Metadata مهمه؟
چون Allocator برای تصمیم گیری به همین اطلاعات وابسته هست
اگر این اطلاعات به هر دلیلی خراب بشن Allocator ممکنه حافظه رو اشتباه مدیریت کنه به همین دلیل بیشتر آسیب پذیری های Heap در گذشته تلاش میکردن Metadata رو هدف قرار بدن
البته Allocator های امروزی نسبت به گذشته محافظت های بیشتری دارن و سواستفاده از این ساختار ها سخت تر شده
همه سیستمها از یک Allocator استفاده میکنن؟ نه
هر سیستم عامل یا حتی هر برنامه ممکنه از Allocator متفاوتی استفاده کنه
چند نمونه معروف:
ptmalloc (داخل لینوکس glibc)
jemalloc
tcmalloc
mimalloc
هر کدوم طراحی و روش مدیریت متفاوتی دارن اما هدف همه یکیه:
مدیریت سریع و بهینه حافظه Heap
چرا شناخت Allocator مهمه؟
وقتی درباره Heap Exploitation صحبت میکنیم فقط با خود Heap سر و کار نداریم
در واقع داریم رفتار Allocator رو بررسی میکنیم اگر ندونیم Allocator چجوری تصمیم میگیره حافظه رو اختصاص بده یا آزاد کنه درک باگ های Heap هم سخت میشه
به همین دلیل قبل از یادگیری تکنیک های Heap Exploitation باید با ساختار Allocator آشنا بشیم
Heap
توسط بخشی به نام Memory Allocator مدیریت میشه این بخش مسئول تخصیص آزاد سازی و استفاده مجدد از حافظه ست حافظه Heap به واحد هایی به نام Chunk تقسیم میشه و هر Chunk اطلاعات مدیریتی مخصوص خودش رو داره شناخت این ساختار پایه ی یادگیری Heap Exploitation هست
@reverseengine
ReverseEngineering
یکی از مهمترین مباحث Heap اگر این بخش رو خوب یاد بگیرید فهمیدن تکنیک های Heap Exploitation خیلی راحت تر میشه Memory Allocator چیه و چجوری Heap رو مدیریت میکنه در پست قبل گفتیم که Heap بخشی از حافظه است که برنامه موقع اجرا از اون استفاده میکنه اما یک…
One of the most important topics of Heap is that if you learn this section well, it will be much easier to understand Heap Exploitation techniques
What is Memory Allocator and how does it manage Heap
In the previous post, we said that Heap is a part of memory that the program uses when it runs, but a question arises?
Who decides where to allocate memory and what happens to it after it is freed
The answer to this question is Memory Allocator
Memory Allocator
A part of the system or standard library that is responsible for managing Heap memory
Every time a program uses functions such as:
it actually sends its request to Allocator
Allocator
Decides:
Where to get memory
How much memory to allocate
How to reuse freed memory
What to do if there is not enough memory
Is the heap a large empty space?
The answer is no
Many people think that the Heap is just a big empty space that the program writes to wherever it wants. In fact, the Heap is made up of small parts called Chunks
Every time malloc() is called, a Chunk is usually handed over to the program
Chunk
The smallest unit that the Allocator manages
Each Chunk has two main parts:
The Metadata part holds Chunk management information
For example:
Information that the Allocator needs to manage memory
After the Metadata is the part that the program actually uses
Why is Metadata important?
Because the Allocator relies on this information to make decisions If this information is corrupted for any reason, the Allocator may mismanage memory, which is why most Heap vulnerabilities in the past tried to target Metadata
Of course, today's Allocators have more protections than in the past and it has become harder to abuse these structures
Do all systems use the same Allocator? No
Each operating system or even each program may use a different Allocator
A few famous examples:
Each has a different design and management method, but the goal is the same:
Fast and efficient management of Heap memory
Why is it important to understand Allocators?
When we talk about Heap Exploitation, we are not just dealing with the Heap itself. We are actually examining the behavior of the Allocator. If we do not know how the Allocator decides to allocate or free memory, it will be difficult to understand Heap bugs. That is why before learning Heap Exploitation techniques, we should familiarize ourselves with the Allocator structure. The Heap is managed by a part called the Memory Allocator. This part is responsible for allocating, freeing, and reusing memory.
The Heap memory set is divided into units called Chunks, and each Chunk has its own management information. Understanding this structure is the basis for learning Heap Exploitation.
@reverseengine
What is Memory Allocator and how does it manage Heap
In the previous post, we said that Heap is a part of memory that the program uses when it runs, but a question arises?
Who decides where to allocate memory and what happens to it after it is freed
The answer to this question is Memory Allocator
Memory Allocator
A part of the system or standard library that is responsible for managing Heap memory
Every time a program uses functions such as:
malloc()
calloc()
realloc()
free()
it actually sends its request to Allocator
Allocator
Decides:
Where to get memory
How much memory to allocate
How to reuse freed memory
What to do if there is not enough memory
Is the heap a large empty space?
The answer is no
Many people think that the Heap is just a big empty space that the program writes to wherever it wants. In fact, the Heap is made up of small parts called Chunks
Every time malloc() is called, a Chunk is usually handed over to the program
Chunk
The smallest unit that the Allocator manages
Each Chunk has two main parts:
Metadata
User Data
The Metadata part holds Chunk management information
For example:
Chunk size
Free or busy status
Information that the Allocator needs to manage memory
After the Metadata is the part that the program actually uses
Why is Metadata important?
Because the Allocator relies on this information to make decisions If this information is corrupted for any reason, the Allocator may mismanage memory, which is why most Heap vulnerabilities in the past tried to target Metadata
Of course, today's Allocators have more protections than in the past and it has become harder to abuse these structures
Do all systems use the same Allocator? No
Each operating system or even each program may use a different Allocator
A few famous examples:
ptmalloc (inside Linux glibc)
jemalloc
tcmalloc
mimalloc
Each has a different design and management method, but the goal is the same:
Fast and efficient management of Heap memory
Why is it important to understand Allocators?
When we talk about Heap Exploitation, we are not just dealing with the Heap itself. We are actually examining the behavior of the Allocator. If we do not know how the Allocator decides to allocate or free memory, it will be difficult to understand Heap bugs. That is why before learning Heap Exploitation techniques, we should familiarize ourselves with the Allocator structure. The Heap is managed by a part called the Memory Allocator. This part is responsible for allocating, freeing, and reusing memory.
The Heap memory set is divided into units called Chunks, and each Chunk has its own management information. Understanding this structure is the basis for learning Heap Exploitation.
@reverseengine
👍1
Opcode Encoding
چرا Opcode ها اینقدر عجیب به نظر میرسن؟
تا اینجا فرض کردیم Opcode ها خیلی ساده هستن
مثلا:
ولی توی ماشینهای مجازی واقعی تقریبا هیچ وقت اوضاع اینقدر ساده نیست
سازنده محافظ نمیخواد تحلیلگر با چند دقیقه نگاه کردن معنی Opcode ها رو بفهمه
به همین خاطر Opcode ها رو به شکل های مختلف مخفی میکنه
مثلا ممکنه Opcode واقعی این باشه:
ولی قبل از اجرا این عملیات روی اون انجام بشه:
بعد از XOR شدن تازه مقدار واقعی به دست میاد
یا ممکنه Opcode اصلا مستقیم داخل بایت کد ذخیره نشده باشه
مثلا قبل از استفاده از روی یک جدول ترجمه عبور کنه
در این حالت اگر فقط به بایتکد نگاه کنید هیچ معنی خاصی نمیبینید
یکی دیگه از روشهای رایج اینه که اندازه Opcode ها ثابت نباشه
مثلا:
یا:
یا حتی:
یعنی هر Opcode تعداد متفاوتی Operand داره اگر تحلیلگر این موضوع رو متوجه نشه از همون دستور اول کل بایت کد رو اشتباه تفسیر میکنه بعضی ماشینهای مجازی حتی Opcode ها رو موقع اجرا تولید میکنن یعنی مقداری که Dispatcher میبینه همون مقداری نیست که داخل فایل ذخیره شده به همین خاطر یکی از اولین کارهای تحلیلگر اینه که مسیر رسیدن Opcode به Dispatcher رو دنبال کنه اگر قبل از Dispatcher عملیاتی مثل XOR، ADD، SUB یا چرخش بیت ها انجام بشه احتمال زیادی وجود داره که Opcode ها رمزگذاری شده باشن هدف از Opcode Encoding فقط سختتر کردن تحلیل نیست باعث میشه ابزارهایی مثل IDA یا Ghidra هم نتونن به راحتی منطق ماشین مجازی رو تشخیص بدن به همین دلیل تحلیلگرها معمولا قبل از اینکه سراغ Handler ها برن سعی میکنن بفهمن Opcode ها دقیقا چطور Decode میشن اگر این مرحله رو درست انجام بدید ادامه فرایند Devirtualization خیلی ساده تر میشه
تمرین:
فرض کنید بایت کد زیر رو دارید:
و میدونید قبل از اجرا هر Opcode با
اول مقدار واقعی هر Opcode رو حساب کنید
بعد فرض کنید نتیجه این جدول باشه:
سعی کنید مسیر اجرای ماشین مجازی رو روی کاغذ باز سازی کنید
@reverseengine
چرا Opcode ها اینقدر عجیب به نظر میرسن؟
تا اینجا فرض کردیم Opcode ها خیلی ساده هستن
مثلا:
01 = LOAD
02 = ADD
03 = JMP
ولی توی ماشینهای مجازی واقعی تقریبا هیچ وقت اوضاع اینقدر ساده نیست
سازنده محافظ نمیخواد تحلیلگر با چند دقیقه نگاه کردن معنی Opcode ها رو بفهمه
به همین خاطر Opcode ها رو به شکل های مختلف مخفی میکنه
مثلا ممکنه Opcode واقعی این باشه:
0x8F
ولی قبل از اجرا این عملیات روی اون انجام بشه:
Opcode ^= 0xA5
بعد از XOR شدن تازه مقدار واقعی به دست میاد
یا ممکنه Opcode اصلا مستقیم داخل بایت کد ذخیره نشده باشه
مثلا قبل از استفاده از روی یک جدول ترجمه عبور کنه
0x3C → LOAD
0x91 → ADD
0xE7 → JMP
در این حالت اگر فقط به بایتکد نگاه کنید هیچ معنی خاصی نمیبینید
یکی دیگه از روشهای رایج اینه که اندازه Opcode ها ثابت نباشه
مثلا:
Opcode
Operand
Operand
یا:
Opcode
Operand
یا حتی:
Opcode
Operand
Operand
Operand
یعنی هر Opcode تعداد متفاوتی Operand داره اگر تحلیلگر این موضوع رو متوجه نشه از همون دستور اول کل بایت کد رو اشتباه تفسیر میکنه بعضی ماشینهای مجازی حتی Opcode ها رو موقع اجرا تولید میکنن یعنی مقداری که Dispatcher میبینه همون مقداری نیست که داخل فایل ذخیره شده به همین خاطر یکی از اولین کارهای تحلیلگر اینه که مسیر رسیدن Opcode به Dispatcher رو دنبال کنه اگر قبل از Dispatcher عملیاتی مثل XOR، ADD، SUB یا چرخش بیت ها انجام بشه احتمال زیادی وجود داره که Opcode ها رمزگذاری شده باشن هدف از Opcode Encoding فقط سختتر کردن تحلیل نیست باعث میشه ابزارهایی مثل IDA یا Ghidra هم نتونن به راحتی منطق ماشین مجازی رو تشخیص بدن به همین دلیل تحلیلگرها معمولا قبل از اینکه سراغ Handler ها برن سعی میکنن بفهمن Opcode ها دقیقا چطور Decode میشن اگر این مرحله رو درست انجام بدید ادامه فرایند Devirtualization خیلی ساده تر میشه
تمرین:
فرض کنید بایت کد زیر رو دارید:
8F 91 E7
و میدونید قبل از اجرا هر Opcode با
0xA5 عمل XOR میشهاول مقدار واقعی هر Opcode رو حساب کنید
بعد فرض کنید نتیجه این جدول باشه:
2A = LOAD
34 = ADD
42 = PRINT
سعی کنید مسیر اجرای ماشین مجازی رو روی کاغذ باز سازی کنید
@reverseengine
ReverseEngineering
Opcode Encoding چرا Opcode ها اینقدر عجیب به نظر میرسن؟ تا اینجا فرض کردیم Opcode ها خیلی ساده هستن مثلا: 01 = LOAD 02 = ADD 03 = JMP ولی توی ماشینهای مجازی واقعی تقریبا هیچ وقت اوضاع اینقدر ساده نیست سازنده محافظ نمیخواد تحلیلگر با چند دقیقه نگاه کردن…
Opcode Encoding
Why do Opcodes look so weird?
So far we have assumed that the Opcodes are very simple
For example:
But in real virtual machines things are almost never that simple
The manufacturer of the protection does not want the analyst to understand the meaning of the Opcodes by looking at them for a few minutes
That is why they hide the Opcodes in different ways
For example, the real Opcode may be:
But before performing this operation on it:
The real value is obtained only after XORing
Or the Opcode may not be stored directly in the bytecode at all
For example, it passes through a translation table before use
In this case, if you only look at the bytecode, you will not see any special meaning
Another common method is to make the Opcodes not have a fixed size
For example:
That is, each Opcode has a different number of Operands. If the analyst does not understand this, he will misinterpret the entire bytecode from the very first instruction. Some virtual machines even generate Opcodes at runtime, meaning that the value that the Dispatcher sees is not the same value that is stored in the file. Therefore, one of the first tasks of the analyst is to follow the path of the Opcode to the Dispatcher. If an operation such as XOR, ADD, SUB, or bit rotation is performed before the Dispatcher, there is a high probability that the Opcodes are encrypted. The purpose of Opcode Encoding is not just to make the analysis more difficult. It also makes tools such as IDA or Ghidra unable to easily recognize the logic of the virtual machine. Therefore, analysts usually try to understand exactly how the Opcodes are decoded before moving on to the Handlers. If you do this step correctly, the rest of the Devirtualization process will be much easier.
Exercise:
Suppose you have the following bytecode:
And you know that before Each Opcode can be XORed with 0xA5
First calculate the actual value of each Opcode
Then assume the result is this table:
Try to reconstruct the path of the virtual machine execution on paper
@reverseengine
Why do Opcodes look so weird?
So far we have assumed that the Opcodes are very simple
For example:
01 = LOAD
02 = ADD
03 = JMP
But in real virtual machines things are almost never that simple
The manufacturer of the protection does not want the analyst to understand the meaning of the Opcodes by looking at them for a few minutes
That is why they hide the Opcodes in different ways
For example, the real Opcode may be:
0x8F
But before performing this operation on it:
Opcode ^= 0xA5
The real value is obtained only after XORing
Or the Opcode may not be stored directly in the bytecode at all
For example, it passes through a translation table before use
0x3C → LOAD
0x91 → ADD
0xE7 → JMP
In this case, if you only look at the bytecode, you will not see any special meaning
Another common method is to make the Opcodes not have a fixed size
For example:
Opcode
Operand
Operand
Or:
Opcode
Operand
Or Even:
Opcode
Operand
Operand
Operand
That is, each Opcode has a different number of Operands. If the analyst does not understand this, he will misinterpret the entire bytecode from the very first instruction. Some virtual machines even generate Opcodes at runtime, meaning that the value that the Dispatcher sees is not the same value that is stored in the file. Therefore, one of the first tasks of the analyst is to follow the path of the Opcode to the Dispatcher. If an operation such as XOR, ADD, SUB, or bit rotation is performed before the Dispatcher, there is a high probability that the Opcodes are encrypted. The purpose of Opcode Encoding is not just to make the analysis more difficult. It also makes tools such as IDA or Ghidra unable to easily recognize the logic of the virtual machine. Therefore, analysts usually try to understand exactly how the Opcodes are decoded before moving on to the Handlers. If you do this step correctly, the rest of the Devirtualization process will be much easier.
Exercise:
Suppose you have the following bytecode:
8F 91 E7
And you know that before Each Opcode can be XORed with 0xA5
First calculate the actual value of each Opcode
Then assume the result is this table:
2A = LOAD
34 = ADD
42 = PRINT
Try to reconstruct the path of the virtual machine execution on paper
@reverseengine
PCB (Process Control Block)
شناسنامه هر Process
تا اینجا یاد گرفتیم Process ساخته میشوه اجرا میشه و بین حالت های مختلف جا به جا میشه
اما یک سؤال مهم:
سیستمعامل از کجا میفهه هر Process چه وضعیتی داره؟
مثلا از کجا میدونه:
الان در حال اجراست؟
چقدر حافظه گرفته؟
چند Thread داره؟
چه فایلهایی رو باز کرده؟
جواب همه این سؤالها یک چیزه:
PCB (Process Control Block)
PCB
رو میتونیم مثل پرونده یا شناسنامه یک Process تصور کنیم
هر Process که ساخته میشه سیستم عامل یک PCB هم برای اون ایجاد میکنه
داخل این ساختار تمام اطلاعات لازم برای مدیریت اون Process نگهداری میشه
داخل PCB چه اطلاعاتی وجود داره؟
شناسه پردازه (PID)
هر Process یک شماره منحصر به فرد داره
مثلا:
سیستم عامل از این شناسه برای تشخیص Process ها استفاده میکنه
وضعیت Process
همان State هایی که پست قبل یاد گرفتیم:
اطلاعات CPU
وقتی سیستم عامل اجرای یک Process رو متوقف میکنه باید بدونه بعدا از کجا ادامه بده
برای همین اطلاعاتی مثل:
مقدار Register ها
Program Counter
Stack Pointer
داخل PCB ذخیره میشن
اطلاعات حافظه
سیستمعامل باید بدونه:
حافظه Process کجاست؟
Heap کجاست؟
Stack کجاست؟
Page Table
مربوط به این Process چیه؟
همه این اطلاعات داخل PCB ثبت میشن
اطلاعات فایل ها
اگر Process چند فایل رو باز کرده باشه سیستم عامل باید اونها رو مدیریت کنه
مثلا:
فایل متنی
سوکت شبکه
پرینتر
این اطلاعات هم در PCB نگهداری میشن
چرا PCB مهمه؟
فرض کنید CPU در حال اجرای Chrome هست
یهو سیستم عامل تصمیم میگیره به Telegram زمان اجرا بده
قبل از این جا به جایی باید وضعیت Chrome ذخیره بشه تا بعدا دقیقا از همون نقطه ادامه بده
این اطلاعات داخل PCB ذخیره میشن
به این کار Context Switch میگن
یک مثال ساده:
فرض کنید داری یک بازی انجام میدید
اگر بازی رو Pause کنید انتظار دارید بعد از Resume دقیقا از همون لحظه ادامه پیدا کنه
سیستم عامل هم برای Process ها همین کار رو انجام میده
قبل از اینکه اجرای یک Process متوقف بشه وضعیت اون رو داخل PCB ذخیره میکنه تا بعدا بدون مشکل ادامه پیدا کنه
چرا برای مهندسی معکوس مهمه؟
در مهندسی معکوس شاید مستقیما با PCB کار نکنید اما بسیاری از مفاهیم مهم به اون وابسته ان:
اگر بدونید سیستم عامل چه اطلاعاتی از هر Process نگه میداره تحلیل رفتار برنامه ها براتون خیلی راحت تر میشه
هر Process یک PCB داره که مثل شناسنامه اون عمل میکنه
داخل PCB اطلاعاتی مثل:
PID
وضعیت Process
Register ها
اطلاعات حافظه
فایلهای باز
اطلاعات زمانبندی
ذخیره میشن
بدون PCB سیستم عامل نمیتونه Process ها رو مدیریت کنه
@reverseengine
شناسنامه هر Process
تا اینجا یاد گرفتیم Process ساخته میشوه اجرا میشه و بین حالت های مختلف جا به جا میشه
اما یک سؤال مهم:
سیستمعامل از کجا میفهه هر Process چه وضعیتی داره؟
مثلا از کجا میدونه:
الان در حال اجراست؟
چقدر حافظه گرفته؟
چند Thread داره؟
چه فایلهایی رو باز کرده؟
جواب همه این سؤالها یک چیزه:
PCB (Process Control Block)
PCB
رو میتونیم مثل پرونده یا شناسنامه یک Process تصور کنیم
هر Process که ساخته میشه سیستم عامل یک PCB هم برای اون ایجاد میکنه
داخل این ساختار تمام اطلاعات لازم برای مدیریت اون Process نگهداری میشه
داخل PCB چه اطلاعاتی وجود داره؟
شناسه پردازه (PID)
هر Process یک شماره منحصر به فرد داره
مثلا:
Process A → PID = 1200
Process B → PID = 2456
سیستم عامل از این شناسه برای تشخیص Process ها استفاده میکنه
وضعیت Process
همان State هایی که پست قبل یاد گرفتیم:
Newوضعیت فعلی Process داخل PCB ذخیره میشه
Ready
Running
Waiting
Terminated
اطلاعات CPU
وقتی سیستم عامل اجرای یک Process رو متوقف میکنه باید بدونه بعدا از کجا ادامه بده
برای همین اطلاعاتی مثل:
مقدار Register ها
Program Counter
Stack Pointer
داخل PCB ذخیره میشن
اطلاعات حافظه
سیستمعامل باید بدونه:
حافظه Process کجاست؟
Heap کجاست؟
Stack کجاست؟
Page Table
مربوط به این Process چیه؟
همه این اطلاعات داخل PCB ثبت میشن
اطلاعات فایل ها
اگر Process چند فایل رو باز کرده باشه سیستم عامل باید اونها رو مدیریت کنه
مثلا:
فایل متنی
سوکت شبکه
پرینتر
این اطلاعات هم در PCB نگهداری میشن
چرا PCB مهمه؟
فرض کنید CPU در حال اجرای Chrome هست
یهو سیستم عامل تصمیم میگیره به Telegram زمان اجرا بده
قبل از این جا به جایی باید وضعیت Chrome ذخیره بشه تا بعدا دقیقا از همون نقطه ادامه بده
این اطلاعات داخل PCB ذخیره میشن
به این کار Context Switch میگن
یک مثال ساده:
فرض کنید داری یک بازی انجام میدید
اگر بازی رو Pause کنید انتظار دارید بعد از Resume دقیقا از همون لحظه ادامه پیدا کنه
سیستم عامل هم برای Process ها همین کار رو انجام میده
قبل از اینکه اجرای یک Process متوقف بشه وضعیت اون رو داخل PCB ذخیره میکنه تا بعدا بدون مشکل ادامه پیدا کنه
چرا برای مهندسی معکوس مهمه؟
در مهندسی معکوس شاید مستقیما با PCB کار نکنید اما بسیاری از مفاهیم مهم به اون وابسته ان:
Context Switch
Process Scheduling
Thread Management
Kernel Debugging
Windows Internals
اگر بدونید سیستم عامل چه اطلاعاتی از هر Process نگه میداره تحلیل رفتار برنامه ها براتون خیلی راحت تر میشه
هر Process یک PCB داره که مثل شناسنامه اون عمل میکنه
داخل PCB اطلاعاتی مثل:
PID
وضعیت Process
Register ها
اطلاعات حافظه
فایلهای باز
اطلاعات زمانبندی
ذخیره میشن
بدون PCB سیستم عامل نمیتونه Process ها رو مدیریت کنه
@reverseengine
ReverseEngineering
PCB (Process Control Block) شناسنامه هر Process تا اینجا یاد گرفتیم Process ساخته میشوه اجرا میشه و بین حالت های مختلف جا به جا میشه اما یک سؤال مهم: سیستمعامل از کجا میفهه هر Process چه وضعیتی داره؟ مثلا از کجا میدونه: الان در حال اجراست؟ چقدر حافظه…
PCB (Process Control Block) The ID of Each Process
So far we have learned that a Process is created, executed, and switched between different states
But an important question:
How does the operating system know what state each Process is in?
For example, how does it know:
Is it currently running?
How much memory does it take?
How many threads does it have?
What files has it opened?
The answer to all these questions is the same:
We can imagine the PCB as a file or ID of a Process.
For each Process that is created, the operating system also creates a PCB for it.
Inside this structure, all the information necessary to manage that Process is stored.
What information is inside the PCB?
Each process has a unique number
For example:
The operating system uses this ID to identify processes
Process State
Same states as we learned in the previous post:
The current state of the process is stored in the PCB
CPU Information
When the operating system stops executing a process, it needs to know where to continue next
For this reason, information such as:
The operating system needs to know:
Where is the process memory?
Where is the heap?
Where is the stack?
Page Table
What is related to this process?
All this information is recorded in the PCB
File information
If a Process has opened several files, the operating system must manage them
For example:
Why is the PCB important?
Suppose the CPU is running Chrome
An operating system decides to give Telegram time to run
Before this, the state of Chrome must be saved somewhere so that it can continue exactly from the same point later
This information is stored in the PCB
This is called Context Switch
A simple example:
Suppose you are playing a game
If you pause the game, you expect it to continue exactly from that moment after Resume
The operating system does the same for Processes
Before a Process is stopped, it saves its state in the PCB so that it can continue later without any problems
Why is it important for reverse engineering?
In reverse engineering, you may not work directly with the PCB, but many important concepts depend on it:
If you know what information the operating system keeps about each process, analyzing the behavior of programs will be much easier for you
Each process has a PCB that acts as its identity card
Inside the PCB, information such as:
Without the PCB, the operating system cannot manage processes
@reverseengine
So far we have learned that a Process is created, executed, and switched between different states
But an important question:
How does the operating system know what state each Process is in?
For example, how does it know:
Is it currently running?
How much memory does it take?
How many threads does it have?
What files has it opened?
The answer to all these questions is the same:
PCB (Process Control Block)
We can imagine the PCB as a file or ID of a Process.
For each Process that is created, the operating system also creates a PCB for it.
Inside this structure, all the information necessary to manage that Process is stored.
What information is inside the PCB?
Process ID (PID)
Each process has a unique number
For example:
Process A → PID = 1200
Process B → PID = 2456
The operating system uses this ID to identify processes
Process State
Same states as we learned in the previous post:
New
Ready
Running
Waiting
Terminated
The current state of the process is stored in the PCB
CPU Information
When the operating system stops executing a process, it needs to know where to continue next
For this reason, information such as:
Register values
Program Counter
Stack Pointer
are stored in the PCB
Memory Information
The operating system needs to know:
Where is the process memory?
Where is the heap?
Where is the stack?
Page Table
What is related to this process?
All this information is recorded in the PCB
File information
If a Process has opened several files, the operating system must manage them
For example:
Text file
Network socket
Printer
This information is also stored in the PCB
Why is the PCB important?
Suppose the CPU is running Chrome
An operating system decides to give Telegram time to run
Before this, the state of Chrome must be saved somewhere so that it can continue exactly from the same point later
This information is stored in the PCB
This is called Context Switch
A simple example:
Suppose you are playing a game
If you pause the game, you expect it to continue exactly from that moment after Resume
The operating system does the same for Processes
Before a Process is stopped, it saves its state in the PCB so that it can continue later without any problems
Why is it important for reverse engineering?
In reverse engineering, you may not work directly with the PCB, but many important concepts depend on it:
Context Switch
Process Scheduling
Thread Management
Kernel Debugging
Windows Internals
If you know what information the operating system keeps about each process, analyzing the behavior of programs will be much easier for you
Each process has a PCB that acts as its identity card
Inside the PCB, information such as:
PID
Process status
Registers
Memory information
Open files
Scheduling information
Without the PCB, the operating system cannot manage processes
@reverseengine
Celebrating our 2025 open-source contributions
https://blog.trailofbits.com/2026/01/30/celebrating-our-2025-open-source-contributions
@reverseengine
https://blog.trailofbits.com/2026/01/30/celebrating-our-2025-open-source-contributions
@reverseengine
The Trail of Bits Blog
Celebrating our 2025 open-source contributions
Trail of Bits engineers contributed over 375 merged pull requests to more than 90 open-source projects in 2025, including significant work on Sigstore rekor-monitor, the Rust compiler and Clippy, pyca/cryptography’s ASN.1 API, hevm performance optimizations…
Binary type inference in Ghidra
https://blog.trailofbits.com/2024/02/07/binary-type-inference-in-ghidra
@reverseengine
https://blog.trailofbits.com/2024/02/07/binary-type-inference-in-ghidra
@reverseengine
The Trail of Bits Blog
Binary type inference in Ghidra
Trail of Bits is releasing BTIGhidra, a Ghidra extension that helps reverse engineers by inferring type information from binaries. The analysis is inter-procedural, propagating and resolving type constraints between functions while consuming user input to…
I Broke My Filesystem 35 Different Ways to Prove My Recovery Tool Works
https://levelup.gitconnected.com/i-broke-my-filesystem-35-different-ways-to-prove-my-recovery-tool-works-dc3d972195e1
https://levelup.gitconnected.com/i-broke-my-filesystem-35-different-ways-to-prove-my-recovery-tool-works-dc3d972195e1
Medium
I Broke My Filesystem 35 Different Ways to Prove My Recovery Tool Works
How I reverse-engineered Apple’s filesystem, built a recovery tool in C and Python, and achieved a 98.5% recovery rate
How I Recovered a Password from a Linux Binary Using Ghidra
https://revs3k.medium.com/how-i-recovered-a-password-from-a-linux-binary-using-ghidra-625b51561751
https://revs3k.medium.com/how-i-recovered-a-password-from-a-linux-binary-using-ghidra-625b51561751
Medium
How I Recovered a Password from a Linux Binary Using Ghidra
Introduction
دوستان ببخشید اگه نمیرسم پست بزارم درگیر یک پروژه هستم سعیمو میکنم از فردا دوباره بزارم خیلی عذر میخام🩶
Sorry guys if I don't have time to post, I'm busy with a project, I'll try to post again tomorrow. I apologize🖤
Sorry guys if I don't have time to post, I'm busy with a project, I'll try to post again tomorrow. I apologize🖤
❤6
Context Switch
وقتی CPU بین Process ها جا به جا میشه
تا اینجا فهمیدیم هر Process یک PCB داره که تمام اطلاعات مهم داخلش ذخیره میشه
حالا سوال اصلی:
اگر CPU در حال اجرای Chrome باشه چجوری یهو میتونه Telegram رو اجرا کنه
دوباره از اول شروع میکنه؟
نه
سیستم عامل از مکانیزمی به نام Context Switch استفاده میکنه
Context Switch
به زبان ساده:
Context Switch
یعنی ذخیره کردن وضعیت Process فعلی و اپلود وضعیت Process بعدی
به این ترتیب هر Process میتونه بعدا دقیقا از همون جایی که متوقف شده بود ادامه پیدا کنه
یک مثال ساده
فرض کنید دارید یک کتاب میخونید
به صفحه 87 میرسید که تلفنتون زنگ میخوره
یک بوکمارک (Bookmark) داخل کتاب میذارید و جواب تلفن رو میدید
بعد از چند دقیقه برمیگردید و دقیقا از صفحه 87 ادامه میدید
اگر بوکمارک نمیذاشتید باید دوباره دنبال صفحه ای که بودید میگشتید
PCB
دقیقا نقش همین بوکمارک رو برای Process ها داره
هنگام Context Switch چه اتفاقی میوفته؟
فرض کنید CPU در حال اجرای Process A هست
سیستم عامل تصمیم میگیره حالا Process B اجرا بشه
مراحل به این صورته:
توقف Process A
اجرای Process A موقتا متوقف میشه
ذخیره وضعیت Process A
سیستمعامل اطلاعات مهم رو داخل PCB اون ذخیره میکنه
مثل:
مقدار Register ها
محل اجرای فعلی برنامه Program Counter
Stack Pointer
وضعیت CPU
اپلود Process B
حالا اطلاعات Process B از PCB اون خونده میشن
Register
ها Stack و بقیه اطلاعات دوباره داخل CPU قرار میگیرن
ادامه اجرای Process B
CPU
اجرای Process B را از همون نقطه ای که قبلا متوقف شده ادامه میده نه از اول
چرا Context Switch لازمه؟
فرض کنید فقط یک هسته CPU داریو و این برنامه ها بازن:
Chrome
Telegram
Spotify
Notepad
اگر Context Switch وجود نداشت اولین برنامه تمام فضای CPU رو میگرفت و بقیه هیچ وقت اجرا نمیشدن
سیستمعامل با جا به جایی سریع بین Process ها باعث میشه احساس کنیم همه برنامه ها همزمان در حال اجرا هستن
هر بار که Context Switch انجام میشه سیستم عامل باید:
اطلاعات Process قبلی رو ذخیره میکنه
اطلاعات Process جدید رو اپلود میکنه
بعضی وقتا کش CPU Cache و TLB هم تحت تاثیر قرار میگیره
این کار زمان و منابع CPU رو مصرف میکنه
به همین دلیل اگر تعداد Context Switch ها خیلی زیاد بشه ممکنه عملکرد سیستم کاهش پیدا کنه
چرا برای مهندسی معکوس مهمه؟
وقتی در حال دیباگ یک برنامه هستید ممکنه اینا رو ببینید:
اجرای برنامه یهو متوقف شد
Thread
دیگه ای شروع به اجرا کرد
دوباره اجرای قبلی ادامه پیدا کرد
درک Context Switch کمک میکنه بفهمید چرا این جا به جایی ها اتفاق میوفتن و چرا ترتیب اجرای کد همیشه خطی نیست
این مفهوم همچنین پایه ای برای درک مباحث پیشرفته تر مثل:
Context Switch:
ذخیره وضعیت Process فعلی در PCB
اپلود وضعیت Process بعدی از PCB
ادامه اجرای Process جدید از همون نقطه قبلی
این مکانیزم باعث میشه چندین برنامه روی یک CPU بهصورت روون اجرا بشن
@reverseengine
وقتی CPU بین Process ها جا به جا میشه
تا اینجا فهمیدیم هر Process یک PCB داره که تمام اطلاعات مهم داخلش ذخیره میشه
حالا سوال اصلی:
اگر CPU در حال اجرای Chrome باشه چجوری یهو میتونه Telegram رو اجرا کنه
دوباره از اول شروع میکنه؟
نه
سیستم عامل از مکانیزمی به نام Context Switch استفاده میکنه
Context Switch
به زبان ساده:
Context Switch
یعنی ذخیره کردن وضعیت Process فعلی و اپلود وضعیت Process بعدی
به این ترتیب هر Process میتونه بعدا دقیقا از همون جایی که متوقف شده بود ادامه پیدا کنه
یک مثال ساده
فرض کنید دارید یک کتاب میخونید
به صفحه 87 میرسید که تلفنتون زنگ میخوره
یک بوکمارک (Bookmark) داخل کتاب میذارید و جواب تلفن رو میدید
بعد از چند دقیقه برمیگردید و دقیقا از صفحه 87 ادامه میدید
اگر بوکمارک نمیذاشتید باید دوباره دنبال صفحه ای که بودید میگشتید
PCB
دقیقا نقش همین بوکمارک رو برای Process ها داره
هنگام Context Switch چه اتفاقی میوفته؟
فرض کنید CPU در حال اجرای Process A هست
سیستم عامل تصمیم میگیره حالا Process B اجرا بشه
مراحل به این صورته:
توقف Process A
اجرای Process A موقتا متوقف میشه
ذخیره وضعیت Process A
سیستمعامل اطلاعات مهم رو داخل PCB اون ذخیره میکنه
مثل:
مقدار Register ها
محل اجرای فعلی برنامه Program Counter
Stack Pointer
وضعیت CPU
اپلود Process B
حالا اطلاعات Process B از PCB اون خونده میشن
Register
ها Stack و بقیه اطلاعات دوباره داخل CPU قرار میگیرن
ادامه اجرای Process B
CPU
اجرای Process B را از همون نقطه ای که قبلا متوقف شده ادامه میده نه از اول
چرا Context Switch لازمه؟
فرض کنید فقط یک هسته CPU داریو و این برنامه ها بازن:
Chrome
Telegram
Spotify
Notepad
اگر Context Switch وجود نداشت اولین برنامه تمام فضای CPU رو میگرفت و بقیه هیچ وقت اجرا نمیشدن
سیستمعامل با جا به جایی سریع بین Process ها باعث میشه احساس کنیم همه برنامه ها همزمان در حال اجرا هستن
هر بار که Context Switch انجام میشه سیستم عامل باید:
اطلاعات Process قبلی رو ذخیره میکنه
اطلاعات Process جدید رو اپلود میکنه
بعضی وقتا کش CPU Cache و TLB هم تحت تاثیر قرار میگیره
این کار زمان و منابع CPU رو مصرف میکنه
به همین دلیل اگر تعداد Context Switch ها خیلی زیاد بشه ممکنه عملکرد سیستم کاهش پیدا کنه
چرا برای مهندسی معکوس مهمه؟
وقتی در حال دیباگ یک برنامه هستید ممکنه اینا رو ببینید:
اجرای برنامه یهو متوقف شد
Thread
دیگه ای شروع به اجرا کرد
دوباره اجرای قبلی ادامه پیدا کرد
درک Context Switch کمک میکنه بفهمید چرا این جا به جایی ها اتفاق میوفتن و چرا ترتیب اجرای کد همیشه خطی نیست
این مفهوم همچنین پایه ای برای درک مباحث پیشرفته تر مثل:
Scheduling
Multi-threading
Race Condition
Kernel Debugging
Context Switch:
ذخیره وضعیت Process فعلی در PCB
اپلود وضعیت Process بعدی از PCB
ادامه اجرای Process جدید از همون نقطه قبلی
این مکانیزم باعث میشه چندین برنامه روی یک CPU بهصورت روون اجرا بشن
@reverseengine
❤2
ReverseEngineering
Context Switch وقتی CPU بین Process ها جا به جا میشه تا اینجا فهمیدیم هر Process یک PCB داره که تمام اطلاعات مهم داخلش ذخیره میشه حالا سوال اصلی: اگر CPU در حال اجرای Chrome باشه چجوری یهو میتونه Telegram رو اجرا کنه دوباره از اول شروع میکنه؟ نه …
Context Switch: When the CPU switches between processes
So far we have understood that each process has a PCB that stores all the important information
Now the main question:
If the CPU is running Chrome, how can it suddenly run Telegram
Starting from scratch again?
No
The operating system uses a mechanism called Context Switch
Context Switch
In simple terms:
Context Switch
means saving the state of the current process and uploading the state of the next process
In this way, each process can continue exactly where it left off
A simple example
Suppose you are reading a book
You reach page 87 when your phone rings
You put a bookmark in the book and see the answer to the phone
After a few minutes, you come back and continue exactly from page 87
If you had not put a bookmark, you would have to look for the page you were on again
PCB
is exactly the role of this bookmark for processes
What happens during Context Switch?
Suppose CPU is executing Process A
The operating system decides to execute Process B now
The steps are as follows:
Stop Process A
Process A execution is temporarily stopped
Save Process A state
The operating system saves important information into its PCB
For example:
Register values
Current execution location of the program Program Counter
Stack Pointer
CPU state
Upload Process B
Now Process B information is read from its PCB
Registers, Stack and other information are put back into the CPU
Continue execution of Process B
CPU resumes execution of Process B from the same point where it was stopped earlier, not from the beginning
Why is Context Switch necessary?
Suppose you have only one CPU core and these programs are open:
Chrome
Telegram
Spotify
Notepad
If there was no Context Switch, the first program would take up all the CPU space and the others would never run
The operating system makes it feel like all programs are running at the same time by quickly switching between processes
Every time a Context Switch is performed, the operating system must:
Save the previous process information
Upload the new process information
Sometimes the CPU Cache and TLB are also affected
This consumes time and CPU resources
Therefore, if the number of Context Switches becomes too large, the system performance may decrease
Why is it important for reverse engineering?
When you are debugging a program, you may see:
Program execution suddenly stops
Thread
starts running again
Previous execution continues
Understanding Context Switch helps you understand why these switchings happen and why the order of code execution is not always linear
This concept also provides a foundation for understanding more advanced topics such as:
Scheduling
Multi-threading
Race Condition
Kernel Debugging
Context Switch:
Save current process state to PCB
Upload next process state from PCB
Continue execution of new process from previous point
This mechanism allows multiple programs to run smoothly on a single CPU
@reverseengine
So far we have understood that each process has a PCB that stores all the important information
Now the main question:
If the CPU is running Chrome, how can it suddenly run Telegram
Starting from scratch again?
No
The operating system uses a mechanism called Context Switch
Context Switch
In simple terms:
Context Switch
means saving the state of the current process and uploading the state of the next process
In this way, each process can continue exactly where it left off
A simple example
Suppose you are reading a book
You reach page 87 when your phone rings
You put a bookmark in the book and see the answer to the phone
After a few minutes, you come back and continue exactly from page 87
If you had not put a bookmark, you would have to look for the page you were on again
PCB
is exactly the role of this bookmark for processes
What happens during Context Switch?
Suppose CPU is executing Process A
The operating system decides to execute Process B now
The steps are as follows:
Stop Process A
Process A execution is temporarily stopped
Save Process A state
The operating system saves important information into its PCB
For example:
Register values
Current execution location of the program Program Counter
Stack Pointer
CPU state
Upload Process B
Now Process B information is read from its PCB
Registers, Stack and other information are put back into the CPU
Continue execution of Process B
CPU resumes execution of Process B from the same point where it was stopped earlier, not from the beginning
Why is Context Switch necessary?
Suppose you have only one CPU core and these programs are open:
Chrome
Telegram
Spotify
Notepad
If there was no Context Switch, the first program would take up all the CPU space and the others would never run
The operating system makes it feel like all programs are running at the same time by quickly switching between processes
Every time a Context Switch is performed, the operating system must:
Save the previous process information
Upload the new process information
Sometimes the CPU Cache and TLB are also affected
This consumes time and CPU resources
Therefore, if the number of Context Switches becomes too large, the system performance may decrease
Why is it important for reverse engineering?
When you are debugging a program, you may see:
Program execution suddenly stops
Thread
starts running again
Previous execution continues
Understanding Context Switch helps you understand why these switchings happen and why the order of code execution is not always linear
This concept also provides a foundation for understanding more advanced topics such as:
Scheduling
Multi-threading
Race Condition
Kernel Debugging
Context Switch:
Save current process state to PCB
Upload next process state from PCB
Continue execution of new process from previous point
This mechanism allows multiple programs to run smoothly on a single CPU
@reverseengine
❤1🥰1
Instruction Handler
داخل هر Handler دقیقا چه اتفاقی میوفته؟
تا اینجا فهمیدیم هر Opcode یک Handler مخصوص خودش رو داره
اما سوال اصلی اینجاست
وقتی Dispatcher کنترل رو به یک Handler میده اون Handler دقیقا چه کاری انجام میده؟
فرض کنید Opcode مربوط به جمع باشه
در ظاهر فقط یک عدد مثل این میبینیم:
0x27
ولی پشت این عدد چندین دستور اسمبلی اجرا میشه
مثلا ممکنه Handler این کار ها رو انجام بده:
خواندن Operand اول
خواندن Operand دوم
انجام عملیات جمع
ذخیره نتیجه
برگشت به Dispatcher
همه این مراحل داخل چندین دستور اسمبلی پیاده سازی میشن
برای همین وقتی Handler رو داخل IDA یا Ghidra باز میکنید معمولا فقط چند دستور نمیبینید
ممکن است ده ها یا حتی صد ها دستور وجود داشته باشه
کاری که تحلیلگر انجام میده این نیست که همه دستورها رو حفظ کنه
بلکه سعی میکنه رفتار کلی Handler رو بفهمه
مثلا بعد از چند دقیقه بررسی به این نتیجه میرسه:
این Handler فقط داده رو جا به جا میکنه
این یکی مقدار ها رو با هم جمع میکنه
این یکی عمل XOR رو انجام میده
این یکی پرش شرطی انجام میده
این یکی مقدار رو داخل Stack یا Virtual Register ذخیره میکنه
وقتی این دسته بندی کامل بشه کم کم جدول Opcode ها ساخته میشه
به این جدول معمولا Opcode Semantics میگن
یعنی مشخص میکنیم هر Opcode چه معنی و چه رفتاری داره
یکی از اشتباهات رایج افراد تازه کار اینه که از اولین دستور اسمبلی شروع میکنن و خط به خط جلو میرن
تحلیلگر های حرفهای برعکس عمل میکنن
اول ورودی Handler رو پیدا میکنن
بعد خروجی Handler رو بررسی میکنن
در آخر مسیر بین این دو رو تحلیل میکنن
این روش باعث میشه خیلی سریع تر بفهمن Handler چه کاری انجام میده
تمرین:
فرض کنید بعد از بررسی یک Handler فقط این اطلاعات رو به دست آوردید:
ورودی:
V0 = 15
V1 = 8
خروجی:
V0 = 23
V1 = 8
بدون اینکه اسمبلی رو ببینید حدس بزنید این Handler چه کاری انجام داده
باید بتونید فقط از روی ورودی و خروجی رفتار Handler رو تشخیص بدید
@reverseengine
داخل هر Handler دقیقا چه اتفاقی میوفته؟
تا اینجا فهمیدیم هر Opcode یک Handler مخصوص خودش رو داره
اما سوال اصلی اینجاست
وقتی Dispatcher کنترل رو به یک Handler میده اون Handler دقیقا چه کاری انجام میده؟
فرض کنید Opcode مربوط به جمع باشه
در ظاهر فقط یک عدد مثل این میبینیم:
0x27
ولی پشت این عدد چندین دستور اسمبلی اجرا میشه
مثلا ممکنه Handler این کار ها رو انجام بده:
خواندن Operand اول
خواندن Operand دوم
انجام عملیات جمع
ذخیره نتیجه
برگشت به Dispatcher
همه این مراحل داخل چندین دستور اسمبلی پیاده سازی میشن
برای همین وقتی Handler رو داخل IDA یا Ghidra باز میکنید معمولا فقط چند دستور نمیبینید
ممکن است ده ها یا حتی صد ها دستور وجود داشته باشه
کاری که تحلیلگر انجام میده این نیست که همه دستورها رو حفظ کنه
بلکه سعی میکنه رفتار کلی Handler رو بفهمه
مثلا بعد از چند دقیقه بررسی به این نتیجه میرسه:
این Handler فقط داده رو جا به جا میکنه
این یکی مقدار ها رو با هم جمع میکنه
این یکی عمل XOR رو انجام میده
این یکی پرش شرطی انجام میده
این یکی مقدار رو داخل Stack یا Virtual Register ذخیره میکنه
وقتی این دسته بندی کامل بشه کم کم جدول Opcode ها ساخته میشه
به این جدول معمولا Opcode Semantics میگن
یعنی مشخص میکنیم هر Opcode چه معنی و چه رفتاری داره
یکی از اشتباهات رایج افراد تازه کار اینه که از اولین دستور اسمبلی شروع میکنن و خط به خط جلو میرن
تحلیلگر های حرفهای برعکس عمل میکنن
اول ورودی Handler رو پیدا میکنن
بعد خروجی Handler رو بررسی میکنن
در آخر مسیر بین این دو رو تحلیل میکنن
این روش باعث میشه خیلی سریع تر بفهمن Handler چه کاری انجام میده
تمرین:
فرض کنید بعد از بررسی یک Handler فقط این اطلاعات رو به دست آوردید:
ورودی:
V0 = 15
V1 = 8
خروجی:
V0 = 23
V1 = 8
بدون اینکه اسمبلی رو ببینید حدس بزنید این Handler چه کاری انجام داده
باید بتونید فقط از روی ورودی و خروجی رفتار Handler رو تشخیص بدید
@reverseengine
❤2
ReverseEngineering
Instruction Handler داخل هر Handler دقیقا چه اتفاقی میوفته؟ تا اینجا فهمیدیم هر Opcode یک Handler مخصوص خودش رو داره اما سوال اصلی اینجاست وقتی Dispatcher کنترل رو به یک Handler میده اون Handler دقیقا چه کاری انجام میده؟ فرض کنید Opcode مربوط به جمع باشه…
Instruction Handler
What exactly happens inside each Handler?
So far we have understood that each Opcode has its own Handler
But here is the main question
When the Dispatcher gives control to a Handler, what exactly does that Handler do?
Suppose the Opcode is related to addition
At first glance, we see just one number like this:
0x27
But behind this number, several assembly instructions are executed
For example, the Handler may do the following:
Read the first Operand
Read the second Operand
Perform the addition operation
Save the result
Return to Dispatcher
All these steps are implemented in several assembly instructions
That's why when you open the Handler in IDA or Ghidra, you usually don't see just a few instructions
There may be dozens or even hundreds of instructions
What the analyst does is not to memorize all the instructions
But it tries to understand the general behavior of the Handler
For example, after a few minutes of examination, it comes to the following conclusion:
This Handler only moves data
This one adds values together
This one performs an XOR operation
This one performs a conditional jump
This one stores a value in the Stack or Virtual Register
When Once this classification is complete, the Opcode table is gradually created
This table is usually called Opcode Semantics
That is, we specify what each Opcode means and what behavior it has
One of the common mistakes of beginners is that they start from the first assembly instruction and proceed line by line
Professional analyzers do the opposite
First they find the Handler input
Then they examine the Handler output
Finally, they analyze the path between the two
This method makes it much faster to understand what the Handler does
Exercise:
Suppose that after examining a Handler, you only obtained this information:
Input:
V0 = 15
V1 = 8
Output:
V0 = 23
V1 = 8
Guess what this Handler did without seeing the assembly
You should be able to recognize the Handler's behavior just from the input and output
@reverseengine
What exactly happens inside each Handler?
So far we have understood that each Opcode has its own Handler
But here is the main question
When the Dispatcher gives control to a Handler, what exactly does that Handler do?
Suppose the Opcode is related to addition
At first glance, we see just one number like this:
0x27
But behind this number, several assembly instructions are executed
For example, the Handler may do the following:
Read the first Operand
Read the second Operand
Perform the addition operation
Save the result
Return to Dispatcher
All these steps are implemented in several assembly instructions
That's why when you open the Handler in IDA or Ghidra, you usually don't see just a few instructions
There may be dozens or even hundreds of instructions
What the analyst does is not to memorize all the instructions
But it tries to understand the general behavior of the Handler
For example, after a few minutes of examination, it comes to the following conclusion:
This Handler only moves data
This one adds values together
This one performs an XOR operation
This one performs a conditional jump
This one stores a value in the Stack or Virtual Register
When Once this classification is complete, the Opcode table is gradually created
This table is usually called Opcode Semantics
That is, we specify what each Opcode means and what behavior it has
One of the common mistakes of beginners is that they start from the first assembly instruction and proceed line by line
Professional analyzers do the opposite
First they find the Handler input
Then they examine the Handler output
Finally, they analyze the path between the two
This method makes it much faster to understand what the Handler does
Exercise:
Suppose that after examining a Handler, you only obtained this information:
Input:
V0 = 15
V1 = 8
Output:
V0 = 23
V1 = 8
Guess what this Handler did without seeing the assembly
You should be able to recognize the Handler's behavior just from the input and output
@reverseengine
❤1
Chunk
دقیقا چیسه و چه ساختاری داره؟
در پست قبل گفتیم Heap از بخش های کوچیکی به اسم Chunk تشکیل شده
هر بار که برنامه از malloc() استفاده میکنه Allocator یک Chunk در اختیارش قرار میده اما این Chunk فقط یه تیکه حافظه ساده نیست اطلاعات مهمی هم داخلش وجود داره
Chunk
فقط داده ی برنامه نیست
فرض کنید از Heap 100 بایت حافظه درخواست کردید شاید فکر کنید Allocator دقیقا همون 100 بایت رو بهتون میده ولی در عمل این اتفاق نمیوفته
قبل از داده ای که برنامه استفاده میکنه Allocator چند بایت برای خودش کنار میذاره این قسمت همون Metadata هست
پس ساختار یک Chunk تقریبا این شکلیه:
+------------------+
| Metadata |
+------------------+
| User Data |
| |
| |
+------------------+
برنامه فقط به بخش User Data دسترسی داره اما Allocator از Metadata برای مدیریت Heap استفاده میکنه
داخل Metadata چه اطلاعاتی وجود داره؟
بسته به نوع Allocator ممکنه فرق داشته باشه اما معمولا اطلاعاتی مثل این ها نگهداری میشن:
اندازهی Chunk
وضعیت آزاد یا اشغال بودن
اطلاعات لازم برای مدیریت حافظه
ارتباط با Chunk های کناری در بعضی Allocator ها
این اطلاعات باعث میشن Allocator بدونه هر قسمت از Heap چه وضعیتی داره
چرا اندازهی واقعی Chunk با چیزی که درخواست کرده فرق داره؟
فرض کنی این درخواست رو نوشتید:
malloc(100);
این به این معنی نیست که دقیقا 100 بایت از Heap اشغال میشه.
چون Allocator باید:
Metadata
رو ذخیره کنه
حافظه رو تراز (Alignment) کنه
اندازه ها رو به واحد های مشخص گرد کنه
برای همین ممکنه در عمل بیشتر از 100 بایت مصرف بشه
این موضوع یکی از چیزهاییه که خیلی از برنامهنویسها در ابتدا بهش توجه نمیکنن
Alignment
پردازنده دوست داره داده ها روی مرز های مشخصی از حافظه قرار بگیرن
مثلا روی سیستمهای 64 بیتی معمولا داده ها روی مضربهای 8 یا 16 بایت تراز میشن این کار باعث میشه دسترسی به حافظه سریع تر و بهینه تر باشه
به همین دلیل Allocator بعضی وقتا اندازه ی درخواستی رو کمی بزرگ تر در نظر میگیره
وقتی free() صدا زده میشه چه اتفاقی میوفته؟
خیلیها فکر میکنن با free() حافظه بلافاصله از بین میره
در واقع معمولا اینطور نیست
بیشتر Allocator ها حافظه رو فقط آزاد علامت گذاری میکنن تا بعدا دوباره از همون Chunk استفاده کنن
یعنی دادههایی که داخل اون Chunk بودن ممکنه هنوز در حافظه باقی مونده باشن فقط برنامه دیگه نباید از اونها استفاده کنه
به همین خاطر باگ هایی مثل Use-After-Free به وجود میان یعنی برنامه بعد از آزاد شدن حافظه اشتباها دوباره به همون بخش دسترسی پیدا میکنه
Heap
از بخشهایی به نام Chunk تشکیل شده که هر کدام علاوه بر فضای مورد استفاده ی برنامه اطلاعات مدیریتی هم دارن این اطلاعات به Allocator کمک میکنه تا حافظه رو مدیریت کنه همچنین حافظه ای که با free() آزاد میشه معمولا بلافاصله پاک نمیشه بلکه برای استفاده ی مجدد آماده نگه داشته میشه
@reverseengine
دقیقا چیسه و چه ساختاری داره؟
در پست قبل گفتیم Heap از بخش های کوچیکی به اسم Chunk تشکیل شده
هر بار که برنامه از malloc() استفاده میکنه Allocator یک Chunk در اختیارش قرار میده اما این Chunk فقط یه تیکه حافظه ساده نیست اطلاعات مهمی هم داخلش وجود داره
Chunk
فقط داده ی برنامه نیست
فرض کنید از Heap 100 بایت حافظه درخواست کردید شاید فکر کنید Allocator دقیقا همون 100 بایت رو بهتون میده ولی در عمل این اتفاق نمیوفته
قبل از داده ای که برنامه استفاده میکنه Allocator چند بایت برای خودش کنار میذاره این قسمت همون Metadata هست
پس ساختار یک Chunk تقریبا این شکلیه:
+------------------+
| Metadata |
+------------------+
| User Data |
| |
| |
+------------------+
برنامه فقط به بخش User Data دسترسی داره اما Allocator از Metadata برای مدیریت Heap استفاده میکنه
داخل Metadata چه اطلاعاتی وجود داره؟
بسته به نوع Allocator ممکنه فرق داشته باشه اما معمولا اطلاعاتی مثل این ها نگهداری میشن:
اندازهی Chunk
وضعیت آزاد یا اشغال بودن
اطلاعات لازم برای مدیریت حافظه
ارتباط با Chunk های کناری در بعضی Allocator ها
این اطلاعات باعث میشن Allocator بدونه هر قسمت از Heap چه وضعیتی داره
چرا اندازهی واقعی Chunk با چیزی که درخواست کرده فرق داره؟
فرض کنی این درخواست رو نوشتید:
malloc(100);
این به این معنی نیست که دقیقا 100 بایت از Heap اشغال میشه.
چون Allocator باید:
Metadata
رو ذخیره کنه
حافظه رو تراز (Alignment) کنه
اندازه ها رو به واحد های مشخص گرد کنه
برای همین ممکنه در عمل بیشتر از 100 بایت مصرف بشه
این موضوع یکی از چیزهاییه که خیلی از برنامهنویسها در ابتدا بهش توجه نمیکنن
Alignment
پردازنده دوست داره داده ها روی مرز های مشخصی از حافظه قرار بگیرن
مثلا روی سیستمهای 64 بیتی معمولا داده ها روی مضربهای 8 یا 16 بایت تراز میشن این کار باعث میشه دسترسی به حافظه سریع تر و بهینه تر باشه
به همین دلیل Allocator بعضی وقتا اندازه ی درخواستی رو کمی بزرگ تر در نظر میگیره
وقتی free() صدا زده میشه چه اتفاقی میوفته؟
خیلیها فکر میکنن با free() حافظه بلافاصله از بین میره
در واقع معمولا اینطور نیست
بیشتر Allocator ها حافظه رو فقط آزاد علامت گذاری میکنن تا بعدا دوباره از همون Chunk استفاده کنن
یعنی دادههایی که داخل اون Chunk بودن ممکنه هنوز در حافظه باقی مونده باشن فقط برنامه دیگه نباید از اونها استفاده کنه
به همین خاطر باگ هایی مثل Use-After-Free به وجود میان یعنی برنامه بعد از آزاد شدن حافظه اشتباها دوباره به همون بخش دسترسی پیدا میکنه
Heap
از بخشهایی به نام Chunk تشکیل شده که هر کدام علاوه بر فضای مورد استفاده ی برنامه اطلاعات مدیریتی هم دارن این اطلاعات به Allocator کمک میکنه تا حافظه رو مدیریت کنه همچنین حافظه ای که با free() آزاد میشه معمولا بلافاصله پاک نمیشه بلکه برای استفاده ی مجدد آماده نگه داشته میشه
@reverseengine
❤1