ReverseEngineering
بخش بیستم بافر اورفلو Use After Free یکی از خطرناکترین باگهای حافظه تا اینجا یاد گرفتیم malloc حافظه میگیره و free اون رو آزاد میکنه حالا میخوایم ببینیم اگر برنامه بعد از free دوباره از همون حافظه استفاده کنه چه اتفاقی میوفته و چطور موقع مهندسی معکوس…
Part 20 Buffer Overflow
Use After Free One of the most dangerous memory bugs
So far we have learned that malloc takes memory and free frees it. Now we want to see what happens if the program uses the same memory again after free and how to detect this bug when reverse engineering
What does Use After Free
mean? Its name is quite specific. First the memory is freed. Then the program uses the same memory again, meaning the program thinks the memory is still valid, while the operating system or Heap Manager may have used that memory for something else.
A simple example:
#include <stdio.h> #include <stdlib.h>Where is the problem with this?
int main() { char *buf = malloc(32); free(buf); puts(buf); return 0; }
Here
free(buf);
The memory is freed, but a few lines later
puts(buf);
The same pointer is used. This is exactly a Use After Free
Why is it dangerous?
Because after free there is no guarantee that the previous data will remain in that memory. The data may have changed, the memory may have been reallocated, the program may crash
What should we look for when reverse engineering?
If you see a pattern like this
free(ptr);
and then a few lines down
ptr->field
or
memcpy(ptr,...)
or
printf("%s", ptr);
or any other use of ptr, you should check whether the same pointer is reused after free
An example in the decompiler
buf = malloc(64); /* ... */ free(buf); /* ... */ strcpy(buf, input);
These few lines are enough to be suspicious
How do you prevent this bug?
One of the easiest ways is to NULL the pointer after free
free(buf); buf = NULL;
Now if buf is used again somewhere
The problem will be identified much earlier
free means memory has been freed
You should not use the same pointer after free
In reverse engineering, always follow the path
malloc → free →
Check for reuse
This is one of the important patterns for finding memory management bugs
Exercise:
Open a simple program that uses malloc and free in Ghidra or IDA and check if the same pointer is used again after free or not. If you find such a pattern, analyze the reason and determine if it is really a Use After Free or it just looks like it.
@reverseengine
❤4
اینجا به یکی از مهمترین مفاهیم Binary Exploitation میرسیم
Info Leak
چرا اولین هدف یک اکسپلویت مدرنه؟
تا اینجا چند تا مکانیزم امنیتی رو شناختیم:
یا Info Leak یعنی چی؟
Info Leak
یعنی برنامه بدون اینکه قرار بوده بخشی از اطلاعات حافظه رو در اختیار کاربر قرار بده
این اطلاعات ممکنه شامل مواردی مثل:
یا حتی داده های حساس داخل حافظه
باشه در ظاهر شاید این اطلاعات مهم به نظر نرسن اما برای یک اکسپلویتر میتونن حکم نقشه ی یک ساختمون رو داشته باشن
چرا اینقدر مهمه؟
فرض کنید ASLR فعاله
یعنی آدرسها در هر بار اجرای برنامه تغییر میکنن
اگر هیچ آدرسی رو ندونید ساختن یک اکسپلویت قابل اعتماد خیلی سخت میشه
اما اگر برنامه فقط یک آدرس رو ناخواسته لو بده مهاجم میتونه از همون آدرس موقعیت بقیه بخش های حافظه رو هم محاسبه کنه
در این حالت بخش بزرگی از مزیت ASLR از بین میره
گاهی اوقات Info Leak خودش یک آسیبپذیری مستقل محسوب میشه و گاهی هم نتیجه ی یک باگ دیگه است
مثلا ممکنه:
برنامه بیشتر از حد لازم داده چاپ کنه
یک اشارهگر (Pointer) رو مستقیم نمایش بده
بخشی از حافظه رو بدون پاک کردن برگردونه
یا یک خطای منطقی باعث افشای اطلاعات بشه
چرا بیشتر اکسپلویتهای امروزی دو مرحلهای هستند؟
خیلی از حملات مدرن این شکلیاند:
مرحله اول پیدا کردن یک Info Leak
مرحله دوم استفاده از اطلاعات بهدست اومده برای دور زدن مکانیزم هایی مثل ASLR و ادامه ی حمله
به همین دلیل تعداد زیادی از تحلیلهای امنیتی که میبینید که اولین هدف نه اجرای کد بلکه پیدا کردن یک راه برای دیدن حافظه است
یک شبیه سازی ساده:
فرض کنید وارد یک شهر ناشناس شدید
اگر هیچ نقشه ای نداشته باشید پیدا کردن یک ساختمون خاص خیلی سخت میشه
اما اگر فقط آدرس یک خیابون رو بهتون بدن کم کم میتونید کل نقشه شهر رو پیدا کنید
Info Leak
دقیقا همین نقش رو در حافظه برنامهها داره
Info Leak
چرا اولین هدف یک اکسپلویت مدرنه؟
تا اینجا چند تا مکانیزم امنیتی رو شناختیم:
Stack CanaryInformation Leak
NX
ASLR
یا Info Leak یعنی چی؟
Info Leak
یعنی برنامه بدون اینکه قرار بوده بخشی از اطلاعات حافظه رو در اختیار کاربر قرار بده
این اطلاعات ممکنه شامل مواردی مثل:
آدرس یک تابع
آدرس یک متغیر
آدرس heap
آدرس stack
آدرس libc
یا حتی داده های حساس داخل حافظه
باشه در ظاهر شاید این اطلاعات مهم به نظر نرسن اما برای یک اکسپلویتر میتونن حکم نقشه ی یک ساختمون رو داشته باشن
چرا اینقدر مهمه؟
فرض کنید ASLR فعاله
یعنی آدرسها در هر بار اجرای برنامه تغییر میکنن
اگر هیچ آدرسی رو ندونید ساختن یک اکسپلویت قابل اعتماد خیلی سخت میشه
اما اگر برنامه فقط یک آدرس رو ناخواسته لو بده مهاجم میتونه از همون آدرس موقعیت بقیه بخش های حافظه رو هم محاسبه کنه
در این حالت بخش بزرگی از مزیت ASLR از بین میره
گاهی اوقات Info Leak خودش یک آسیبپذیری مستقل محسوب میشه و گاهی هم نتیجه ی یک باگ دیگه است
مثلا ممکنه:
برنامه بیشتر از حد لازم داده چاپ کنه
یک اشارهگر (Pointer) رو مستقیم نمایش بده
بخشی از حافظه رو بدون پاک کردن برگردونه
یا یک خطای منطقی باعث افشای اطلاعات بشه
چرا بیشتر اکسپلویتهای امروزی دو مرحلهای هستند؟
خیلی از حملات مدرن این شکلیاند:
مرحله اول پیدا کردن یک Info Leak
مرحله دوم استفاده از اطلاعات بهدست اومده برای دور زدن مکانیزم هایی مثل ASLR و ادامه ی حمله
به همین دلیل تعداد زیادی از تحلیلهای امنیتی که میبینید که اولین هدف نه اجرای کد بلکه پیدا کردن یک راه برای دیدن حافظه است
یک شبیه سازی ساده:
فرض کنید وارد یک شهر ناشناس شدید
اگر هیچ نقشه ای نداشته باشید پیدا کردن یک ساختمون خاص خیلی سخت میشه
اما اگر فقط آدرس یک خیابون رو بهتون بدن کم کم میتونید کل نقشه شهر رو پیدا کنید
Info Leak
دقیقا همین نقش رو در حافظه برنامهها داره
مکانیزم هایی مثل ASLR آدرس ها رو مخفی میکنن اما اگر برنامه حتی مقدار کمی از اطلاعات حافظه رو افشا کنه این محافظت تا حد زیادی بی اثر میشه به همین دلیل پیدا کردن Info Leak یکی از مهم ترین مراحل در تحلیل و درک اکسپلویت های مدرنه@reverseengine
ReverseEngineering
اینجا به یکی از مهمترین مفاهیم Binary Exploitation میرسیم Info Leak چرا اولین هدف یک اکسپلویت مدرنه؟ تا اینجا چند تا مکانیزم امنیتی رو شناختیم: Stack Canary NX ASLR Information Leak یا Info Leak یعنی چی؟ Info Leak یعنی برنامه بدون اینکه قرار بوده…
Here we come to one of the most important concepts of Binary Exploitation
Info Leak
Why is it the first target of a modern exploit?
So far we have recognized several security mechanisms:
Stack Canary
NX
ASLR
What is Information Leak or Info Leak?
Info Leak
means that the program provides some memory information to the user without being supposed to
This information may include things like:
Or even sensitive data in memory
On the surface, this information may not seem important, but for an exploiter, it can be like a blueprint of a building
Why is it so important?
Let's say ASLR is enabled
This means that the addresses change every time the program is run
If you don't know any addresses, it's very difficult to create a reliable exploit
But if the program accidentally leaks just one address, the attacker can calculate the location of other memory locations from that address
In this case, a large part of the benefit of ASLR is lost
Sometimes an Info Leak is a standalone vulnerability, and sometimes it's the result of another bug
For example, it's possible to:
The program prints more data than necessary
Display a pointer directly
Return a portion of memory without clearing
Or a logical error causes information to be leaked
Why are most exploits today two-step?
Many modern attacks look like this:
The first step is to find an Info Leak
The second step is to use the information obtained to bypass mechanisms such as ASLR and continue the attack
That is why a lot of security analysis that you see the first goal is not to execute code but to find a way to see the memory
A simple simulation:
Suppose you enter an unknown city
If you have no map, it will be very difficult to find a specific building
But if you are only given the address of a street, you can gradually find the entire map of the city
Info Leak
It plays exactly the same role in the memory of programs
Info Leak
Why is it the first target of a modern exploit?
So far we have recognized several security mechanisms:
Stack Canary
NX
ASLR
What is Information Leak or Info Leak?
Info Leak
means that the program provides some memory information to the user without being supposed to
This information may include things like:
The address of a function
The address of a variable
The heap address
The stack address
The libc address
Or even sensitive data in memory
On the surface, this information may not seem important, but for an exploiter, it can be like a blueprint of a building
Why is it so important?
Let's say ASLR is enabled
This means that the addresses change every time the program is run
If you don't know any addresses, it's very difficult to create a reliable exploit
But if the program accidentally leaks just one address, the attacker can calculate the location of other memory locations from that address
In this case, a large part of the benefit of ASLR is lost
Sometimes an Info Leak is a standalone vulnerability, and sometimes it's the result of another bug
For example, it's possible to:
The program prints more data than necessary
Display a pointer directly
Return a portion of memory without clearing
Or a logical error causes information to be leaked
Why are most exploits today two-step?
Many modern attacks look like this:
The first step is to find an Info Leak
The second step is to use the information obtained to bypass mechanisms such as ASLR and continue the attack
That is why a lot of security analysis that you see the first goal is not to execute code but to find a way to see the memory
A simple simulation:
Suppose you enter an unknown city
If you have no map, it will be very difficult to find a specific building
But if you are only given the address of a street, you can gradually find the entire map of the city
Info Leak
It plays exactly the same role in the memory of programs
Mechanisms such as ASLR hide addresses, but if the program reveals even a small amount of memory information, this protection becomes largely ineffective. This is why finding Info Leak is one of the most important steps in analyzing and understanding modern exploits@reverseengine
Virtual Register رجیستر های مجازی
تا اینجا فهمیدیم Dispatcher تصمیم میگیره کدوم Handler اجرا بشه و Handler هم کار اصلی رو انجام میده
Handler
ها داده ها رو کجا نگه میدارن؟
روی رجیستر های واقعی CPU مثل RAX و RBX؟ معمولا نه
بیشتر ماشینهای مجازی از چیزی به اسم Virtual Register استفاده میکنن
Virtual Register
یه فضای حافظه ست که نقش رجیستر های CPU رو بازی میکنه
مثلا فرض کنید ماشین مجازی 8 تا رجیستر داره:
وقتی Opcode زیر اجرا میشه:
مقدار 10 داخل V0 قرار میگیره
بعد Opcode بعدی:
عدد 20 داخل V1 ذخیره میشه
حالا Opcode بعدی:
ماشین مجازی مقدار V0 و V1 رو جمع میکنه
نتیجه دوباره داخل V0 ذخیره میشه
اگر بعدش این Opcode اجرا بشه:
عدد 30 چاپ میشه
اینجا هیچکدوم از رجیسترهای واقعی CPU مستقیما دیده نمیشن
ممکنه داخل Handlerه ا فقط چند دستور مثل این ببینید:
در نگاه اول انگار برنامه فقط داره با حافظه کار میکنه
ولی در واقع:
یعنی این آدرس های حافظه همون رجیستر های مجازی هستن به همین دلیل یکی از اولین کارهای تحلیلگر اینه که محل نگهداری Virtual Register ها رو پیدا کنه
وقتی بفهمید هر Offset مربوط به کدوم Virtual Register هست خوندن Handler ها چند برابر راحت تر میشه
خیلی وقت ها حتی اسمگذاری هم میکنن:
از این لحظه به بعد به جای آدرس های حافظه ذهن تحلیلگر با رجیسترهای مجازی کار میکنه این کار باعث میشه منطق ماشین مجازی کم کم شبیه یک CPU معمولی به نظر برسه
تمرین:
فرض کنید یک ماشین مجازی چهار رجیستر داره:
و Opcode های زیر اجرا میشن:
بدون اینکه کدی بنویسید مرحله به مرحله مشخص کنید بعد از اجرای هر Opcode مقدار هر Virtual Register چقدر میشه
Virtual Register So far we have understood that Dispatcher decides which Handler to execute and Handler does the main work
Where do Handlers store data?
On real CPU registers like RAX and RBX? Usually not
Most virtual machines use something called Virtual Register
Virtual Register
is a memory space that acts as CPU registers
For example, suppose the virtual machine has 8 registers:
When the following Opcode is executed:
10 is placed into V0
Then the next Opcode:
20 is stored into V1
Now the next Opcode:
The virtual machine adds V0 and V1
The result is stored back into V0
If this Opcode is then executed:
30 is printed
Here none of the actual CPU registers are directly visible
You may only see a few instructions inside the Handlers like this:
At first glance, it seems like the program is only working with memory
But in fact:
That is, these memory addresses are the same virtual registers, which is why one of the first tasks of the analyst is to find the location of the Virtual Registers.
When you understand which Virtual Register each Offset belongs to, reading the Handlers becomes much easier.
Many times they even give them names:
From this moment on, instead of memory addresses, the analyst's mind works with virtual registers. This makes the logic of the virtual machine look a little like a regular CPU.
Exercise:
Suppose a virtual machine It has four registers:
And the following Opcodes are executed:
Without writing any code, determine step by step what the value of each Virtual Register will be after executing each Opcode.
@reverseengine
تا اینجا فهمیدیم Dispatcher تصمیم میگیره کدوم Handler اجرا بشه و Handler هم کار اصلی رو انجام میده
Handler
ها داده ها رو کجا نگه میدارن؟
روی رجیستر های واقعی CPU مثل RAX و RBX؟ معمولا نه
بیشتر ماشینهای مجازی از چیزی به اسم Virtual Register استفاده میکنن
Virtual Register
یه فضای حافظه ست که نقش رجیستر های CPU رو بازی میکنه
مثلا فرض کنید ماشین مجازی 8 تا رجیستر داره:
V0 V1 V2 V3 V4 V5 V6 V7
وقتی Opcode زیر اجرا میشه:
LOAD V0, 10
مقدار 10 داخل V0 قرار میگیره
بعد Opcode بعدی:
LOAD V1, 20
عدد 20 داخل V1 ذخیره میشه
حالا Opcode بعدی:
ADD V0, V1
ماشین مجازی مقدار V0 و V1 رو جمع میکنه
نتیجه دوباره داخل V0 ذخیره میشه
اگر بعدش این Opcode اجرا بشه:
PRINT V0
عدد 30 چاپ میشه
اینجا هیچکدوم از رجیسترهای واقعی CPU مستقیما دیده نمیشن
ممکنه داخل Handlerه ا فقط چند دستور مثل این ببینید:
mov rax, [rdi+18h]
mov rcx, [rdi+20h]
add rax, rcx
mov [rdi+18h], raxدر نگاه اول انگار برنامه فقط داره با حافظه کار میکنه
ولی در واقع:
[rdi+18h] = V0 [rdi+20h] = V1 یعنی این آدرس های حافظه همون رجیستر های مجازی هستن به همین دلیل یکی از اولین کارهای تحلیلگر اینه که محل نگهداری Virtual Register ها رو پیدا کنه
وقتی بفهمید هر Offset مربوط به کدوم Virtual Register هست خوندن Handler ها چند برابر راحت تر میشه
خیلی وقت ها حتی اسمگذاری هم میکنن:
[rdi+18h] → V0 [rdi+20h] → V1 [rdi+28h] → V2 از این لحظه به بعد به جای آدرس های حافظه ذهن تحلیلگر با رجیسترهای مجازی کار میکنه این کار باعث میشه منطق ماشین مجازی کم کم شبیه یک CPU معمولی به نظر برسه
تمرین:
فرض کنید یک ماشین مجازی چهار رجیستر داره:
V0 V1 V2 V3
و Opcode های زیر اجرا میشن:
LOAD V0, 15
LOAD V1, 5
SUB V0, V1
PRINT V0بدون اینکه کدی بنویسید مرحله به مرحله مشخص کنید بعد از اجرای هر Opcode مقدار هر Virtual Register چقدر میشه
Virtual Register So far we have understood that Dispatcher decides which Handler to execute and Handler does the main work
Where do Handlers store data?
On real CPU registers like RAX and RBX? Usually not
Most virtual machines use something called Virtual Register
Virtual Register
is a memory space that acts as CPU registers
For example, suppose the virtual machine has 8 registers:
V0 V1 V2 V3 V4 V5 V6 V7
When the following Opcode is executed:
LOAD V0, 10
10 is placed into V0
Then the next Opcode:
LOAD V1, 20
20 is stored into V1
Now the next Opcode:
ADD V0, V1
The virtual machine adds V0 and V1
The result is stored back into V0
If this Opcode is then executed:
PRINT V0
30 is printed
Here none of the actual CPU registers are directly visible
You may only see a few instructions inside the Handlers like this:
mov rax, [rdi+18h]
mov rcx, [rdi+20h]
add rax, rcx
mov [rdi+18h], rax
At first glance, it seems like the program is only working with memory
But in fact:
[rdi+18h] = V0 [rdi+20h] = V1
That is, these memory addresses are the same virtual registers, which is why one of the first tasks of the analyst is to find the location of the Virtual Registers.
When you understand which Virtual Register each Offset belongs to, reading the Handlers becomes much easier.
Many times they even give them names:
[rdi+18h] → V0 [rdi+20h] → V1 [rdi+28h] → V2
From this moment on, instead of memory addresses, the analyst's mind works with virtual registers. This makes the logic of the virtual machine look a little like a regular CPU.
Exercise:
Suppose a virtual machine It has four registers:
V0 V1 V2 V3
And the following Opcodes are executed:
LOAD V0, 15
LOAD V1, 5
SUB V0, V1
PRINT V0
Without writing any code, determine step by step what the value of each Virtual Register will be after executing each Opcode.
@reverseengine
RustyWater ShellCode Dropper has emerged as a key component in recent Static Kitten (MuddyWater) operations targeting organizations in the Gulf and broader Middle East.
Written in Rust and disguised as a legitimate looking reddit.exe, this implant serves as the main payload and backbone of their attacks. It uses a multi-stage dropper (CertificationKit.ini) that decrypts and deploys the payload at runtime, establishes registry persistence, and injects shellcode into explorer.exe for stealth.
What makes it particularly effective is its robust 8-layer anti-analysis system checking for virtual machines, debuggers, sandboxes, low resources, and analysis tools before execution. This ensures it only activates on real victim systems.
A clear example of how Iranian APT groups continue to evolve their tooling with Rust for better evasion and persistence in the region.
Full project details: https://github.com/S3N4T0R-0X0/RustyWater-ShellCode-Dropper
Written in Rust and disguised as a legitimate looking reddit.exe, this implant serves as the main payload and backbone of their attacks. It uses a multi-stage dropper (CertificationKit.ini) that decrypts and deploys the payload at runtime, establishes registry persistence, and injects shellcode into explorer.exe for stealth.
What makes it particularly effective is its robust 8-layer anti-analysis system checking for virtual machines, debuggers, sandboxes, low resources, and analysis tools before execution. This ensures it only activates on real victim systems.
A clear example of how Iranian APT groups continue to evolve their tooling with Rust for better evasion and persistence in the region.
Full project details: https://github.com/S3N4T0R-0X0/RustyWater-ShellCode-Dropper
Human Expertise Still Matters: AI Vulnerability Discovery, Triage, and Patching
https://www.youtube.com/watch?v=ayaINQn7K08
https://www.youtube.com/watch?v=ayaINQn7K08
YouTube
Human Expertise Still Matters AI Vulnerability Discovery Triage and Patching
In this talk we will examine the growing gap between AI-generated security findings and real, exploitable vulnerabilities. As AI tools become better at source code review, web application testing, and vulnerability discovery, they also produce false positives…
This is so overkill
A labeled reverse-engineering dataset over 4,598 crackmes from crackmes.one, built for benchmarking automated RE (AI agents + decompilers)
https://github.com/crackmesone/crackmes-re-dataset
A labeled reverse-engineering dataset over 4,598 crackmes from crackmes.one, built for benchmarking automated RE (AI agents + decompilers)
https://github.com/crackmesone/crackmes-re-dataset
GitHub
GitHub - crackmesone/crackmes-re-dataset: Labeled reverse-engineering dataset over 4,598 crackmes: flags, verifier scripts, and…
Labeled reverse-engineering dataset over 4,598 crackmes: flags, verifier scripts, and normalized obfuscation tags. - crackmesone/crackmes-re-dataset
Forwarded from Source Byte
I made a useful MCP wrapper to help agents analyse multiple binaries in IDA in parallel, and easily switching them, check this out, if you will like it, it would be cool to get feedback and maybe a post in your channel 🙏
https://github.com/whoisqwerz/pocket_disasm
https://github.com/whoisqwerz/pocket_disasm
GitHub
GitHub - whoisqwerz/pocket_disasm: Multi-session IDALib MCP router for coding agents. Analyze multiple binaries in parallel with…
Multi-session IDALib MCP router for coding agents. Analyze multiple binaries in parallel with IDA-compatible reverse engineering tools. - whoisqwerz/pocket_disasm
SSTIC2025_Slides_windows_kernel_shadow_stack_mitigation_aulnette.pdf
2.8 MB
Analyzing the Windows kernel shadow stack mitigation
Jailbreaking Subaru StarLink
https://github.com/sgayou/subaru-starlink-research/blob/master/doc/README.md
https://github.com/sgayou/subaru-starlink-research/blob/master/doc/README.md
GitHub
subaru-starlink-research/doc/README.md at master · sgayou/subaru-starlink-research
Subaru StarLink persistent root code execution. Contribute to sgayou/subaru-starlink-research development by creating an account on GitHub.
IDA 9.4 Release: A New Dyld Shared Cache, Swift Analysis, New Teams add-on, and more
https://hex-rays.com/blog/ida-9.4-release-a-new-dyld-shared-cache-swift-analysis-new-teams-add-on-and-more-
@reverseengine
https://hex-rays.com/blog/ida-9.4-release-a-new-dyld-shared-cache-swift-analysis-new-teams-add-on-and-more-
@reverseengine
Virtual Stack
حافظه ای که ماشین مجازی روی اون کار میکنه
تا اینجا با Dispatcher Handler و Virtual Register آشنا شدیم
اما همه ماشین های مجازی از رجیستر استفاده نمیکنن
بعضیها تقریبا تمام عملیاتشون رو روی یک Stack مجازی انجام میدن
اگر قبلا با اسمبلی کار کرده باشید احتمالا Stack واقعی CPU رو میشناسید
ماشینهای مجازی هم دقیقا همین ایده رو پیاده میکنن با این تفاوت که Stack خودشون رو داخل حافظه میسازن
فرض کنید Stack مجازی در ابتدا خالی باشه
اولین Opcode اجرا میشه:
حالا Stack این شکلیه:
بعد:
Stack:
حالا Opcode بعدی:
Handler
مربوط به ADD دو مقدار بالای Stack رو برمیداره
اونها رو با هم جمع میکنه
یعنی 30 دوباره روی Stack قرار میگیره
Stack
حالا این شکلیه:
بعد:
Handler
مقدار بالای Stack رو میخونه و چاپ میکنه
خیلی از ماشین های مجازی واقعی هم تقریبا با همین منطق کار میکنن
به جای اینکه Opcode ها بنویسن:
مینویسن:
چون طراحی Stack-Based معمولا ساده تره و تولید بایت کد براش راحت تره
وقتی دارید یک VM رو تحلیل میکنید یکی از اولین سوال هایی که باید از خودتون بپرسید اینه:
این VM رجیستر محوره یا Stack-Based؟
جواب این سوال مسیر ادامه تحلیل رو مشخص میکنه
اگر مدام میبینید Handler ها دادهها رو از یک بافر مشخص برمیدارن مقدار جدید داخل همون بافر قرار میدن و یک اشاره گر مدام بالا و پایین میره احتمال زیادی وجود داره که با یک Virtual Stack طرف باشید
در مقابل اگر بیشتر عملیات روی چند خونه ثابت حافظه انجام میشه احتمالا VM از Virtual Register استفاده میکنه
یکی از مهارتهای مهم در Devirtualization اینه که خیلی زود تشخیص بدید معماری ماشین مجازی از کدوم نوعه
این کار باعث میشه ساعت ها وقتتون صرف تحلیل اشتباه نشه
تمرین:
فرض کنید Stack مجازی در ابتدا خالیه و این Opcode ها اجرا میشن:
Virtual Stack The memory that the virtual machine works on
So far we have been introduced to the Dispatcher Handler and Virtual Register
But not all virtual machines use registers
Some perform almost all their operations on a virtual stack
If you have worked with assembly before, you probably know the real CPU stack
Virtual machines implement exactly the same idea, except that they create their own stack in memory
Assume that the virtual stack is initially empty
The first Opcode is executed:
Now the Stack looks like this:
Next:
Stack:
Now the next Opcode:
The Handler
removes the top two values of the Stack
Adds them together
That is, 30 is placed back on the Stack
Stack
Now this Figure:
Next:
Handler
Reads and prints the top of the stack
Many real virtual machines work with almost the same logic
Instead of writing Opcodes:
They write:
Because Stack-Based design is usually simpler and bytecode generation is easier
When you are analyzing a VM, one of the first questions you should ask yourself is:
Is this VM register-based or stack-based?
The answer to this question will determine the path of further analysis
If you constantly see Handlers taking data from a specific buffer, putting a new value into the same buffer, and a pointer constantly moving up and down, there is a high probability that you are dealing with a Virtual Stack
On the other hand, if most of the operations are performed on a few fixed memory locations, the VM is probably using Virtual Registers
One of the important skills in Devirtualization is to quickly identify what type of virtual machine architecture you have
This will save you hours of time on incorrect analysis
Exercise:
Assume that the virtual stack is initially empty and these opcodes are executed:
Without running the program, write down the state of the stack after each opcode
@reverseengine
حافظه ای که ماشین مجازی روی اون کار میکنه
تا اینجا با Dispatcher Handler و Virtual Register آشنا شدیم
اما همه ماشین های مجازی از رجیستر استفاده نمیکنن
بعضیها تقریبا تمام عملیاتشون رو روی یک Stack مجازی انجام میدن
اگر قبلا با اسمبلی کار کرده باشید احتمالا Stack واقعی CPU رو میشناسید
ماشینهای مجازی هم دقیقا همین ایده رو پیاده میکنن با این تفاوت که Stack خودشون رو داخل حافظه میسازن
فرض کنید Stack مجازی در ابتدا خالی باشه
اولین Opcode اجرا میشه:
PUSH 10
حالا Stack این شکلیه:
10
بعد:
PUSH 20
Stack:
20
10
حالا Opcode بعدی:
ADD
Handler
مربوط به ADD دو مقدار بالای Stack رو برمیداره
20
10
اونها رو با هم جمع میکنه
یعنی 30 دوباره روی Stack قرار میگیره
Stack
حالا این شکلیه:
30
بعد:
Handler
مقدار بالای Stack رو میخونه و چاپ میکنه
خیلی از ماشین های مجازی واقعی هم تقریبا با همین منطق کار میکنن
به جای اینکه Opcode ها بنویسن:
ADD V0, V1
مینویسن:
PUSH V0
PUSH V1
ADD
چون طراحی Stack-Based معمولا ساده تره و تولید بایت کد براش راحت تره
وقتی دارید یک VM رو تحلیل میکنید یکی از اولین سوال هایی که باید از خودتون بپرسید اینه:
این VM رجیستر محوره یا Stack-Based؟
جواب این سوال مسیر ادامه تحلیل رو مشخص میکنه
اگر مدام میبینید Handler ها دادهها رو از یک بافر مشخص برمیدارن مقدار جدید داخل همون بافر قرار میدن و یک اشاره گر مدام بالا و پایین میره احتمال زیادی وجود داره که با یک Virtual Stack طرف باشید
در مقابل اگر بیشتر عملیات روی چند خونه ثابت حافظه انجام میشه احتمالا VM از Virtual Register استفاده میکنه
یکی از مهارتهای مهم در Devirtualization اینه که خیلی زود تشخیص بدید معماری ماشین مجازی از کدوم نوعه
این کار باعث میشه ساعت ها وقتتون صرف تحلیل اشتباه نشه
تمرین:
فرض کنید Stack مجازی در ابتدا خالیه و این Opcode ها اجرا میشن:
PUSH 8بدون اجرای برنامه روی کاغذ وضعیت Stack رو بعد از هر Opcode بنویسید
PUSH 12
ADD
PUSH 3
MUL
Virtual Stack The memory that the virtual machine works on
So far we have been introduced to the Dispatcher Handler and Virtual Register
But not all virtual machines use registers
Some perform almost all their operations on a virtual stack
If you have worked with assembly before, you probably know the real CPU stack
Virtual machines implement exactly the same idea, except that they create their own stack in memory
Assume that the virtual stack is initially empty
The first Opcode is executed:
PUSH 10
Now the Stack looks like this:
10
Next:
PUSH 20
Stack:
20
10
Now the next Opcode:
ADD
The Handler
removes the top two values of the Stack
20
10
Adds them together
That is, 30 is placed back on the Stack
Stack
Now this Figure:
30
Next:
Handler
Reads and prints the top of the stack
Many real virtual machines work with almost the same logic
Instead of writing Opcodes:
ADD V0, V1
They write:
PUSH V0
PUSH V1
ADD
Because Stack-Based design is usually simpler and bytecode generation is easier
When you are analyzing a VM, one of the first questions you should ask yourself is:
Is this VM register-based or stack-based?
The answer to this question will determine the path of further analysis
If you constantly see Handlers taking data from a specific buffer, putting a new value into the same buffer, and a pointer constantly moving up and down, there is a high probability that you are dealing with a Virtual Stack
On the other hand, if most of the operations are performed on a few fixed memory locations, the VM is probably using Virtual Registers
One of the important skills in Devirtualization is to quickly identify what type of virtual machine architecture you have
This will save you hours of time on incorrect analysis
Exercise:
Assume that the virtual stack is initially empty and these opcodes are executed:
PUSH 8
PUSH 12
ADD
PUSH 3
MUL
Without running the program, write down the state of the stack after each opcode
@reverseengine
یکی از مهمترین بخشهای Binary Exploitation
Heap هست
تا الان بیشتر دربارهی Stack و کنترل جریان اجرا صحبت کردیم اما خیلی از باگهای جدی امروزی داخل Heap اتفاق میوفتن
Heap
چیه و چرا در Binary Exploitation مهمه؟
وقتی یک برنامه اجرا میشه حافظه ی اون به چند بخش مختلف تقسیم میشه که دو بخش هستن:
Stack
بیشتر برای اطلاعات موقت توابع استفاده میشه اما Heap برای زمانیه که برنامه در زمان اجرا نیاز دارد حافظه ای رو خودش مدیریت کنه
Heap چیه؟
Heap
یک بخش از حافظه است که برنامه میتونه در زمان اجرا از اون درخواست حافظه کنه و بعدا اونو ازاد کنه
مثلا وقتی برنامه نمیدونه چقدر داده قراره دریافت کنه نمیتونه از یک فضای ثابت روی Stack استفاده کنه در این حالت معمولا سراغ Heap میروه
در زبانهایی مثل C و C++ مدیریت Heap معمولا با توابعی مثل:
انجام میشه
تفاوت Stack و Heap
Stack:
ساختار منظم تر و سریع تر داره
عمر دادهها معمولا وابسته به تابع هست
مدیریت اون بیشتر توسط کامپایلر انجام میشه
مثلا:
وقتی تابع تموم بشه این داده از بین میره
Heap:
توسط خود برنامه مدیریت میشه
دادهها میتونن مدت بیشتری باقی بمونن
برنامه خودش تصمیم میگیره چه زمانی حافظه بگیره و چه زمانی ازاد کنه
مثلا:
اینجا برنامه درخواست میکنه 100 بایت حافظه از Heap بگیره
چرا Heap در امنیت مهمه؟
چون مدیریت دستی حافظه مخصوصا در زبانهایی مثل C و C++ پیچیدگی زیادی داره
برنامهنویس باید همیشه مراقب باشه:
چه زمانی حافظه گرفته شده؟
چه زمانی آزاد شده؟
آیا هنوز از حافظه آزاد شده استفاده میشه؟
آیا اندازهی داده با اندازهی حافظه هماهنگه؟
یک اشتباه کوچیک میتونه باعث ایجاد آسیب پذیری بشه
چه نوع باگهایی در Heap دیده میشن؟
چند مورد معروف:
وقتی برنامه بیشتر از اندازهی اختصاص داده شده روی Heap داده مینویسه
وقتی برنامه حافظه ای رو آزاد میکنه اما بعدا همچنان از اون استفاده میکنه
وقتی برنامه یک بخش از حافظه رو بیشتر از یک بار آزاد میکنه
وقتی برنامه حافظه میگیره ولی هیچ وقت آزادش نمیکنه
چرا Heap سختتر از Stack است؟
Stack
ساختار نسبتا مشخصی دارده اما Heap پیچیده تره
چون Heap توسط یک Memory Allocator مدیریت میشه
مثلا در لینوکس یکی از allocator های معروف:
ptmalloc (در glibc)
هست
این allocator تصمیم میگیره:
چه بخشی از حافظه اختصاص داده بشه
کدام حافظه آزاد باشه
درخواست های جدید کجا قرار بگیرن
به همین دلیل تحلیل Heap نیاز به درک عمیق تری از مدیریت حافظه داره
چرا هکر ها یا مهندسان معکوس به Heap علاقه دارن؟
چون Heap معمولا شامل داده های مهم برنامه هست
مثلا:
ساختار های داده
Object
ها در ++C
اطلاعات session
pointer ها
داده های برنامه
اگر مدیریت Heap اشتباه باشه ممکنه باعث تغییر رفتار برنامه بشه
Heap
یکی از مهمترین بخش های حافظه در برنامه هاست که برای ذخیره سازی داده های پویا استفاده میشه
برخلاف Stack مدیریت Heap بیشتر بر عهده ی برنامه نویسه و همین موضوع باعث بهوجود اومدن باگهای پیچیدهای مثل Heap Overflow و Use-After-Free میشه
برای درک اکسپلویت های مدرن شناخت Heap و نحوهی کار Memory Allocator ها ضروریه
@reverseengine
Heap هست
تا الان بیشتر دربارهی Stack و کنترل جریان اجرا صحبت کردیم اما خیلی از باگهای جدی امروزی داخل Heap اتفاق میوفتن
Heap
چیه و چرا در Binary Exploitation مهمه؟
وقتی یک برنامه اجرا میشه حافظه ی اون به چند بخش مختلف تقسیم میشه که دو بخش هستن:
Stack
Heap
Stack
بیشتر برای اطلاعات موقت توابع استفاده میشه اما Heap برای زمانیه که برنامه در زمان اجرا نیاز دارد حافظه ای رو خودش مدیریت کنه
Heap چیه؟
Heap
یک بخش از حافظه است که برنامه میتونه در زمان اجرا از اون درخواست حافظه کنه و بعدا اونو ازاد کنه
مثلا وقتی برنامه نمیدونه چقدر داده قراره دریافت کنه نمیتونه از یک فضای ثابت روی Stack استفاده کنه در این حالت معمولا سراغ Heap میروه
در زبانهایی مثل C و C++ مدیریت Heap معمولا با توابعی مثل:
malloc()
calloc()
realloc()
free()
انجام میشه
تفاوت Stack و Heap
Stack:
ساختار منظم تر و سریع تر داره
عمر دادهها معمولا وابسته به تابع هست
مدیریت اون بیشتر توسط کامپایلر انجام میشه
مثلا:
void function(){
char buffer[64];
}
وقتی تابع تموم بشه این داده از بین میره
Heap:
توسط خود برنامه مدیریت میشه
دادهها میتونن مدت بیشتری باقی بمونن
برنامه خودش تصمیم میگیره چه زمانی حافظه بگیره و چه زمانی ازاد کنه
مثلا:
char *data = malloc(100);
اینجا برنامه درخواست میکنه 100 بایت حافظه از Heap بگیره
چرا Heap در امنیت مهمه؟
چون مدیریت دستی حافظه مخصوصا در زبانهایی مثل C و C++ پیچیدگی زیادی داره
برنامهنویس باید همیشه مراقب باشه:
چه زمانی حافظه گرفته شده؟
چه زمانی آزاد شده؟
آیا هنوز از حافظه آزاد شده استفاده میشه؟
آیا اندازهی داده با اندازهی حافظه هماهنگه؟
یک اشتباه کوچیک میتونه باعث ایجاد آسیب پذیری بشه
چه نوع باگهایی در Heap دیده میشن؟
چند مورد معروف:
Heap Overflow
وقتی برنامه بیشتر از اندازهی اختصاص داده شده روی Heap داده مینویسه
Use-After-Free (UAF)
وقتی برنامه حافظه ای رو آزاد میکنه اما بعدا همچنان از اون استفاده میکنه
Double Free
وقتی برنامه یک بخش از حافظه رو بیشتر از یک بار آزاد میکنه
Memory Leak
وقتی برنامه حافظه میگیره ولی هیچ وقت آزادش نمیکنه
چرا Heap سختتر از Stack است؟
Stack
ساختار نسبتا مشخصی دارده اما Heap پیچیده تره
چون Heap توسط یک Memory Allocator مدیریت میشه
مثلا در لینوکس یکی از allocator های معروف:
ptmalloc (در glibc)
هست
این allocator تصمیم میگیره:
چه بخشی از حافظه اختصاص داده بشه
کدام حافظه آزاد باشه
درخواست های جدید کجا قرار بگیرن
به همین دلیل تحلیل Heap نیاز به درک عمیق تری از مدیریت حافظه داره
چرا هکر ها یا مهندسان معکوس به Heap علاقه دارن؟
چون Heap معمولا شامل داده های مهم برنامه هست
مثلا:
ساختار های داده
Object
ها در ++C
اطلاعات session
pointer ها
داده های برنامه
اگر مدیریت Heap اشتباه باشه ممکنه باعث تغییر رفتار برنامه بشه
Heap
یکی از مهمترین بخش های حافظه در برنامه هاست که برای ذخیره سازی داده های پویا استفاده میشه
برخلاف Stack مدیریت Heap بیشتر بر عهده ی برنامه نویسه و همین موضوع باعث بهوجود اومدن باگهای پیچیدهای مثل Heap Overflow و Use-After-Free میشه
برای درک اکسپلویت های مدرن شناخت Heap و نحوهی کار Memory Allocator ها ضروریه
@reverseengine
ReverseEngineering
یکی از مهمترین بخشهای Binary Exploitation Heap هست تا الان بیشتر دربارهی Stack و کنترل جریان اجرا صحبت کردیم اما خیلی از باگهای جدی امروزی داخل Heap اتفاق میوفتن Heap چیه و چرا در Binary Exploitation مهمه؟ وقتی یک برنامه اجرا میشه حافظه ی اون…
One of the most important parts of Binary Exploitation Is the Heap
So far we have talked mostly about the Stack and execution flow control, but many of today's serious bugs occur in the Heap
What is the Heap and why is it important in Binary Exploitation?
When a program is executed, its memory is divided into several different parts, which are two parts:
Stack is mostly used for temporary information of functions, but the Heap is for when the program needs to manage memory itself at runtime
What is the Heap?
Heap is a part of memory that a program can request memory from at runtime and free it later
For example, when a program does not know how much data it is going to receive, it cannot use a fixed space on the Stack, in this case it usually goes to the Heap
In languages like C and C++, Heap management is usually done with functions like:
Stack:
It has a more organized and faster structure
The lifetime of the data is usually dependent on the function
Its management is mostly done by the compiler
For example:
When the function ends, this data is destroyed
Heap:
Managed by the program itself
Data can remain for a longer period
The program itself decides when to get memory and when to free it
For example:
Here the program requests to get 100 bytes of memory from the Heap
Why is the Heap important in security?
Because manual memory management is very complicated, especially in languages like C and C++
The programmer should always be careful:
When was the memory taken?
When was it freed?
Is the freed memory still in use?
Is the data size consistent with the memory size?
A small mistake can cause a vulnerability
What types of bugs are seen in the Heap?
Some famous cases:
When the program writes more data to the Heap than the allocated size
When the program frees memory but then continues to use it
When the program frees a section of memory more than once
When the program takes memory but never frees it
Why is the Heap harder than the Stack?
has a relatively well-defined structure, but the Heap is more complex
Because the Heap is managed by a Memory Allocator
For example, in Linux, one of the famous allocators is:
ptmalloc (in glibc)
This allocator decides:
What part of the memory to allocate
Which memory to free
Where new requests should be placed
That is why Heap analysis requires a deeper understanding of memory management
Why are hackers or reverse engineers interested in the Heap?
Because the Heap usually contains important program data
For example:
If the Heap management is wrong, it may change the behavior of the program.
The Heap is one of the most important memory areas in programs, used to store dynamic data.
Unlike the Stack, the Heap management is more the responsibility of the programmer, which leads to complex bugs such as Heap Overflow and Use-After-Free.
To understand modern exploits, it is essential to understand the Heap and how Memory Allocators work.
@reverseengine
So far we have talked mostly about the Stack and execution flow control, but many of today's serious bugs occur in the Heap
What is the Heap and why is it important in Binary Exploitation?
When a program is executed, its memory is divided into several different parts, which are two parts:
Stack
Heap
Stack is mostly used for temporary information of functions, but the Heap is for when the program needs to manage memory itself at runtime
What is the Heap?
Heap is a part of memory that a program can request memory from at runtime and free it later
For example, when a program does not know how much data it is going to receive, it cannot use a fixed space on the Stack, in this case it usually goes to the Heap
In languages like C and C++, Heap management is usually done with functions like:
malloc()The difference between Stack and Heap
calloc()
realloc()
free()
Stack:
It has a more organized and faster structure
The lifetime of the data is usually dependent on the function
Its management is mostly done by the compiler
For example:
void function(){
char buffer[64];
}
When the function ends, this data is destroyed
Heap:
Managed by the program itself
Data can remain for a longer period
The program itself decides when to get memory and when to free it
For example:
char *data = malloc(100);
Here the program requests to get 100 bytes of memory from the Heap
Why is the Heap important in security?
Because manual memory management is very complicated, especially in languages like C and C++
The programmer should always be careful:
When was the memory taken?
When was it freed?
Is the freed memory still in use?
Is the data size consistent with the memory size?
A small mistake can cause a vulnerability
What types of bugs are seen in the Heap?
Some famous cases:
Heap Overflow
When the program writes more data to the Heap than the allocated size
Use-After-Free (UAF)
When the program frees memory but then continues to use it
Double Free
When the program frees a section of memory more than once
Memory Leak
When the program takes memory but never frees it
Why is the Heap harder than the Stack?
Stack
has a relatively well-defined structure, but the Heap is more complex
Because the Heap is managed by a Memory Allocator
For example, in Linux, one of the famous allocators is:
ptmalloc (in glibc)
This allocator decides:
What part of the memory to allocate
Which memory to free
Where new requests should be placed
That is why Heap analysis requires a deeper understanding of memory management
Why are hackers or reverse engineers interested in the Heap?
Because the Heap usually contains important program data
For example:
Data structures
Objects in C++
Session information
Pointers
Program data
If the Heap management is wrong, it may change the behavior of the program.
The Heap is one of the most important memory areas in programs, used to store dynamic data.
Unlike the Stack, the Heap management is more the responsibility of the programmer, which leads to complex bugs such as Heap Overflow and Use-After-Free.
To understand modern exploits, it is essential to understand the Heap and how Memory Allocators work.
@reverseengine