Fork
چطور یه Process جدید ساخته میشه؟
تا اینجا درباره Process API حرف زدیم و گفتیم برنامه ها چطوری میتونن با Process ها کار کنن
یکی از مهم ترین چیز هایی که اینجا باهاش سروکار داریم
ولی اصلا fork() چیکار میکنه؟
خیلی ساده بخوایم بگیم:
fork()
از Process فعلی یه Process جدید به اسم Child میسازه
یعنی قبل از fork() فقط یه Process داریم:
Parent
بعد از fork():
Parent
+
Child
حالا دو تا Process داریم که هر دو از همون جایی که fork() اجرا شده به اجرای برنامه ادامه میدن
اما یه نکته مهم این وسط هست
Child
صرفا یه کپی ساده از فایل برنامه نیست
سیستم عامل برای Child یه Process مستقل ایجاد میکنه
Child
معمولا:
فضای آدرس مجازی خودش رو داره
PID
متفاوتی داره
وضعیت اجرای خودش رو داره
و منابع قابل مدیریت خودش رو داره
ولی در لحظهای که ساخته میشه وضعیت حافظه و اجرای اون خیلی شبیه Parent هست
پس چرا میگیم Child شبیه کپی Parent هست
چون وقتی Child ساخته میشه سیستمعامل کاری میکنه که Child تقریبا همون وضعیت Parent رو داشته باشه
مثلا فرض کنید Parent قبل از fork() این متغیر رو داشته باشه:
int x = 10;
Child
هم در فضای آدرس خودش مقدار مشابه ای برای
ولی این به این معنی نیست که Parent و Child دارن از یه متغیر معمولی مشترک استفاده میکنن
هر Process فضای حافظه مجازی خودش رو داره
یه ویژگی خیلی جالب fork()
توی Parent و Child مقدار برگشتی یکسانی نداره
توی Parent مقدار برگشتی معمولا PID مربوط به Child
توی Child مقدار برگشتی
اگر هم مشکلی موقع ساخت Child پیش بیاد، Parent یه مقدار منفی دریافت میکنه
پس برنامه میتونه بفهمه:
من Parent ام؟
یا Child؟
و بر اساس اون مسیر متفاوتی رو اجرا کنه
مثلا:
pid = fork();
if (pid == 0)
// Child
else
// Parent
بعد از اجرای fork() هر دو Process از همون نقطه به اجرای برنامه ادامه میدن
ولی مقدار
حالا یه سوال مهمتر:
اگه Child تقریبا شبیه Parent ساخته شده چطوری میتونه یه برنامه کاملا
متفاوت رو اجرا کنه؟
اینجاست که
معمولا توی سیستمهای Unix یه الگوی خیلی معروف داریم:
Parent
↓
fork()
↓
Child
↓
exec()
↓
New Program
یعنی:
fork() → ساختن Child
و بعد:
exec() → جایگزین کردن برنامه داخل
Child با برنامهای که میخوایم اجرا بشه
مثلا یه Shell رو در نظر بگیرید
Shell
میتونه یه Child بسازه و بعد Child رو با
حالا این موضوع چرا برای مهندسی معکوس مهمه؟
وقتی توی یه برنامه لینوکسی به
یه Process جدید وارد ماجرا شده
مثلا:
Parent
│
├── fork()
│
├────────► Child
│
▼
ادامه Parent
یعنی ممکنه رفتار برنامه بین Parent و Child تقسیم شده باشه پس موقع تحلیل یه برنامه اگه به
این موضوع بعدا توی تحلیل Process ها Debugging و بررسی رفتار برنامهها خیلی به دردتون میخوره
یه Child Process جدید میسازه
Parent
و Child هر دو به اجرای برنامه ادامه میدن
Child
یه PID متفاوت داره
هر Process فضای آدرس مجازی خودش رو داره مقدار برگشتی
و اینجا یکی از ایدههای مهم رو خیلی خوب میشه دید:
سیستمعامل فقط برنامهها رو اجرا نمیکنه بلکه یه محیط میسازه که چند تا Process بتونن مستقل از هم اجرا بشن
@reverseengine
چطور یه Process جدید ساخته میشه؟
تا اینجا درباره Process API حرف زدیم و گفتیم برنامه ها چطوری میتونن با Process ها کار کنن
یکی از مهم ترین چیز هایی که اینجا باهاش سروکار داریم
fork() هستولی اصلا fork() چیکار میکنه؟
خیلی ساده بخوایم بگیم:
fork()
از Process فعلی یه Process جدید به اسم Child میسازه
یعنی قبل از fork() فقط یه Process داریم:
Parent
بعد از fork():
Parent
+
Child
حالا دو تا Process داریم که هر دو از همون جایی که fork() اجرا شده به اجرای برنامه ادامه میدن
اما یه نکته مهم این وسط هست
Child
صرفا یه کپی ساده از فایل برنامه نیست
سیستم عامل برای Child یه Process مستقل ایجاد میکنه
Child
معمولا:
فضای آدرس مجازی خودش رو داره
PID
متفاوتی داره
وضعیت اجرای خودش رو داره
و منابع قابل مدیریت خودش رو داره
ولی در لحظهای که ساخته میشه وضعیت حافظه و اجرای اون خیلی شبیه Parent هست
پس چرا میگیم Child شبیه کپی Parent هست
چون وقتی Child ساخته میشه سیستمعامل کاری میکنه که Child تقریبا همون وضعیت Parent رو داشته باشه
مثلا فرض کنید Parent قبل از fork() این متغیر رو داشته باشه:
int x = 10;
Child
هم در فضای آدرس خودش مقدار مشابه ای برای
x دارهولی این به این معنی نیست که Parent و Child دارن از یه متغیر معمولی مشترک استفاده میکنن
هر Process فضای حافظه مجازی خودش رو داره
یه ویژگی خیلی جالب fork()
fork()توی Parent و Child مقدار برگشتی یکسانی نداره
توی Parent مقدار برگشتی معمولا PID مربوط به Child
توی Child مقدار برگشتی
0 هستاگر هم مشکلی موقع ساخت Child پیش بیاد، Parent یه مقدار منفی دریافت میکنه
پس برنامه میتونه بفهمه:
من Parent ام؟
یا Child؟
و بر اساس اون مسیر متفاوتی رو اجرا کنه
مثلا:
pid = fork();
if (pid == 0)
// Child
else
// Parent
بعد از اجرای fork() هر دو Process از همون نقطه به اجرای برنامه ادامه میدن
ولی مقدار
pid بهشون میگه که الان داخل Parent هستن یا Childحالا یه سوال مهمتر:
اگه Child تقریبا شبیه Parent ساخته شده چطوری میتونه یه برنامه کاملا
متفاوت رو اجرا کنه؟
اینجاست که
exec() وارد داستان میشهمعمولا توی سیستمهای Unix یه الگوی خیلی معروف داریم:
Parent
↓
fork()
↓
Child
↓
exec()
↓
New Program
یعنی:
fork() → ساختن Child
و بعد:
exec() → جایگزین کردن برنامه داخل
Child با برنامهای که میخوایم اجرا بشه
مثلا یه Shell رو در نظر بگیرید
Shell
میتونه یه Child بسازه و بعد Child رو با
exec() تبدیل کنه به برنامهای که کاربر درخواست کردهحالا این موضوع چرا برای مهندسی معکوس مهمه؟
وقتی توی یه برنامه لینوکسی به
fork() برخورد میکنید باید حواستون باشه که از اینجا به بعد دیگه فقط با یه مسیر اجرای برنامه طرف نیستیدیه Process جدید وارد ماجرا شده
مثلا:
Parent
│
├── fork()
│
├────────► Child
│
▼
ادامه Parent
یعنی ممکنه رفتار برنامه بین Parent و Child تقسیم شده باشه پس موقع تحلیل یه برنامه اگه به
fork() رسیدید حتما باید این احتمال رو در نظر بگیرید که از این نقطه به بعد دو تا Process جداگانه دارید که ممکنه هرکدوم رفتار متفاوتی داشته باشناین موضوع بعدا توی تحلیل Process ها Debugging و بررسی رفتار برنامهها خیلی به دردتون میخوره
fork()یه Child Process جدید میسازه
Parent
و Child هر دو به اجرای برنامه ادامه میدن
Child
یه PID متفاوت داره
هر Process فضای آدرس مجازی خودش رو داره مقدار برگشتی
fork() کمک میکنه بفهمیم داخل Parent هستیم یا Child و معمولا fork() رو در کنار exec() میبینیمو اینجا یکی از ایدههای مهم رو خیلی خوب میشه دید:
سیستمعامل فقط برنامهها رو اجرا نمیکنه بلکه یه محیط میسازه که چند تا Process بتونن مستقل از هم اجرا بشن
@reverseengine
ReverseEngineering
Fork چطور یه Process جدید ساخته میشه؟ تا اینجا درباره Process API حرف زدیم و گفتیم برنامه ها چطوری میتونن با Process ها کار کنن یکی از مهم ترین چیز هایی که اینجا باهاش سروکار داریم fork() هست ولی اصلا fork() چیکار میکنه؟ خیلی ساده بخوایم بگیم: fork()…
Fork
How do you create a new Process?
So far we've talked about the Process API and how programs can work with Processes
One of the most important things we're dealing with here is fork()
But what does fork() do?
To put it very simply:
fork()
creates a new Process called Child from the current Process
That is, before fork() we only have one Process:
Parent
After fork():
Parent
+
Child
Now we have two Processes, both of which continue to execute the program from the same place where fork() was executed
But there is an important point here
Child
is not just a simple copy of the program file
The operating system creates an independent Process for Child
Child
usually:
It has its own virtual address space
It has a different PID
It has its own execution state
And it has its own manageable resources
But at the moment it is created, its memory and execution state are very similar to Parent
So why do we say that Child is like a copy of Parent
Because when Child is created, the operating system makes Child have almost the same state as Parent
For example, suppose Parent has this variable before fork():
int x = 10;
Child
also has the same value for x in its address space
But this does not mean that Parent and Child are using a common common variable
Each Process has its own virtual memory space
A very interesting feature of fork()
fork()
does not have the same return value in Parent and Child
In Parent the return value is usually the PID of Child
In Child the return value is 0
If there is a problem while creating Child, Parent gets a negative value
So the program can understand:
Am I Parent?
Or Child?
And based on that execute a different path
For example:
pid = fork();
if (pid == 0) // Child
else // Parent
After executing fork() both Processes continue executing the program from the same point
But the pid value tells them whether they are now inside Parent or Child
Now a more important question:
If Child is created almost exactly like Parent, how can it execute a completely different program?
This is where exec() comes into play.
Usually, in Unix systems, we have a very famous pattern:
Parent
↓
fork()
↓
Child
↓
exec()
↓
New Program
That is:
fork() → create Child
And then:
exec() → replace the program inside
Child with the program we want to run
For example, consider a Shell
The Shell
can create a Child and then use exec() to convert the Child into the program requested by the user
Now why is this important for reverse engineering?
When you encounter fork() in a Linux program, you should be aware that from here on, you are no longer dealing with just one path of program execution
A new process has entered the story
For example:
Parent
│
├── fork()
│
├───────► Child
│
▼
Continuation of Parent
This means that the behavior of the program may be divided between Parent and Child
So when analyzing a program, if you reach fork(), you must definitely consider the possibility that from this point on, you have two separate processes, each of which may have different behavior. This will be very useful later in analyzing processes, debugging, and examining program behavior
fork()
creates a new Child Process
Parent
and Child
both continue executing the program
Child
has a different PID
Each Process has its own virtual address space
The return value of fork() helps us know whether we are inside Parent or Child
And we usually see fork() next to exec()
And here one of the important ideas is very important It's easy to see:
The operating system doesn't just run programs, it creates an environment where multiple processes can run independently
@reverseengine
How do you create a new Process?
So far we've talked about the Process API and how programs can work with Processes
One of the most important things we're dealing with here is fork()
But what does fork() do?
To put it very simply:
fork()
creates a new Process called Child from the current Process
That is, before fork() we only have one Process:
Parent
After fork():
Parent
+
Child
Now we have two Processes, both of which continue to execute the program from the same place where fork() was executed
But there is an important point here
Child
is not just a simple copy of the program file
The operating system creates an independent Process for Child
Child
usually:
It has its own virtual address space
It has a different PID
It has its own execution state
And it has its own manageable resources
But at the moment it is created, its memory and execution state are very similar to Parent
So why do we say that Child is like a copy of Parent
Because when Child is created, the operating system makes Child have almost the same state as Parent
For example, suppose Parent has this variable before fork():
int x = 10;
Child
also has the same value for x in its address space
But this does not mean that Parent and Child are using a common common variable
Each Process has its own virtual memory space
A very interesting feature of fork()
fork()
does not have the same return value in Parent and Child
In Parent the return value is usually the PID of Child
In Child the return value is 0
If there is a problem while creating Child, Parent gets a negative value
So the program can understand:
Am I Parent?
Or Child?
And based on that execute a different path
For example:
pid = fork();
if (pid == 0) // Child
else // Parent
After executing fork() both Processes continue executing the program from the same point
But the pid value tells them whether they are now inside Parent or Child
Now a more important question:
If Child is created almost exactly like Parent, how can it execute a completely different program?
This is where exec() comes into play.
Usually, in Unix systems, we have a very famous pattern:
Parent
↓
fork()
↓
Child
↓
exec()
↓
New Program
That is:
fork() → create Child
And then:
exec() → replace the program inside
Child with the program we want to run
For example, consider a Shell
The Shell
can create a Child and then use exec() to convert the Child into the program requested by the user
Now why is this important for reverse engineering?
When you encounter fork() in a Linux program, you should be aware that from here on, you are no longer dealing with just one path of program execution
A new process has entered the story
For example:
Parent
│
├── fork()
│
├───────► Child
│
▼
Continuation of Parent
This means that the behavior of the program may be divided between Parent and Child
So when analyzing a program, if you reach fork(), you must definitely consider the possibility that from this point on, you have two separate processes, each of which may have different behavior. This will be very useful later in analyzing processes, debugging, and examining program behavior
fork()
creates a new Child Process
Parent
and Child
both continue executing the program
Child
has a different PID
Each Process has its own virtual address space
The return value of fork() helps us know whether we are inside Parent or Child
And we usually see fork() next to exec()
And here one of the important ideas is very important It's easy to see:
The operating system doesn't just run programs, it creates an environment where multiple processes can run independently
@reverseengine
User Mode
در برابر Kernel Mode
برای اینکه بفهمید چرا تکنیک هایی مثل Direct Syscall اصلا تعریف شدن اول باید بدونید ویندوز دو تعریف مهم داره:
┌─────────────────────────┐
│ User Mode │
│ Applications / DLLs │
└───────────┬─────────────┘
│
System Call
│
▼
┌─────────────────────────┐
│ Kernel Mode │
│ Windows Kernel / Drivers│
└─────────────────────────┘
User Mode
برنامههای معمولی اینجا اجرا میشن:
Browser
PowerShell
Game
Your Program
↓
kernel32.dll
↓
ntdll.dll
↓
System Call
EDR
ها میتونن در این لایه telemetry جمع آوری کنن و رفتار برنامه رو بررسی کنن
Kernel Mode
اینجا بخش های حساس سیستم عامل قرار دارن
مثلا مدیریت:
Processes
Threads
Memory
Drivers
File System
Network
به همین دلیل اگر فقط یک لایه از User Mode رو دور بزنید به معنی نامرئی شدن نیست.
-
Direct Syscall
چه مفهومی داره؟
ایده اصلی اینه که بهجای طی کردن مسیر معمول User-Mode API برنامه مستقیما به مرز System Call نزدیک بشید
مفهوم:
Normal:
Application
↓
Win32 API
↓
ntdll
↓
System Call
↓
Kernel
Direct Syscall:
Application
↓
System Call
↓
Kernel
اما نکته مهم همینجاست:
Direct Syscall
به معنی دور زدن کامل EDR نیست
چون Kernel و سایر منابع telemetry همچنان میتونن رفتار اتفاق افتاده رو ببینن
پس چرا هکر ها بهش علاقه دارن؟
چون اگر یک مکانیزم دفاعی مشخص در User Mode قرار گرفته باشه تغییر مسیر اجرای برنامه میتونه روی همان مکانیزم اثر بذاره
ولی EDR مدرن فقط به یک نقطه وابسته نیست:
┌── User Mode
│
Process ─────┼── Memory
│
├── ETW / Telemetry
│
├── Kernel
│
└── Network
بنابراین Evasion واقعی یک بازی چند لایه ست نه پیدا کردن یک API عجیب و غریب
نکتهای که باید یاد بگیرید
اگر بخواید AV/EDR Evasion رو واقعا بفهمید نباید از حفظ کردن تکنیکها شروع کنید
باید بفهمید:
دفاع کجاست
چه چیزی رو میتونه ببینه
چه telemetry دریافت میکنه
مهاجم تلاش میکنه کدوم visibility رو کاهش بده
User Mode vs Kernel Mode
To understand why techniques like Direct Syscall were defined at all, you first need to know that Windows has two important definitions:
┌───────────────────────────┐
│ User Mode │
│ Applications / DLLs │
└─
│
System Call
│
▼
┌────────────� Management:
Processes
Threads
Memory
Drivers
File System
Network
That's why bypassing just one layer of User Mode doesn't mean you'll be invisible.
Direct Syscall
What does it mean?
The main idea is to approach the System Call boundary directly instead of going through the usual User-Mode API path.
Conceptual:
Normal:
Application
↓
Win32 API
↓
ntdll
↓
System Call
↓
Kernel
Direct Syscall:
Application
↓
System Call
↓
Kernel
But here's the important point:
Direct Syscall
does not mean bypassing EDR completely
Because the Kernel and other telemetry sources can still see the behavior that happened
So why are hackers interested in it?
Because if a specific defense mechanism is in User Mode, rerouting the program can affect that mechanism
But modern EDR is not just about one point:
┌── User Mode
│
Process ──────┼── Memory
│
├── ETW / Telemetry
│
├── Kernel
│
└── Network
So real evasion is a multi-layered game, not about finding a weird API
What you need to learn
If you really want to understand AV/EDR evasion, you shouldn't start by memorizing techniques
You need to understand:
Where is the defense
What can it see
What telemetry is it receiving
What visibility is the attacker trying to reduce
@reverseengine
در برابر Kernel Mode
برای اینکه بفهمید چرا تکنیک هایی مثل Direct Syscall اصلا تعریف شدن اول باید بدونید ویندوز دو تعریف مهم داره:
┌─────────────────────────┐
│ User Mode │
│ Applications / DLLs │
└───────────┬─────────────┘
│
System Call
│
▼
┌─────────────────────────┐
│ Kernel Mode │
│ Windows Kernel / Drivers│
└─────────────────────────┘
User Mode
برنامههای معمولی اینجا اجرا میشن:
Browser
PowerShell
Game
Your Program
↓
kernel32.dll
↓
ntdll.dll
↓
System Call
EDR
ها میتونن در این لایه telemetry جمع آوری کنن و رفتار برنامه رو بررسی کنن
Kernel Mode
اینجا بخش های حساس سیستم عامل قرار دارن
مثلا مدیریت:
Processes
Threads
Memory
Drivers
File System
Network
به همین دلیل اگر فقط یک لایه از User Mode رو دور بزنید به معنی نامرئی شدن نیست.
-
Direct Syscall
چه مفهومی داره؟
ایده اصلی اینه که بهجای طی کردن مسیر معمول User-Mode API برنامه مستقیما به مرز System Call نزدیک بشید
مفهوم:
Normal:
Application
↓
Win32 API
↓
ntdll
↓
System Call
↓
Kernel
Direct Syscall:
Application
↓
System Call
↓
Kernel
اما نکته مهم همینجاست:
Direct Syscall
به معنی دور زدن کامل EDR نیست
چون Kernel و سایر منابع telemetry همچنان میتونن رفتار اتفاق افتاده رو ببینن
پس چرا هکر ها بهش علاقه دارن؟
چون اگر یک مکانیزم دفاعی مشخص در User Mode قرار گرفته باشه تغییر مسیر اجرای برنامه میتونه روی همان مکانیزم اثر بذاره
ولی EDR مدرن فقط به یک نقطه وابسته نیست:
┌── User Mode
│
Process ─────┼── Memory
│
├── ETW / Telemetry
│
├── Kernel
│
└── Network
بنابراین Evasion واقعی یک بازی چند لایه ست نه پیدا کردن یک API عجیب و غریب
نکتهای که باید یاد بگیرید
اگر بخواید AV/EDR Evasion رو واقعا بفهمید نباید از حفظ کردن تکنیکها شروع کنید
باید بفهمید:
دفاع کجاست
چه چیزی رو میتونه ببینه
چه telemetry دریافت میکنه
مهاجم تلاش میکنه کدوم visibility رو کاهش بده
User Mode vs Kernel Mode
To understand why techniques like Direct Syscall were defined at all, you first need to know that Windows has two important definitions:
┌───────────────────────────┐
│ User Mode │
│ Applications / DLLs │
└─
│
System Call
│
▼
┌────────────� Management:
Processes
Threads
Memory
Drivers
File System
Network
That's why bypassing just one layer of User Mode doesn't mean you'll be invisible.
Direct Syscall
What does it mean?
The main idea is to approach the System Call boundary directly instead of going through the usual User-Mode API path.
Conceptual:
Normal:
Application
↓
Win32 API
↓
ntdll
↓
System Call
↓
Kernel
Direct Syscall:
Application
↓
System Call
↓
Kernel
But here's the important point:
Direct Syscall
does not mean bypassing EDR completely
Because the Kernel and other telemetry sources can still see the behavior that happened
So why are hackers interested in it?
Because if a specific defense mechanism is in User Mode, rerouting the program can affect that mechanism
But modern EDR is not just about one point:
┌── User Mode
│
Process ──────┼── Memory
│
├── ETW / Telemetry
│
├── Kernel
│
└── Network
So real evasion is a multi-layered game, not about finding a weird API
What you need to learn
If you really want to understand AV/EDR evasion, you shouldn't start by memorizing techniques
You need to understand:
Where is the defense
What can it see
What telemetry is it receiving
What visibility is the attacker trying to reduce
@reverseengine
Heap Overflow
یکی از معروفترین باگهای Heap:
Heap Overflow
,
ایدهاش خیلی شبیه Buffer Overflow روی Stack هست با این تفاوت که این بار سر ریز داخل Heap اتفاق میوفته
Heap Overflow
فرض کنید برنامه از Heap یک Chunk با ظرفیت مشخص میگیره:
C
یعنی برنامه فضایی برای نگهداری داده در اختیار داره حالا اگر برنامه بیشتر از ظرفیتی که برای این Chunk در نظر گرفته شده داخلش بنویسه داده از محدوده خودش خارج میشه
به این اتفاق میگیم:
Heap Overflow
به شکل ساده:
[ Buffer A ][ Buffer B ]
████████████████████
↓
نوشتن بیش از ظرفیت A
↓
[ Buffer A ][AAAAAAA...]
↑
وارد محدوده B
چرا این اتفاق خطرناکه؟
چون Chunk ها معمولا کنار هم قرار میگیرن پس اگر برنامه از محدوده ی Chunk خودش خارج بشه ممکنه داده های مربوط به حافظه ی مجاور رو تغییر بده
اون حافظه ی مجاور میتونه متعلق به:
یک Object دیگه
یک ساختار داده
اطلاعات مدیریتی Heap
یا دادههای مهم برنامه باشه
در نتیجه یک اشتباه ساده در اندازه ی داده میتونه روی بخش دیگری از برنامه تاثیر بذاره
تفاوت Heap Overflow و Stack Overflow
هر دو از یک ایدهی کلی پیروی میکنن:
نوشتن بیشتر از ظرفیت حافظه ای که در اختیار برنامه قرار گرفته
اما محل اتفاق فرق داره
Stack Overflow:
Stack
↓
Buffer
↓
داده بیشتر از ظرفیت
↓
اطلاعات اطراف Buffer
Heap Overflow:
Heap
↓
Chunk
↓
داده بیشتر از ظرفیت
↓
Chunk یا Metadata مجاور
پس نباید این دو رو یکی بدونیم
چرا Chunkهای مجاور اهمیت دارن؟
فرض کنید Heap این شکلیه:
+----------------+
| Chunk A |
+----------------+
| Chunk B |
+----------------+
| Chunk C |
+----------------+
اگر برنامه داخل Chunk A بیشتر از ظرفیتش بنویسه ممکنه نوشتن وارد محدوده ی Chunk B بشه در نتیجه ممکنه دادهی B تغییر کنه بدون اینکه برنامه مستقیما قصد تغییر B رو داشته باشه این دقیقا همون چیزی هست که Heap Overflow رو خطرناک میکنه
Metadata
هم میتونه مهم باشه یادتون هست گفتیم Chunk علاوه بر User Data اطلاعات مدیریتی هم داره اگر یک Overflow از محدوده ی خودش خارج بشه در بعضی شرایط ممکنه به Metadata مربوط به بخشهای مجاور هم برسه اینجاست که موضوع پیچیده تر میشه چون دیگه فقط دادهی برنامه تغییر نکرده ممکنه اطلاعاتی که Allocator برای مدیریت Heap استفاده میکنه هم تحت تاثیر قرار گرفته باشه البته Allocator های مدرن بررسی های مختلفی دارن و خیلی از روش های قدیمی Heap Exploitation دیگه به سادگی گذشته قابل استفاده نیستن
یک نکته خیلی مهم
هر Heap Overflow حتما قابل تبدیل شدن به Exploit نیست
ممکنه فقط باعث بشه:
برنامه Crash کنه
یک مقدار اشتباه تغییر کنه
دادهای خراب بشه
یا رفتار برنامه غیرقابل پیشبینی بشه
برای اینکه یک Memory Corruption واقعا قابل سواستفاده بشه باید بررسی کنیم چه چیزی قابل تغییره و این تغییر چه اثری روی برنامه داره
پس:
وجود باگ فقط نقطهی شروع تحلیلمونه
@reverseengine
یکی از معروفترین باگهای Heap:
Heap Overflow
,
ایدهاش خیلی شبیه Buffer Overflow روی Stack هست با این تفاوت که این بار سر ریز داخل Heap اتفاق میوفته
Heap Overflow
فرض کنید برنامه از Heap یک Chunk با ظرفیت مشخص میگیره:
C
char *buffer = malloc(64);
یعنی برنامه فضایی برای نگهداری داده در اختیار داره حالا اگر برنامه بیشتر از ظرفیتی که برای این Chunk در نظر گرفته شده داخلش بنویسه داده از محدوده خودش خارج میشه
به این اتفاق میگیم:
Heap Overflow
به شکل ساده:
[ Buffer A ][ Buffer B ]
████████████████████
↓
نوشتن بیش از ظرفیت A
↓
[ Buffer A ][AAAAAAA...]
↑
وارد محدوده B
چرا این اتفاق خطرناکه؟
چون Chunk ها معمولا کنار هم قرار میگیرن پس اگر برنامه از محدوده ی Chunk خودش خارج بشه ممکنه داده های مربوط به حافظه ی مجاور رو تغییر بده
اون حافظه ی مجاور میتونه متعلق به:
یک Object دیگه
یک ساختار داده
اطلاعات مدیریتی Heap
یا دادههای مهم برنامه باشه
در نتیجه یک اشتباه ساده در اندازه ی داده میتونه روی بخش دیگری از برنامه تاثیر بذاره
تفاوت Heap Overflow و Stack Overflow
هر دو از یک ایدهی کلی پیروی میکنن:
نوشتن بیشتر از ظرفیت حافظه ای که در اختیار برنامه قرار گرفته
اما محل اتفاق فرق داره
Stack Overflow:
Stack
↓
Buffer
↓
داده بیشتر از ظرفیت
↓
اطلاعات اطراف Buffer
Heap Overflow:
Heap
↓
Chunk
↓
داده بیشتر از ظرفیت
↓
Chunk یا Metadata مجاور
پس نباید این دو رو یکی بدونیم
چرا Chunkهای مجاور اهمیت دارن؟
فرض کنید Heap این شکلیه:
+----------------+
| Chunk A |
+----------------+
| Chunk B |
+----------------+
| Chunk C |
+----------------+
اگر برنامه داخل Chunk A بیشتر از ظرفیتش بنویسه ممکنه نوشتن وارد محدوده ی Chunk B بشه در نتیجه ممکنه دادهی B تغییر کنه بدون اینکه برنامه مستقیما قصد تغییر B رو داشته باشه این دقیقا همون چیزی هست که Heap Overflow رو خطرناک میکنه
Metadata
هم میتونه مهم باشه یادتون هست گفتیم Chunk علاوه بر User Data اطلاعات مدیریتی هم داره اگر یک Overflow از محدوده ی خودش خارج بشه در بعضی شرایط ممکنه به Metadata مربوط به بخشهای مجاور هم برسه اینجاست که موضوع پیچیده تر میشه چون دیگه فقط دادهی برنامه تغییر نکرده ممکنه اطلاعاتی که Allocator برای مدیریت Heap استفاده میکنه هم تحت تاثیر قرار گرفته باشه البته Allocator های مدرن بررسی های مختلفی دارن و خیلی از روش های قدیمی Heap Exploitation دیگه به سادگی گذشته قابل استفاده نیستن
یک نکته خیلی مهم
هر Heap Overflow حتما قابل تبدیل شدن به Exploit نیست
ممکنه فقط باعث بشه:
برنامه Crash کنه
یک مقدار اشتباه تغییر کنه
دادهای خراب بشه
یا رفتار برنامه غیرقابل پیشبینی بشه
برای اینکه یک Memory Corruption واقعا قابل سواستفاده بشه باید بررسی کنیم چه چیزی قابل تغییره و این تغییر چه اثری روی برنامه داره
پس:
Bug ≠ Exploit
وجود باگ فقط نقطهی شروع تحلیلمونه
Heap Overflow
زمانی اتفاق میوفته که برنامه بیشتر از ظرفیت یک بخش اختصاص یافته روی Heap بنویسه و از محدودهی خودش خارج بشه چون Chunk ها در Heap کنار هم قرار میگیرن این Overflow ممکنه دادهها یا ساختارهای مجاور رو تحت تاثیر قرار بده.برای تحلیل Heap Exploitation باید بفهمیم این Overflow دقیقا چه چیزی رو میتونه تغییر بده و اون تغییر چه اثری روی رفتار برنامه داره
@reverseengine
ReverseEngineering
Heap Overflow یکی از معروفترین باگهای Heap: Heap Overflow , ایدهاش خیلی شبیه Buffer Overflow روی Stack هست با این تفاوت که این بار سر ریز داخل Heap اتفاق میوفته Heap Overflow فرض کنید برنامه از Heap یک Chunk با ظرفیت مشخص میگیره: C char *buffer = malloc(64);…
Heap Overflow
One of the most famous Heap bugs:
Heap Overflow
Its idea is very similar to Buffer Overflow on Stack except that this time the overflow occurs inside the Heap
Heap Overflow
Suppose the program gets a Chunk with a certain capacity from the Heap:
C
char *buffer = malloc(64);
That is, the program has space to store data. Now if the program writes more than the capacity intended for this Chunk, the data will go out of its bounds
We call this event:
Heap Overflow
Simply put:
[ Buffer A ][ Buffer B ]
████████████████████
↓
Writing more than the capacity of A
↓
[ Buffer A ][AAAAAAAA...]
↑
Entering B range
Why is this dangerous?
Because chunks are usually placed next to each other, if the program goes out of its own chunk, it may change the data in the adjacent memory. That adjacent memory could belong to:
Another object
A data structure
Heap management information
Or important program data
As a result, a simple mistake in the size of the data can affect another part of the program
Difference between Heap Overflow and Stack Overflow
Both follow the same general idea:
Writing more than the memory capacity provided to the program
But the location of the event is different
Stack Overflow:
Stack
↓
Buffer
↓
Data more than capacity
↓
Information around Buffer
Heap Overflow:
Heap
↓
Chunk
↓
Data more than capacity
↓
Adjacent Chunk or Metadata
So we should not consider these two as the same
Why are adjacent chunks important?
Suppose the Heap looks like this:
+----------------+
| Chunk A |
+----------------+
| Chunk B |
+----------------+
| Chunk C |
+----------------+
If the program writes more than its capacity into Chunk A, the write may enter the Chunk B area, as a result, the data in B may change without the program directly intending to change B. This is exactly what makes Heap Overflow dangerous
Metadata
can also be important. Remember that we said that Chunk has management information in addition to User Data. If an Overflow goes out of its own scope, in some situations it may also reach the Metadata related to adjacent sections. This is where the matter becomes more complicated because not only the program data has changed, the information that the Allocator uses to manage the Heap may also be affected. Of course, modern Allocators have different checks and many of the old Heap Exploitation methods are no longer as simple as they used to be.
A very important point
Not every Heap Overflow can be turned into an Exploit. It may only cause:
The program to crash
A wrong value to change
Data to be corrupted
Or the program to behave unpredictable
For a Memory Corruption to really To be exploitable, we need to examine what can be changed and what effect this change has on the program
So:
The existence of a bug is only the starting point for our analysis
@reverseengine
One of the most famous Heap bugs:
Heap Overflow
Its idea is very similar to Buffer Overflow on Stack except that this time the overflow occurs inside the Heap
Heap Overflow
Suppose the program gets a Chunk with a certain capacity from the Heap:
C
char *buffer = malloc(64);
That is, the program has space to store data. Now if the program writes more than the capacity intended for this Chunk, the data will go out of its bounds
We call this event:
Heap Overflow
Simply put:
[ Buffer A ][ Buffer B ]
████████████████████
↓
Writing more than the capacity of A
↓
[ Buffer A ][AAAAAAAA...]
↑
Entering B range
Why is this dangerous?
Because chunks are usually placed next to each other, if the program goes out of its own chunk, it may change the data in the adjacent memory. That adjacent memory could belong to:
Another object
A data structure
Heap management information
Or important program data
As a result, a simple mistake in the size of the data can affect another part of the program
Difference between Heap Overflow and Stack Overflow
Both follow the same general idea:
Writing more than the memory capacity provided to the program
But the location of the event is different
Stack Overflow:
Stack
↓
Buffer
↓
Data more than capacity
↓
Information around Buffer
Heap Overflow:
Heap
↓
Chunk
↓
Data more than capacity
↓
Adjacent Chunk or Metadata
So we should not consider these two as the same
Why are adjacent chunks important?
Suppose the Heap looks like this:
+----------------+
| Chunk A |
+----------------+
| Chunk B |
+----------------+
| Chunk C |
+----------------+
If the program writes more than its capacity into Chunk A, the write may enter the Chunk B area, as a result, the data in B may change without the program directly intending to change B. This is exactly what makes Heap Overflow dangerous
Metadata
can also be important. Remember that we said that Chunk has management information in addition to User Data. If an Overflow goes out of its own scope, in some situations it may also reach the Metadata related to adjacent sections. This is where the matter becomes more complicated because not only the program data has changed, the information that the Allocator uses to manage the Heap may also be affected. Of course, modern Allocators have different checks and many of the old Heap Exploitation methods are no longer as simple as they used to be.
A very important point
Not every Heap Overflow can be turned into an Exploit. It may only cause:
The program to crash
A wrong value to change
Data to be corrupted
Or the program to behave unpredictable
For a Memory Corruption to really To be exploitable, we need to examine what can be changed and what effect this change has on the program
So:
Bug ≠ Exploit
The existence of a bug is only the starting point for our analysis
Heap Overflow
Occurs when a program writes more than the capacity of an allocated section on the Heap and goes out of its bounds because the Chunks in the Heap are placed next to each other. This Overflow may affect adjacent data or structures. To analyze Heap Exploitation, we need to understand what exactly this Overflow can change and what effect that change has on the behavior of the program.
@reverseengine
Forwarded from club1337
Devman-ArticleXakep.txt
17.9 KB
Вымогатель-болтун. Как Devman прошел путь от новичка до преступника в розыске Интерпола
👑 Статья для подписчиков
31 июля 2025 года Джон Ди Маджо открыл сообщение в зашифрованном мессенджере. Преступники обычно не любят, когда их деятельность расследуют, но этот написал сам. К сообщению была приложена фотография: дорогие часы, спортивные автомобили. Отправителя Ди Маджо знал.
https://xakep.ru/2026/08/13/devman/
Telegram✉️ @club1337
X (Twitter)🕊 @club31337
👑 Статья для подписчиков
31 июля 2025 года Джон Ди Маджо открыл сообщение в зашифрованном мессенджере. Преступники обычно не любят, когда их деятельность расследуют, но этот написал сам. К сообщению была приложена фотография: дорогие часы, спортивные автомобили. Отправителя Ди Маджо знал.
https://xakep.ru/2026/08/13/devman/
Telegram
X (Twitter)
Please open Telegram to view this post
VIEW IN TELEGRAM
String Obfuscation
وقتی رشتهها هم مخفی میشن
تا اینجا درباره Control Flow Flattening و Opaque Predicate صحبت کردیم
حالا بریم سراغ یکی از چیزهایی که توی تحلیل استاتیک خیلی زود باهاش برخورد میکنید
رشتهها
وقتی یک برنامه رو با IDA یا Ghidra باز میکنیم معمولا یکی از اولین کارها اینه که Strings رو بررسی کنیم
چون رشته هایی مثل این خیلی اطلاعات میدن:
Login failed
Access denied
https://example.com
config.json
username
اما Obfuscation میتونه همین رشتهها رو هم از حالت واضح خارج کنه
مثلا به جای اینکه داخل باینری داشته باشیم:
Access denied
ممکنه فقط یک سری بایت ببینیم:
12 37 21 04 55 19 ...
و برنامه موقع اجرا اونها رو به رشته اصلی تبدیل کنه
یک روش ساده برای این کار XOR هست
مثلا رشته اصلی:
HELLO
با یک کلید مشخص XOR میشه و نتیجه داخل فایل قرار میگیره
وقتی برنامه اجرا میشه دوباره همون عملیات انجام میشه و رشته اصلی برمیگرده
در نتیجه اگر فقط Strings رو روی فایل اجرا کنیم ممکنه اصلا HELLO رو نبینیم
اما اینجا یک نکته خیلی مهم وجود داره
هدف ما این نیست که فقط دنبال رشته خوندن بگردیم
باید بفهمهیم رشته کجا ساخته میشه
مثلا ممکنه داخل دیساسمبل ببینیم:
داده رمزگذاریشده
↓
Decode
↓
Buffer
↓
استفاده توسط برنامه
اگر تابع Decode رو پیدا کنیم میتونیم بفهمیم برنامه چطور رشتهها رو در زمان اجرا بازسازی میکنه
بعضی برنامهها حتی رشتهها رو از اول به صورت کامل در حافظه نگه نمیدارن
ممکنه فقط زمانی که لازم دارن رشته رو Decode کنن و بعد دوباره پاکش کنن
برای همین تحلیل داینامیک اینجا خیلی کمک میکنه
میتونیم ببینیم چه زمانی Buffer ساخته میشه و چه زمانی محتوای قابل خوندن داخلش قرار میگیره
یک نکته مهم دیگه هم اینه که هر دادهای که شبیه متن نیست لزوما رمزگذاری نشده
ممکنه فشرده شده باشه ممکنه یک ساختار باینری باشه یا حتی فقط دادهای باشه که با یک Encoding خاص ذخیره شده
پس قبل از اینکه بگیم این رشته رمزگذاری شده باید مسیر استفاده از اون داده رو بررسی کنیم
تمرین:
یک برنامه ساده بسازید که یک رشته مشخص داشته باشه
بعد رشته رو با یک XOR ساده قبل از ذخیره شدن تغییر بدید و در زمان اجرا دوباره Decode کنید
حالا برنامه رو داخل Ghidra باز کنید
اول Strings رو بررسی کنید
بعد تابعی که داده رو Decode میکنه پیدا کنید
در اخر سعی کنید بدون اجرای برنامه الگوریتم Decode رو از روی اسمبلی بازسازی کنید
اینجا یک چیز مهم به دست میارید:
به جای اینکه فقط دنبال چیزی که برنامه نشون میده بگردید یاد میگیرید بفهمید اون داده چطور ساخته شده
@reverseengine
وقتی رشتهها هم مخفی میشن
تا اینجا درباره Control Flow Flattening و Opaque Predicate صحبت کردیم
حالا بریم سراغ یکی از چیزهایی که توی تحلیل استاتیک خیلی زود باهاش برخورد میکنید
رشتهها
وقتی یک برنامه رو با IDA یا Ghidra باز میکنیم معمولا یکی از اولین کارها اینه که Strings رو بررسی کنیم
چون رشته هایی مثل این خیلی اطلاعات میدن:
Login failed
Access denied
https://example.com
config.json
username
اما Obfuscation میتونه همین رشتهها رو هم از حالت واضح خارج کنه
مثلا به جای اینکه داخل باینری داشته باشیم:
Access denied
ممکنه فقط یک سری بایت ببینیم:
12 37 21 04 55 19 ...
و برنامه موقع اجرا اونها رو به رشته اصلی تبدیل کنه
یک روش ساده برای این کار XOR هست
مثلا رشته اصلی:
HELLO
با یک کلید مشخص XOR میشه و نتیجه داخل فایل قرار میگیره
وقتی برنامه اجرا میشه دوباره همون عملیات انجام میشه و رشته اصلی برمیگرده
در نتیجه اگر فقط Strings رو روی فایل اجرا کنیم ممکنه اصلا HELLO رو نبینیم
اما اینجا یک نکته خیلی مهم وجود داره
هدف ما این نیست که فقط دنبال رشته خوندن بگردیم
باید بفهمهیم رشته کجا ساخته میشه
مثلا ممکنه داخل دیساسمبل ببینیم:
داده رمزگذاریشده
↓
Decode
↓
Buffer
↓
استفاده توسط برنامه
اگر تابع Decode رو پیدا کنیم میتونیم بفهمیم برنامه چطور رشتهها رو در زمان اجرا بازسازی میکنه
بعضی برنامهها حتی رشتهها رو از اول به صورت کامل در حافظه نگه نمیدارن
ممکنه فقط زمانی که لازم دارن رشته رو Decode کنن و بعد دوباره پاکش کنن
برای همین تحلیل داینامیک اینجا خیلی کمک میکنه
میتونیم ببینیم چه زمانی Buffer ساخته میشه و چه زمانی محتوای قابل خوندن داخلش قرار میگیره
یک نکته مهم دیگه هم اینه که هر دادهای که شبیه متن نیست لزوما رمزگذاری نشده
ممکنه فشرده شده باشه ممکنه یک ساختار باینری باشه یا حتی فقط دادهای باشه که با یک Encoding خاص ذخیره شده
پس قبل از اینکه بگیم این رشته رمزگذاری شده باید مسیر استفاده از اون داده رو بررسی کنیم
تمرین:
یک برنامه ساده بسازید که یک رشته مشخص داشته باشه
بعد رشته رو با یک XOR ساده قبل از ذخیره شدن تغییر بدید و در زمان اجرا دوباره Decode کنید
حالا برنامه رو داخل Ghidra باز کنید
اول Strings رو بررسی کنید
بعد تابعی که داده رو Decode میکنه پیدا کنید
در اخر سعی کنید بدون اجرای برنامه الگوریتم Decode رو از روی اسمبلی بازسازی کنید
اینجا یک چیز مهم به دست میارید:
به جای اینکه فقط دنبال چیزی که برنامه نشون میده بگردید یاد میگیرید بفهمید اون داده چطور ساخته شده
@reverseengine
ReverseEngineering
String Obfuscation وقتی رشتهها هم مخفی میشن تا اینجا درباره Control Flow Flattening و Opaque Predicate صحبت کردیم حالا بریم سراغ یکی از چیزهایی که توی تحلیل استاتیک خیلی زود باهاش برخورد میکنید رشتهها وقتی یک برنامه رو با IDA یا Ghidra باز میکنیم معمولا…
String Obfuscation
When strings are also hidden
So far we have talked about Control Flow Flattening and Opaque Predicate
Now let's move on to one of the things you will encounter very soon in static analysis
Strings
When we open a program with IDA or Ghidra, one of the first things we usually do is check the Strings
Because strings like this give a lot of information:
Login failed
Access denied
https://example.com
config.json
username
But Obfuscation can also remove these strings from the clear state
For example, instead of having inside the binary:
Access denied
We may only see a series of bytes:
12 37 21 04 55 19 ...
And the program converts them into the original string when it runs
A simple way to do this is XOR
For example, the original string:
HELLO
It is XORed with a specific key and the result is placed in the file takes
When the program is run, the same operation is performed again and the original string is returned
As a result, if we just run Strings on the file, we may not see HELLO at all
But there is a very important point here
Our goal is not to just look for the string to read
We need to understand where the string is created
For example, we may see in the disassembler:
Encrypted data
↓
Decode
↓
Buffer
↓
Used by the program
If we find the Decode function, we can understand how the program reconstructs the strings at runtime
Some programs do not even keep the strings completely in memory from the beginning
They may only need to decode the string and then clear it again
That is why dynamic analysis is very helpful here
We can see when the Buffer is created and when the readable content is placed in it
Another important point is that any data that does not look like text is necessarily encrypted Not
It could be compressed, it could be a binary structure, or even just data stored with a specific encoding
So before we say this string is encrypted, we need to look at how that data is used
Exercise:
Write a simple program that takes a given string
Then modify the string with a simple XOR before saving it and decode it again at runtime
Now open the program in Ghidra
First examine the Strings
Then find the function that decodes the data
Finally, try to recreate the Decode algorithm from assembly without running the program
Here you will learn something important:
Instead of just looking for what the program shows, you will learn to understand how that data is constructed
@reverseengine
When strings are also hidden
So far we have talked about Control Flow Flattening and Opaque Predicate
Now let's move on to one of the things you will encounter very soon in static analysis
Strings
When we open a program with IDA or Ghidra, one of the first things we usually do is check the Strings
Because strings like this give a lot of information:
Login failed
Access denied
https://example.com
config.json
username
But Obfuscation can also remove these strings from the clear state
For example, instead of having inside the binary:
Access denied
We may only see a series of bytes:
12 37 21 04 55 19 ...
And the program converts them into the original string when it runs
A simple way to do this is XOR
For example, the original string:
HELLO
It is XORed with a specific key and the result is placed in the file takes
When the program is run, the same operation is performed again and the original string is returned
As a result, if we just run Strings on the file, we may not see HELLO at all
But there is a very important point here
Our goal is not to just look for the string to read
We need to understand where the string is created
For example, we may see in the disassembler:
Encrypted data
↓
Decode
↓
Buffer
↓
Used by the program
If we find the Decode function, we can understand how the program reconstructs the strings at runtime
Some programs do not even keep the strings completely in memory from the beginning
They may only need to decode the string and then clear it again
That is why dynamic analysis is very helpful here
We can see when the Buffer is created and when the readable content is placed in it
Another important point is that any data that does not look like text is necessarily encrypted Not
It could be compressed, it could be a binary structure, or even just data stored with a specific encoding
So before we say this string is encrypted, we need to look at how that data is used
Exercise:
Write a simple program that takes a given string
Then modify the string with a simple XOR before saving it and decode it again at runtime
Now open the program in Ghidra
First examine the Strings
Then find the function that decodes the data
Finally, try to recreate the Decode algorithm from assembly without running the program
Here you will learn something important:
Instead of just looking for what the program shows, you will learn to understand how that data is constructed
@reverseengine
DWARF as a Shared Reverse Engineering Format
https://lief.re/blog/2025-05-27-dwarf-editor
@reverseengine
https://lief.re/blog/2025-05-27-dwarf-editor
@reverseengine
LIEF
DWARF as a Shared Reverse Engineering Format · LIEF
Create DWARF debug information with the LIEF Extended DWARF editor and share reverse-engineered types and functions between Binary Ninja and Ghidra plugins.
ReverseEngineering
String Obfuscation When strings are also hidden So far we have talked about Control Flow Flattening and Opaque Predicate Now let's move on to one of the things you will encounter very soon in static analysis Strings When we open a program with IDA or Ghidra…
Understanding String Obfuscation
https://medium.com/mobilepeople/understanding-string-obfuscation-8f25f3ff0416
https://medium.com/mobilepeople/understanding-string-obfuscation-8f25f3ff0416
Medium
Understanding String Obfuscation
I love the way of understanding a topic with some background, history, etc. When I came across string obfuscation and searched it over the…
Constant Obfuscation وقتی حتی عدد ها هم مخفی میشن
تا اینجا دیدیم که چطور رشتهها رو میشه مخفی کردولی فقط رشتهها نیستن که Obfuscate میشن عددها و ثابت های برنامه هم میتونن مخفی بشن
فرض کنید برنامه باید مقدار
حالا برنامه نویس یا ابزار Obfuscation میتونه همون مقدار رو به شکل پیچیده تری تولید کنه
مثلا:
در نهایت مقدار
اما یک قدم جلوتر:
نتیجه XOR دوباره میتونه یک مقدار مشخص باشه در این حالت وقتی فقط یک دستور رو میبینید مقدار واقعی ثابت فورا مشخص نمیشه بعصی وقتا حتی محاسبات بیشتری استفاده میشه:
و همه اینها فقط برای تولید یک مقدار ثابت انجام میشن مثلا ممکنه برنامه برای ساختن یک عدد ساده چند تا دستور اجرا کنه
اینجا کاری که ما انجام میدیم اینه که به جای نگاه کردن به تک تک دستورها مسیر تولید مقدار رو دنبال میکنیم
یعنی میپرسیم:
این مقدار از کجا اومد؟
چه عملیاتی روش انجام شده؟
در نهایت کجا استفاده شده؟
مثلا:
مقدار اولیه
اگر بتونید این زنجیره رو ساده کنید مقدار واقعی ثابت دوباره مشخص میشه
این کار مخصوصا وقتی مهم میشه که ثابتها بخشی از یک الگوریتم باشن
مثلا یک برنامه ممکنه یک مقدار ثابت رو برای مقایسه محاسبه یا ساختن یک جدول استفاده کنه اگر مقدار اصلی مخفی شده باشه فهمیدن الگوریتم هم سخت تر میشه
یک نکته جالب اینه که بعضی وقتا Decompiler خودش میتونه این محاسبات رو ساده کنه
مثلا چند دستور اسمبلی رو تبدیل کنه به:
اما همیشه نباید به خروجی Decompiler اعتماد کرد گاهی Obfuscation باعث میشه خروجی چیزی کاملا پیچیده و گمراهکننده باشه
پس یکی از مهارتهای مهم Reverse Engineer اینه که بتونه بین سه چیز حرکت کنه:
تمرین:
این عبارت رو بدون اجرای برنامه ساده کنید:
بعد یک مثال پیچیده تر برای خودتون بسازید که در نهایت به یک عدد مشخص برسه بعد همون برنامه رو Compile کنید و داخل Ghidra ببینید Compiler چه شکلی از اون ساخته
هدف تمرین این نیست که فقط جواب عدد رو پیدا کنید هدف اینه که یاد بگیرید یک مقدار رو از مسیر محاسباتش دنبال کنید
@reverseengine
تا اینجا دیدیم که چطور رشتهها رو میشه مخفی کردولی فقط رشتهها نیستن که Obfuscate میشن عددها و ثابت های برنامه هم میتونن مخفی بشن
فرض کنید برنامه باید مقدار
100 رو استفاده کنه در حالت عادی ممکنه توی اسمبلی چیزی شبیه این ببینید:mov eax, 100
حالا برنامه نویس یا ابزار Obfuscation میتونه همون مقدار رو به شکل پیچیده تری تولید کنه
مثلا:
mov eax, 73
add eax, 27
در نهایت مقدار
eax میشه 100اما یک قدم جلوتر:
mov eax, 0x12345678
xor eax, 0x12345614
نتیجه XOR دوباره میتونه یک مقدار مشخص باشه در این حالت وقتی فقط یک دستور رو میبینید مقدار واقعی ثابت فورا مشخص نمیشه بعصی وقتا حتی محاسبات بیشتری استفاده میشه:
XOR
ADD
SUB
ROL
ROR
و همه اینها فقط برای تولید یک مقدار ثابت انجام میشن مثلا ممکنه برنامه برای ساختن یک عدد ساده چند تا دستور اجرا کنه
اینجا کاری که ما انجام میدیم اینه که به جای نگاه کردن به تک تک دستورها مسیر تولید مقدار رو دنبال میکنیم
یعنی میپرسیم:
این مقدار از کجا اومد؟
چه عملیاتی روش انجام شده؟
در نهایت کجا استفاده شده؟
مثلا:
مقدار اولیه
↓مقدار نهایی
XOR
↓
ADD
↓
SUB
↓
اگر بتونید این زنجیره رو ساده کنید مقدار واقعی ثابت دوباره مشخص میشه
این کار مخصوصا وقتی مهم میشه که ثابتها بخشی از یک الگوریتم باشن
مثلا یک برنامه ممکنه یک مقدار ثابت رو برای مقایسه محاسبه یا ساختن یک جدول استفاده کنه اگر مقدار اصلی مخفی شده باشه فهمیدن الگوریتم هم سخت تر میشه
یک نکته جالب اینه که بعضی وقتا Decompiler خودش میتونه این محاسبات رو ساده کنه
مثلا چند دستور اسمبلی رو تبدیل کنه به:
x = 100;
اما همیشه نباید به خروجی Decompiler اعتماد کرد گاهی Obfuscation باعث میشه خروجی چیزی کاملا پیچیده و گمراهکننده باشه
پس یکی از مهارتهای مهم Reverse Engineer اینه که بتونه بین سه چیز حرکت کنه:
Assembly
Decompiler
منطق واقعی برنامه
تمرین:
این عبارت رو بدون اجرای برنامه ساده کنید:
int x = 73;
x = x + 27;
x = x ^ 0;
بعد یک مثال پیچیده تر برای خودتون بسازید که در نهایت به یک عدد مشخص برسه بعد همون برنامه رو Compile کنید و داخل Ghidra ببینید Compiler چه شکلی از اون ساخته
هدف تمرین این نیست که فقط جواب عدد رو پیدا کنید هدف اینه که یاد بگیرید یک مقدار رو از مسیر محاسباتش دنبال کنید
@reverseengine
ReverseEngineering
Constant Obfuscation وقتی حتی عدد ها هم مخفی میشن تا اینجا دیدیم که چطور رشتهها رو میشه مخفی کردولی فقط رشتهها نیستن که Obfuscate میشن عددها و ثابت های برنامه هم میتونن مخفی بشن فرض کنید برنامه باید مقدار 100 رو استفاده کنه در حالت عادی ممکنه توی اسمبلی…
Constant Obfuscation When Even Numbers Are Hiding
So far we have seen how strings can be hidden, but it is not only strings that are obfuscated, numbers and program constants can also be hidden
Suppose the program needs to use the value 100, in normal case you might see something like this in the assembly:
Now the programmer or the Obfuscation tool can generate the same value in a more complex way
For example:
Finally the value of eax becomes 100
But one step further:
The result of XOR can again be a specific value. In this case, when you see only one instruction, the actual value of the constant is not immediately clear. Sometimes even more calculations are used:
And all this
They are only used to produce a constant value. For example, a program might execute several instructions to generate a simple number. What we do here is that instead of looking at each instruction individually, we follow the path of the value. That is, we ask:
Where did this value come from?
What operation did the method perform?
Where was it finally used?
For example:
Initial value
If you can simplify this chain, the actual value of the constant is revealed again. This is especially important when constants are part of an algorithm. For example, a program might use a constant value to compare calculations or build a table. If the original value is hidden, it becomes harder to understand the algorithm. One interesting thing is that sometimes the decompiler itself can simplify these calculations.
For example, it can convert a few assembly instructions to:
But you shouldn't always trust the output of the Decompiler. Sometimes Obfuscation can make the output of something quite complex and misleading.
So one of the important skills of a Reverse Engineer is to be able to move between three things:
Assembly
Decompiler
The actual logic of the program
Exercise:
Simplify this expression without running the program:
Then create a more complex example for yourself that will eventually reach a specific number. Then compile the same program and see what the Compiler makes of it in Ghidra.
The goal of the exercise is not to just find the answer to the number. The goal is to learn to follow a value through its calculation path.
@reverseengine
So far we have seen how strings can be hidden, but it is not only strings that are obfuscated, numbers and program constants can also be hidden
Suppose the program needs to use the value 100, in normal case you might see something like this in the assembly:
mov eax, 100
Now the programmer or the Obfuscation tool can generate the same value in a more complex way
For example:
mov eax, 73
add eax, 27
Finally the value of eax becomes 100
But one step further:
mov eax, 0x12345678
xor eax, 0x12345614
The result of XOR can again be a specific value. In this case, when you see only one instruction, the actual value of the constant is not immediately clear. Sometimes even more calculations are used:
XOR
ADD
SUB
ROL
ROR
And all this
They are only used to produce a constant value. For example, a program might execute several instructions to generate a simple number. What we do here is that instead of looking at each instruction individually, we follow the path of the value. That is, we ask:
Where did this value come from?
What operation did the method perform?
Where was it finally used?
For example:
Initial value
↓
XOR
↓
ADD
↓
SUB
↓
Final value
If you can simplify this chain, the actual value of the constant is revealed again. This is especially important when constants are part of an algorithm. For example, a program might use a constant value to compare calculations or build a table. If the original value is hidden, it becomes harder to understand the algorithm. One interesting thing is that sometimes the decompiler itself can simplify these calculations.
For example, it can convert a few assembly instructions to:
x = 100;
But you shouldn't always trust the output of the Decompiler. Sometimes Obfuscation can make the output of something quite complex and misleading.
So one of the important skills of a Reverse Engineer is to be able to move between three things:
Assembly
Decompiler
The actual logic of the program
Exercise:
Simplify this expression without running the program:
int x = 73;
x = x + 27;
x = x ^ 0;
Then create a more complex example for yourself that will eventually reach a specific number. Then compile the same program and see what the Compiler makes of it in Ghidra.
The goal of the exercise is not to just find the answer to the number. The goal is to learn to follow a value through its calculation path.
@reverseengine
Pointer
چیه و چرا در Binary Exploitation مهمه؟
تا اینجا دربارهی Heap و باگ هایی مثل UAF Double Free و Heap Overflow صحبت کردیم
اما برای اینکه بفهمیم این باگها چطور میتونن روی رفتار برنامه تاثیر بذارن باید یک مفهوم پایه ای رو خیلی خوب بلد باشیم: Pointer
Pointer
به زبون ساده Pointer متغیریه که به جای خود داده آدرس اون داده در حافظه رو نگه میداره
مثلا:
اینجا:
پس ptr خودش مقدارش 100 نیست
بلکه میدونه 100 کجای حافظه قرار گرفته
چرا Pointer در C و C++ اینقدر مهمه؟
چون خیلی از کارهای سطح پایین با Pointer انجام میشن
مثلا وقتی مینویسیم:
متغیر buffer یک Pointer هست
یعنی malloc() یک بخش از Heap رو در اختیار برنامه قرار میده و آدرس اون بخش داخل buffer قرار میگیره
به شکل ساده:
وقتی Pointer دیگه به یک حافظه ی معتبر اشاره نکنه
مثلا:
بعد از free() استفاده از buffer میتونه مشکلساز باشه
چون Pointer هنوز ممکنه یک آدرس داشته باشه اما حافظهی مربوط به اون دیگه معتبر نیست
این همون مفهومی بود که در Use-After-Free دیدیم
Dangling Pointer
به Pointer ی که به یک ناحیهی حافظهی دیگه معتبر نیست معمولا Dangling Pointer میگیم
مثلا:
اما اون آدرس دیگه نباید به عنوان یک Object معتبر استفاده بشه
Pointer و Heap Overflow
چه ارتباطی دارن
فرض کنید داخل Heap چند Object وجود داره:
یعنی یک خطای نوشتن میتونه در ادامه باعث یک خطای خوندن یا نوشتن در جای دیگری از حافظه بشه
به همین دلیل Pointer ها در تحلیل Memory Corruption اهمیت زیادی دارن
Pointer با آدرس یکیه؟
تقریبا ولی بهتره دقیقتر بگیم:
Pointer
یک متغیر یا مقدار دادهایه که یک آدرس حافظه رو نگه میداره
مثلا:
اینجا ptr یک Pointer هست و مقدارش یک آدرس محسوب میشه
خود Pointer هم در حافظه ذخیره میشه
یعنی:
چرا این موضوع برای Exploitation مهمه؟
چون در Binary Exploitation فقط دنبال خراب کردن دادهها نیستیم
یکی از سوالهای مهم اینه:
آیا میتونیم چیزی رو تغییر بدیم که برنامه بعدا از اون به عنوان یک آدرس استفاده کنه؟
اگر جواب مثبت باشه یک Memory Corruption میتونه اثر خیلی بیشتری از تغییر یک عدد معمولی داشته باشه
به همین دلیل هنگام تحلیل Heap باید به Pointer های موجود در:
Objectها
ساختارهای داده
Heap metadata
جدولها و reference ها
توجه ویژه داشته باشیم
Pointer
متغیریه که یک آدرس حافظه رو نگه میداره در برنامههای C و C++ پوینتر ها نقش بسیار مهمی در مدیریت Heap و دسترسی به Objectها دارن اگر یک Pointer خراب یا نامعتبر بشه ممکنه برنامه بعدا به آدرس اشتباهی دسترسی پیدا کنه به همین دلیل شناخت Pointer ها برای درک UAF Heap Overflow و بسیاری از Memory Corruption ها ضروریه
@reverseengine
چیه و چرا در Binary Exploitation مهمه؟
تا اینجا دربارهی Heap و باگ هایی مثل UAF Double Free و Heap Overflow صحبت کردیم
اما برای اینکه بفهمیم این باگها چطور میتونن روی رفتار برنامه تاثیر بذارن باید یک مفهوم پایه ای رو خیلی خوب بلد باشیم: Pointer
Pointer
به زبون ساده Pointer متغیریه که به جای خود داده آدرس اون داده در حافظه رو نگه میداره
مثلا:
int value = 100;
int *ptr = &value;
اینجا:
value
│
│ 100
▼
[ 100 ]
ptr
│
└──────────► آدرس value
پس ptr خودش مقدارش 100 نیست
بلکه میدونه 100 کجای حافظه قرار گرفته
چرا Pointer در C و C++ اینقدر مهمه؟
چون خیلی از کارهای سطح پایین با Pointer انجام میشن
مثلا وقتی مینویسیم:
char *buffer = malloc(100);
متغیر buffer یک Pointer هست
یعنی malloc() یک بخش از Heap رو در اختیار برنامه قرار میده و آدرس اون بخش داخل buffer قرار میگیره
به شکل ساده:
bufferمشکل از کجا شروع میشه؟
│
▼
Heap
+----------------+
| 100 bytes |
+----------------+
وقتی Pointer دیگه به یک حافظه ی معتبر اشاره نکنه
مثلا:
char *buffer = malloc(100);
free(buffer);
بعد از free() استفاده از buffer میتونه مشکلساز باشه
چون Pointer هنوز ممکنه یک آدرس داشته باشه اما حافظهی مربوط به اون دیگه معتبر نیست
این همون مفهومی بود که در Use-After-Free دیدیم
Dangling Pointer
به Pointer ی که به یک ناحیهی حافظهی دیگه معتبر نیست معمولا Dangling Pointer میگیم
مثلا:
bufferخود buffer هنوز ممکنه مقدار قبلی رو داشته باشه
│
▼
[ Chunk ]
│
free()
│
▼
[ آزاد شده ]
اما اون آدرس دیگه نباید به عنوان یک Object معتبر استفاده بشه
Pointer و Heap Overflow
چه ارتباطی دارن
فرض کنید داخل Heap چند Object وجود داره:
+----------------+اگر یک Memory Corruption باعث تغییر یک Pointer بشه برنامه ممکنه بعدا از اون Pointer برای دسترسی به یک آدرس متفاوت استفاده کنه
| Object A |
+----------------+
| Pointer |
+----------------+
| Object B |
+----------------+
یعنی یک خطای نوشتن میتونه در ادامه باعث یک خطای خوندن یا نوشتن در جای دیگری از حافظه بشه
به همین دلیل Pointer ها در تحلیل Memory Corruption اهمیت زیادی دارن
Pointer با آدرس یکیه؟
تقریبا ولی بهتره دقیقتر بگیم:
Pointer
یک متغیر یا مقدار دادهایه که یک آدرس حافظه رو نگه میداره
مثلا:
ptr = 0x12345678
اینجا ptr یک Pointer هست و مقدارش یک آدرس محسوب میشه
خود Pointer هم در حافظه ذخیره میشه
یعنی:
Pointerپس Pointer هم خودش یک داده است و طبیعتا میتونه تحت تاثیر Memory Corruption قرار بگیره
│
▼
[ address ]
چرا این موضوع برای Exploitation مهمه؟
چون در Binary Exploitation فقط دنبال خراب کردن دادهها نیستیم
یکی از سوالهای مهم اینه:
آیا میتونیم چیزی رو تغییر بدیم که برنامه بعدا از اون به عنوان یک آدرس استفاده کنه؟
اگر جواب مثبت باشه یک Memory Corruption میتونه اثر خیلی بیشتری از تغییر یک عدد معمولی داشته باشه
به همین دلیل هنگام تحلیل Heap باید به Pointer های موجود در:
Objectها
ساختارهای داده
Heap metadata
جدولها و reference ها
توجه ویژه داشته باشیم
Pointer
متغیریه که یک آدرس حافظه رو نگه میداره در برنامههای C و C++ پوینتر ها نقش بسیار مهمی در مدیریت Heap و دسترسی به Objectها دارن اگر یک Pointer خراب یا نامعتبر بشه ممکنه برنامه بعدا به آدرس اشتباهی دسترسی پیدا کنه به همین دلیل شناخت Pointer ها برای درک UAF Heap Overflow و بسیاری از Memory Corruption ها ضروریه
@reverseengine
ReverseEngineering
Pointer چیه و چرا در Binary Exploitation مهمه؟ تا اینجا دربارهی Heap و باگ هایی مثل UAF Double Free و Heap Overflow صحبت کردیم اما برای اینکه بفهمیم این باگها چطور میتونن روی رفتار برنامه تاثیر بذارن باید یک مفهوم پایه ای رو خیلی خوب بلد باشیم: Pointer…
What is Pointer
and why is it important in Binary Exploitation?
So far we have talked about Heap and bugs like UAF Double Free and Heap Overflow
But to understand how these bugs can affect the behavior of the program, we need to know a basic concept very well: Pointer
Pointer
In simple terms, a Pointer is a variable that holds the address of that data in memory instead of its own data
For example:
Here:
So ptr itself does not have the value 100
But it knows where 100 is located in memory
Why is Pointer so important in C and C++?
Because many low-level tasks are done with Pointers
For example, when we write:
The buffer variable is a Pointer
That is, malloc() provides a section of the Heap to the program and the address of that section is placed in the buffer
In simple terms:
When the Pointer no longer points to a valid memory
For example:
Using a buffer after free() can be problematic
Because the Pointer may still have an address, but the memory associated with it is no longer valid
This was the same concept we saw in Use-After-Free
Dangling Pointer
A Pointer that is not valid to another memory area is usually called a Dangling Pointer
For example:
But that address should no longer be used as a valid Object
What is the relationship between Pointer and Heap Overflow
Suppose there are several Objects in the Heap:
That is, a write error can subsequently cause a read or write error elsewhere in memory
That is why pointers are very important in memory corruption analysis
Is a pointer the same as an address?
Roughly, but to be more precise:
A pointer
is a variable or data value that holds a memory address
For example:
Here ptr is a pointer and its value is considered an address
The pointer itself is also stored in memory
That is:
Why is this important for exploitation?
Because in Binary Exploitation we are not just looking to corrupt data
One of the important questions is:
Can we change something that the program will later use as an address?
If the answer is yes, a Memory Corruption can have a much greater effect than changing a regular number
That is why when analyzing the Heap, we should pay special attention to the Pointers in:
Pointer is a variable that holds a memory address. In C and C++ programs, pointers play a very important role in Heap management and accessing objects. If a Pointer becomes corrupted or invalid, the program may later access the wrong address. That is why understanding Pointers is essential to understanding UAF Heap Overflow and many Memory Corruptions
@reverseengine
and why is it important in Binary Exploitation?
So far we have talked about Heap and bugs like UAF Double Free and Heap Overflow
But to understand how these bugs can affect the behavior of the program, we need to know a basic concept very well: Pointer
Pointer
In simple terms, a Pointer is a variable that holds the address of that data in memory instead of its own data
For example:
int value = 100;
int *ptr = &value;
Here:
value
│
│ 100
▼
[ 100 ]
ptr
│
└────────► address of value
So ptr itself does not have the value 100
But it knows where 100 is located in memory
Why is Pointer so important in C and C++?
Because many low-level tasks are done with Pointers
For example, when we write:
char *buffer = malloc(100);
The buffer variable is a Pointer
That is, malloc() provides a section of the Heap to the program and the address of that section is placed in the buffer
In simple terms:
bufferWhere does the problem start?
│
▼
Heap
+----------------+
| 100 bytes |
+----------------+
When the Pointer no longer points to a valid memory
For example:
char *buffer = malloc(100);
free(buffer);
Using a buffer after free() can be problematic
Because the Pointer may still have an address, but the memory associated with it is no longer valid
This was the same concept we saw in Use-After-Free
Dangling Pointer
A Pointer that is not valid to another memory area is usually called a Dangling Pointer
For example:
bufferThe buffer itself may still have the previous value
│
▼
[ Chunk ]
│
free()
│
▼
[ Freed ]
But that address should no longer be used as a valid Object
What is the relationship between Pointer and Heap Overflow
Suppose there are several Objects in the Heap:
+----------------+If a memory corruption causes a pointer to change, the program may later use that pointer to access a different address
| Object A |
+----------------+
| Pointer |
+----------------+
| Object B |
+----------------+
That is, a write error can subsequently cause a read or write error elsewhere in memory
That is why pointers are very important in memory corruption analysis
Is a pointer the same as an address?
Roughly, but to be more precise:
A pointer
is a variable or data value that holds a memory address
For example:
ptr = 0x12345678
Here ptr is a pointer and its value is considered an address
The pointer itself is also stored in memory
That is:
PointerSo the pointer itself is data and can naturally be affected by memory corruption
│
▼
[ address ]
Why is this important for exploitation?
Because in Binary Exploitation we are not just looking to corrupt data
One of the important questions is:
Can we change something that the program will later use as an address?
If the answer is yes, a Memory Corruption can have a much greater effect than changing a regular number
That is why when analyzing the Heap, we should pay special attention to the Pointers in:
Objects
Data structures
Heap metadata
Tables and references
Pointer is a variable that holds a memory address. In C and C++ programs, pointers play a very important role in Heap management and accessing objects. If a Pointer becomes corrupted or invalid, the program may later access the wrong address. That is why understanding Pointers is essential to understanding UAF Heap Overflow and many Memory Corruptions
@reverseengine
Direct Syscall vs Indirect Syscall
تا اینجا فهمیدیم Direct Syscall یعنی برنامه تلاش میکنه مسیر معمول User-Mode API رو کوتاه تر کنه
حالا سوال:
Indirect Syscall
ایده اصلی
در Direct Syscall اجرای دستور
بهصورت مفهومی:
اما در Indirect Syscall ایده اینه که اجرای
هدف مفهومی این تکنیک تغییر شکل Call Stack و مسیر User-Mode execution نسبت به Direct Syscall هست
چرا این موضوع برای EDR مهمه؟
EDR
فقط نمیپرسه:
syscall اتفاق افتاد؟
بلکه میتونه سوالهای بیشتری بپرسه:
چه Process ی syscall رو انجام داده؟
Call Stack چطور بوده؟
Thread از کجا شروع شده؟
Memory مربوط به کجاست؟
قبل و بعد از syscall چه اتفاقی افتاده؟
پس:
Direct Syscall
≠
Invisible
و:
Indirect Syscall
EDR Bypass
این دو بیشتر تغییر مسیر اجرای User Mode هستن نه حذف کامل visibility
حالا یک لایه پایینتر
اینجا میرسیم به یکی از مهمترین چیزهایی که باید برای فهم EDR Evasion بلد باشید:
Memory Permissions
هر Memory Region میتونه مجوزهایی مثل این داشته باشه:
R = Read
W = Write
X = Execute
مثلا:
RW
یعنی قابل خوندن و نوشتنه ولی نباید به عنوان کد اجرا بشه
و:
RX
یعنی قابل خوندن و اجراست ولی نوشتن روی اون مجاز نیست
پس RWX چیه؟
RWX
یعنی یک ناحیه همزمان:
داره
این نوع Memory Permission میتونه برای ما جالب باشه چون کدی که همزمان قابل تغییر و اجراست در بعضی سناریوها ریسک بیشتری ایجاد میکنه البته RWX بهتنهایی به معنی بدافزار بودن نیست بعضی نرمافزارها کاملا legitimate هم ممکنه چنین Memory هایی داشته باشن
نکته مهم این قسمت
EDR
ها معمولا فقط به این نگاه نمیکنن که این Memory چه Permission ی داره
بلکه تغییرات Permission و رفتار اطرافش هم اهمیت داره
مثلا از دید تحلیلی
همین زنجیره میتونه مهمتر از دیدن یک Memory Region به تنهایی باشه
زنجیره مون اینه:
@reverseengine
تا اینجا فهمیدیم Direct Syscall یعنی برنامه تلاش میکنه مسیر معمول User-Mode API رو کوتاه تر کنه
حالا سوال:
Indirect Syscall
ایده اصلی
در Direct Syscall اجرای دستور
syscall از کدی انجام میشه که خود برنامه یا یک Stub مشخص فراهم کردهبهصورت مفهومی:
Application
│
▼
Custom Syscall Stub
│
▼
syscall
│
▼
Kernel
اما در Indirect Syscall ایده اینه که اجرای
syscall از یک مسیر/Stub موجود در فضای User Mode انجام بشهApplication
│
▼
Indirect Path
│
▼
Known syscall stub
│
▼
Kernel
هدف مفهومی این تکنیک تغییر شکل Call Stack و مسیر User-Mode execution نسبت به Direct Syscall هست
چرا این موضوع برای EDR مهمه؟
EDR
فقط نمیپرسه:
syscall اتفاق افتاد؟
بلکه میتونه سوالهای بیشتری بپرسه:
چه Process ی syscall رو انجام داده؟
Call Stack چطور بوده؟
Thread از کجا شروع شده؟
Memory مربوط به کجاست؟
قبل و بعد از syscall چه اتفاقی افتاده؟
پس:
Direct Syscall
≠
Invisible
و:
Indirect Syscall
EDR Bypass
این دو بیشتر تغییر مسیر اجرای User Mode هستن نه حذف کامل visibility
حالا یک لایه پایینتر
اینجا میرسیم به یکی از مهمترین چیزهایی که باید برای فهم EDR Evasion بلد باشید:
Memory Permissions
هر Memory Region میتونه مجوزهایی مثل این داشته باشه:
R = Read
W = Write
X = Execute
مثلا:
RW
یعنی قابل خوندن و نوشتنه ولی نباید به عنوان کد اجرا بشه
و:
RX
یعنی قابل خوندن و اجراست ولی نوشتن روی اون مجاز نیست
پس RWX چیه؟
RWX
یعنی یک ناحیه همزمان:
Read + Write + Execute
داره
این نوع Memory Permission میتونه برای ما جالب باشه چون کدی که همزمان قابل تغییر و اجراست در بعضی سناریوها ریسک بیشتری ایجاد میکنه البته RWX بهتنهایی به معنی بدافزار بودن نیست بعضی نرمافزارها کاملا legitimate هم ممکنه چنین Memory هایی داشته باشن
نکته مهم این قسمت
EDR
ها معمولا فقط به این نگاه نمیکنن که این Memory چه Permission ی داره
بلکه تغییرات Permission و رفتار اطرافش هم اهمیت داره
مثلا از دید تحلیلی
Memory Allocation
↓
Write
↓
Permission Change
↓
Execution
همین زنجیره میتونه مهمتر از دیدن یک Memory Region به تنهایی باشه
زنجیره مون اینه:
API Hooking
↓
User Mode / Kernel Mode
↓
Direct Syscall
↓
Indirect Syscall
↓
Memory Permissions
@reverseengine
ReverseEngineering
Direct Syscall vs Indirect Syscall تا اینجا فهمیدیم Direct Syscall یعنی برنامه تلاش میکنه مسیر معمول User-Mode API رو کوتاه تر کنه حالا سوال: Indirect Syscall ایده اصلی در Direct Syscall اجرای دستور syscall از کدی انجام میشه که خود برنامه یا یک Stub…
Direct Syscall vs Indirect Syscall
So far we have understood that Direct Syscall means that the application tries to shorten the usual path of the User-Mode API
Now the question:
Indirect Syscall
The main idea
In Direct Syscall, the execution of the syscall command is done from the code that the application itself or a specific stub provides
Conceptually:
But in Indirect Syscall, the idea is to execute the syscall from a path/stub existing in the User Mode space
The conceptual goal of this technique is to change the Call Stack shape and the User-Mode execution path compared to Direct Syscall
Why is this important for EDR?
EDR
does not just ask:
Did the syscall happen?
It can ask more questions:
What process made the syscall?
What was the call stack like?
Where did the thread start?
Where is the memory?
What happened before and after the syscall?
So:
Direct Syscall
≠
Invisible
And:
These two are more of a redirection of User Mode execution than a complete removal of visibility
Now one layer lower
Here we come to one of the most important things you need to know to understand EDR Evasion:
Memory Permissions
Each Memory Region can have permissions like this:
For example:
RW
It can be read and written, but it should not be executed as code
And:
RX
It can be read and executed, but writing to it is not allowed
So what is RWX?
RWX
It means a simultaneous region:
Read + Write + Execute
This type of Memory Permission can be interesting for us because code that can be modified and executed simultaneously poses a higher risk in some scenarios. Of course, RWX alone does not mean it is malware. Some completely legitimate software may also have such memories.
The important point of this section
EDRs usually do not only look at what permission this memory has.
But the permission changes and the behavior around it are also important.
For example, from an analytical point of view
This chain can be more important than looking at a Memory Region alone.
Our chain is:
@reverseengine
So far we have understood that Direct Syscall means that the application tries to shorten the usual path of the User-Mode API
Now the question:
Indirect Syscall
The main idea
In Direct Syscall, the execution of the syscall command is done from the code that the application itself or a specific stub provides
Conceptually:
Application
│
▼
Custom Syscall Stub
│
▼
syscall
│
▼
Kernel
But in Indirect Syscall, the idea is to execute the syscall from a path/stub existing in the User Mode space
Application
│
▼
Indirect Path
│
▼
Known syscall stub
│
▼
Kernel
The conceptual goal of this technique is to change the Call Stack shape and the User-Mode execution path compared to Direct Syscall
Why is this important for EDR?
EDR
does not just ask:
Did the syscall happen?
It can ask more questions:
What process made the syscall?
What was the call stack like?
Where did the thread start?
Where is the memory?
What happened before and after the syscall?
So:
Direct Syscall
≠
Invisible
And:
Indirect Syscall
EDR Bypass
These two are more of a redirection of User Mode execution than a complete removal of visibility
Now one layer lower
Here we come to one of the most important things you need to know to understand EDR Evasion:
Memory Permissions
Each Memory Region can have permissions like this:
R = Read
W = Write
X = Execute
For example:
RW
It can be read and written, but it should not be executed as code
And:
RX
It can be read and executed, but writing to it is not allowed
So what is RWX?
RWX
It means a simultaneous region:
Read + Write + Execute
This type of Memory Permission can be interesting for us because code that can be modified and executed simultaneously poses a higher risk in some scenarios. Of course, RWX alone does not mean it is malware. Some completely legitimate software may also have such memories.
The important point of this section
EDRs usually do not only look at what permission this memory has.
But the permission changes and the behavior around it are also important.
For example, from an analytical point of view
Memory Allocation
↓
Write
↓
Permission Change
↓
Execution
This chain can be more important than looking at a Memory Region alone.
Our chain is:
API Hooking
↓
User Mode / Kernel Mode
↓
Direct Syscall
↓
Indirect Syscall
↓
Memory Permissions
@reverseengine
بخش بیست و شیشم بافر اورفلو
Valgrind
چیه و چه فرقی با ASan داره
توی قسمت قبل با ASan آشنا شدیم
اینجا میخوایم بریم سراغ Valgrind و ببینیم چطور میتونه Memory Bugها رو پیدا کنه
بعد هم خیلی ساده ASan و Valgrind رو با هم مقایسه میکنیم
Valgrind
برنامه رو زیر نظر میگیره و دسترسی های حافظه رو بررسی میکنه
مثلا میتونه مواردی مثل اینا رو پیدا کنه
یک مثال ساده:
فایل
C
اینجا فقط برای 4 عدد حافظه گرفتیم
ولی داریم عضو شماره 5 رو مینویسیم
پس یک Out of Bounds Write داریم
کامپایل
shell
بعد با Valgrind اجراش میکنیم
shell
Valgrind
گزارش میده که برنامه یک دسترسی غیرمجاز به حافظه داشته
قسمت مهمش برای Reverse Engineering
فرض کنید یک برنامه پیچیده دارید
برنامه کرش میکنه ولی هنوز نمیدونید مشکل دقیقا کجاست
Valgrind
میتونه Stack Trace و اطلاعات مربوط به دسترسی اشتباه رو نشون بده
بعد
میتونید همون تابع رو داخل Ghidra یا IDA باز کنید و Assembly اون قسمت رو بررسی کنید
یعنی دوباره این مسیر رو داریم
برنامه
در مقابل ASan
خیلی ساده بخوایم بگیم
ASan
معمولا سریع تره و برای Fuzzing و تستهای مداوم خیلی کاربردیه
Valgrind
نیازی به کامپایل با ASan نداره و ابزارهای مختلفی برای تحلیل رفتار برنامه در اختیارمون میذاره البته Valgrind معمولا سربار اجرایی بیشتری داره
یک نکته مهم:
Valgrind
و ASan جای Reverse Engineering رو نمیگیرن
اونا فقط کمک میکنن سریع تر بفهمیم
کجا باید دنبال مشکل بگردیم
بعد کار اصلی ما شروع میشه
یعنی رفتن داخل Ghidra یا IDA و فهمیدن اینکه چرا این Memory Bug اتفاق افتاده
تا اینجا سه ابزار مهم رو داریم
Fuzzer
↓
Crash پیدا میکنه
ASan / Valgrind
↓
Memory Bug رو تحلیل میکنن
Ghidra / IDA
↓
علت Bug رو از روی Binary بررسی میکنیم این دقیقا همون ترکیبیه که یک Reverse Engineer برای تحلیل Memory Bug ها باید کم کم بهش مسلط بشه
@reverseengine
Valgrind
چیه و چه فرقی با ASan داره
توی قسمت قبل با ASan آشنا شدیم
اینجا میخوایم بریم سراغ Valgrind و ببینیم چطور میتونه Memory Bugها رو پیدا کنه
بعد هم خیلی ساده ASan و Valgrind رو با هم مقایسه میکنیم
Valgrind
برنامه رو زیر نظر میگیره و دسترسی های حافظه رو بررسی میکنه
مثلا میتونه مواردی مثل اینا رو پیدا کنه
Invalid Readاستفاده نادرست از حافظه
Invalid Write
Use After Free
Memory Leak
یک مثال ساده:
فایل
C
#include <stdio.h>
#include <stdlib.h>
int main()
{
int *data = malloc(4 * sizeof(int));
data[5] = 100;
free(data);
return 0;
}
اینجا فقط برای 4 عدد حافظه گرفتیم
ولی داریم عضو شماره 5 رو مینویسیم
پس یک Out of Bounds Write داریم
کامپایل
shell
gcc -g file22_demo.c -o file22_demo
بعد با Valgrind اجراش میکنیم
shell
valgrind ./file22_demo
Valgrind
گزارش میده که برنامه یک دسترسی غیرمجاز به حافظه داشته
قسمت مهمش برای Reverse Engineering
فرض کنید یک برنامه پیچیده دارید
برنامه کرش میکنه ولی هنوز نمیدونید مشکل دقیقا کجاست
Valgrind
میتونه Stack Trace و اطلاعات مربوط به دسترسی اشتباه رو نشون بده
بعد
میتونید همون تابع رو داخل Ghidra یا IDA باز کنید و Assembly اون قسمت رو بررسی کنید
یعنی دوباره این مسیر رو داریم
برنامه
↓Valgrind
Valgrind
↓
Memory Error
↓
Stack Trace
↓
Ghidra / IDA
↓
Assembly Analysis
در مقابل ASan
خیلی ساده بخوایم بگیم
ASan
معمولا سریع تره و برای Fuzzing و تستهای مداوم خیلی کاربردیه
Valgrind
نیازی به کامپایل با ASan نداره و ابزارهای مختلفی برای تحلیل رفتار برنامه در اختیارمون میذاره البته Valgrind معمولا سربار اجرایی بیشتری داره
یک نکته مهم:
Valgrind
و ASan جای Reverse Engineering رو نمیگیرن
اونا فقط کمک میکنن سریع تر بفهمیم
کجا باید دنبال مشکل بگردیم
بعد کار اصلی ما شروع میشه
یعنی رفتن داخل Ghidra یا IDA و فهمیدن اینکه چرا این Memory Bug اتفاق افتاده
تا اینجا سه ابزار مهم رو داریم
Fuzzer
↓
Crash پیدا میکنه
ASan / Valgrind
↓
Memory Bug رو تحلیل میکنن
Ghidra / IDA
↓
علت Bug رو از روی Binary بررسی میکنیم این دقیقا همون ترکیبیه که یک Reverse Engineer برای تحلیل Memory Bug ها باید کم کم بهش مسلط بشه
@reverseengine
❤2
Part 26 Buffer Overflow
What is Valgrind and how is it different from ASan?
We met ASan in the previous section.
Here we want to go to Valgrind and see how it can find Memory Bugs.
Then we will compare ASan and Valgrind very simply.
Valgrind
monitors the program and checks memory accesses.
For example, it can find things like these:
Invalid Read
Invalid Write
Use After Free
Memory Leak
Incorrect memory usage
A simple example:
C file
#include <stdio.h>
#include <stdlib.h>
int main()
{
int *data = malloc(4 * sizeof(int));
data[5] = 100;
free(data);
return 0;
}
Here we only got 4 memory slots
But we are writing member number 5
So we have an Out of Bounds Write
Compile
shell
gcc -g file22_demo.c -o file22_demo
Then we run it with Valgrind
shell
valgrind ./file22_demo
Valgrind
reports that the program has an illegal memory access
The important part is for Reverse Engineering
Suppose you have a complex program
The program crashes but you still don't know exactly where the problem is
Valgrind
can show Stack Trace and information about the incorrect access
Then
You can open the same function in Ghidra or IDA and check the Assembly of that part
That means we have this path again
Program
↓
Valgrind
↓
Memory Error
↓
Stack Trace
↓
Ghidra / IDA
↓
Assembly Analysis
Valgrind vs. ASan
To put it simply
ASan
is usually faster and for Fuzzing and continuous testing are very useful
Valgrind
Does not require compilation with ASan and provides us with various tools to analyze the behavior of the program, of course Valgrind usually has more execution overhead
An important point:
Valgrind
and ASan do not replace Reverse Engineering
They only help us understand faster
Where to look for the problem
Then our main work begins
That is, going into Ghidra or IDA and understanding why this Memory Bug occurred
So far we have three important
tools
Fuzzer
↓
Finds a crash
ASan / Valgrind
↓
Analyzes the Memory Bug
Ghidra / IDA
↓
We examine the cause of the Bug from the Binary This is exactly the combination that a Reverse Engineer should gradually master to analyze Memory Bugs
@reverseengine