ReverseEngineering
1.32K subscribers
50 photos
11 videos
106 files
888 links
Download Telegram
Fork
چطور یه Process جدید ساخته میشه؟
تا اینجا درباره Process API حرف زدیم و گفتیم برنامه‌ ها چطوری میتونن با Process ها کار کنن

یکی از مهم‌ ترین چیز هایی که اینجا باهاش سروکار داریم fork() هست

ولی اصلا fork() چیکار میکنه؟

خیلی ساده بخوایم بگیم:

fork()
از Process فعلی یه Process جدید به اسم Child میسازه


یعنی قبل از fork() فقط یه Process داریم:
Parent
بعد از fork():
Parent

+

Child
حالا دو تا Process داریم که هر دو از همون جایی که fork() اجرا شده به اجرای برنامه ادامه میدن
اما یه نکته مهم این وسط هست

Child
صرفا یه کپی ساده از فایل برنامه نیست

سیستم‌ عامل برای Child یه Process مستقل ایجاد میکنه
Child
معمولا:

فضای آدرس مجازی خودش رو داره

PID
متفاوتی داره
وضعیت اجرای خودش رو داره
و منابع قابل مدیریت خودش رو داره
ولی در لحظه‌ای که ساخته میشه وضعیت حافظه و اجرای اون خیلی شبیه Parent هست

پس چرا میگیم Child شبیه کپی Parent هست

چون وقتی Child ساخته میشه سیستم‌عامل کاری می‌کنه که Child تقریبا همون وضعیت Parent رو داشته باشه
مثلا فرض کنید Parent قبل از fork() این متغیر رو داشته باشه:


int x = 10;

Child
هم در فضای آدرس خودش مقدار مشابه ای برای x داره
ولی این به این معنی نیست که Parent و Child دارن از یه متغیر معمولی مشترک استفاده میکنن

هر Process فضای حافظه مجازی خودش رو داره

یه ویژگی خیلی جالب fork()

fork()
توی Parent و Child مقدار برگشتی یکسانی نداره

توی Parent مقدار برگشتی معمولا PID مربوط به Child

توی Child مقدار برگشتی 0 هست
اگر هم مشکلی موقع ساخت Child پیش بیاد، Parent یه مقدار منفی دریافت میکنه

پس برنامه میتونه بفهمه:

من Parent ام؟

یا Child؟

و بر اساس اون مسیر متفاوتی رو اجرا کنه

مثلا:

pid = fork();

if (pid == 0)
// Child

else
// Parent

بعد از اجرای fork() هر دو Process از همون نقطه به اجرای برنامه ادامه میدن
ولی مقدار pid بهشون میگه که الان داخل Parent هستن یا Child

حالا یه سوال مهم‌تر:

اگه Child تقریبا شبیه Parent ساخته شده چطوری میتونه یه برنامه کاملا
متفاوت رو اجرا کنه؟

اینجاست که exec() وارد داستان میشه
معمولا توی سیستم‌های Unix یه الگوی خیلی معروف داریم:
Parent
↓
fork()
↓
Child
↓
exec()
↓
New Program
یعنی:
fork() → ساختن Child

و بعد:

exec() → جایگزین کردن برنامه داخل
Child با برنامه‌ای که میخوایم اجرا بشه

مثلا یه Shell رو در نظر بگیرید

Shell
میتونه یه Child بسازه و بعد Child رو با exec() تبدیل کنه به برنامه‌ای که کاربر درخواست کرده

حالا این موضوع چرا برای مهندسی معکوس مهمه؟

وقتی توی یه برنامه لینوکسی به fork() برخورد میکنید باید حواستون باشه که از اینجا به بعد دیگه فقط با یه مسیر اجرای برنامه طرف نیستید

یه Process جدید وارد ماجرا شده

مثلا:
Parent

│

├── fork()

│

├────────► Child

│

▼

ادامه Parent
یعنی ممکنه رفتار برنامه بین Parent و Child تقسیم شده باشه پس موقع تحلیل یه برنامه اگه به fork() رسیدید حتما باید این احتمال رو در نظر بگیرید که از این نقطه به بعد دو تا Process جداگانه دارید که ممکنه هرکدوم رفتار متفاوتی داشته باشن
این موضوع بعدا توی تحلیل Process ها Debugging و بررسی رفتار برنامه‌ها خیلی به دردتون میخوره


fork()
یه Child Process جدید میسازه

Parent
و Child هر دو به اجرای برنامه ادامه میدن
Child
یه PID متفاوت داره
هر Process فضای آدرس مجازی خودش رو داره مقدار برگشتی fork() کمک میکنه بفهمیم داخل Parent هستیم یا Child و معمولا fork() رو در کنار exec() میبینیم

و اینجا یکی از ایده‌های مهم رو خیلی خوب میشه دید:

سیستم‌عامل فقط برنامه‌ها رو اجرا نمیکنه بلکه یه محیط میسازه که چند تا Process بتونن مستقل از هم اجرا بشن

@reverseengine
ReverseEngineering
Fork چطور یه Process جدید ساخته میشه؟ تا اینجا درباره Process API حرف زدیم و گفتیم برنامه‌ ها چطوری میتونن با Process ها کار کنن یکی از مهم‌ ترین چیز هایی که اینجا باهاش سروکار داریم fork() هست ولی اصلا fork() چیکار میکنه؟ خیلی ساده بخوایم بگیم: fork()…
Fork

How do you create a new Process?
So far we've talked about the Process API and how programs can work with Processes

One of the most important things we're dealing with here is fork()

But what does fork() do?

To put it very simply:

fork()
creates a new Process called Child from the current Process

That is, before fork() we only have one Process:
Parent

After fork():
Parent

+
Child
Now we have two Processes, both of which continue to execute the program from the same place where fork() was executed

But there is an important point here

Child
is not just a simple copy of the program file

The operating system creates an independent Process for Child

Child
usually:

It has its own virtual address space

It has a different PID

It has its own execution state

And it has its own manageable resources

But at the moment it is created, its memory and execution state are very similar to Parent

So why do we say that Child is like a copy of Parent

Because when Child is created, the operating system makes Child have almost the same state as Parent

For example, suppose Parent has this variable before fork():

int x = 10;

Child
also has the same value for x in its address space
But this does not mean that Parent and Child are using a common common variable
Each Process has its own virtual memory space
A very interesting feature of fork()

fork()
does not have the same return value in Parent and Child

In Parent the return value is usually the PID of Child

In Child the return value is 0
If there is a problem while creating Child, Parent gets a negative value

So the program can understand:

Am I Parent?

Or Child?

And based on that execute a different path

For example:

pid = fork();

if (pid == 0) // Child

else // Parent

After executing fork() both Processes continue executing the program from the same point
But the pid value tells them whether they are now inside Parent or Child

Now a more important question:

If Child is created almost exactly like Parent, how can it execute a completely different program?

This is where exec() comes into play.
Usually, in Unix systems, we have a very famous pattern:
Parent
↓
fork()
↓
Child
↓
exec()
↓
New Program

That is:
fork() → create Child

And then:

exec() → replace the program inside
Child with the program we want to run
For example, consider a Shell

The Shell
can create a Child and then use exec() to convert the Child into the program requested by the user

Now why is this important for reverse engineering?

When you encounter fork() in a Linux program, you should be aware that from here on, you are no longer dealing with just one path of program execution

A new process has entered the story

For example:
Parent

│

├── fork()

│

├───────► Child

│

▼

Continuation of Parent
This means that the behavior of the program may be divided between Parent and Child
So when analyzing a program, if you reach fork(), you must definitely consider the possibility that from this point on, you have two separate processes, each of which may have different behavior. This will be very useful later in analyzing processes, debugging, and examining program behavior

fork()
creates a new Child Process

Parent
and Child
both continue executing the program
Child
has a different PID
Each Process has its own virtual address space
The return value of fork() helps us know whether we are inside Parent or Child
And we usually see fork() next to exec()

And here one of the important ideas is very important It's easy to see:

The operating system doesn't just run programs, it creates an environment where multiple processes can run independently

@reverseengine
User Mode
در برابر Kernel Mode

برای اینکه بفهمید چرا تکنیک‌ هایی مثل Direct Syscall اصلا تعریف شدن اول باید بدونید ویندوز دو تعریف مهم داره:

┌─────────────────────────┐
│ User Mode │
│ Applications / DLLs │
└───────────┬─────────────┘
│
System Call
│
▼
┌─────────────────────────┐
│ Kernel Mode │
│ Windows Kernel / Drivers│
└─────────────────────────┘

User Mode

برنامه‌های معمولی اینجا اجرا میشن:

Browser
PowerShell
Game
Your Program
↓
kernel32.dll
↓
ntdll.dll
↓
System Call

EDR
ها میتونن در این لایه telemetry جمع‌ آوری کنن و رفتار برنامه رو بررسی کنن




Kernel Mode

اینجا بخش‌ های حساس سیستم‌ عامل قرار دارن

مثلا مدیریت:

Processes

Threads

Memory

Drivers

File System

Network


به همین دلیل اگر فقط یک لایه از User Mode رو دور بزنید به معنی نامرئی شدن نیست.


-

Direct Syscall
چه مفهومی داره؟

ایده اصلی اینه که به‌جای طی کردن مسیر معمول User-Mode API برنامه مستقیما به مرز System Call نزدیک بشید


مفهوم:

Normal:

Application
↓
Win32 API
↓
ntdll
↓
System Call
↓
Kernel


Direct Syscall:

Application
↓
System Call
↓
Kernel

اما نکته مهم همینجاست:

Direct Syscall
به معنی دور زدن کامل EDR نیست

چون Kernel و سایر منابع telemetry همچنان میتونن رفتار اتفاق‌ افتاده رو ببینن




پس چرا هکر ها بهش علاقه دارن؟

چون اگر یک مکانیزم دفاعی مشخص در User Mode قرار گرفته باشه تغییر مسیر اجرای برنامه میتونه روی همان مکانیزم اثر بذاره

ولی EDR مدرن فقط به یک نقطه وابسته نیست:

┌── User Mode
│
Process ─────┼── Memory
│
├── ETW / Telemetry
│
├── Kernel
│
└── Network

بنابراین Evasion واقعی یک بازی چند لایه ست نه پیدا کردن یک API عجیب و غریب

نکته‌ای که باید یاد بگیرید

اگر بخواید AV/EDR Evasion رو واقعا بفهمید نباید از حفظ کردن تکنیک‌ها شروع کنید

باید بفهمید:

دفاع کجاست
چه چیزی رو میتونه ببینه
چه telemetry دریافت میکنه
مهاجم تلاش میکنه کدوم visibility رو کاهش بده





User Mode vs Kernel Mode

To understand why techniques like Direct Syscall were defined at all, you first need to know that Windows has two important definitions:

┌───────────────────────────┐
│ User Mode │
│ Applications / DLLs │
└─
│
System Call
│
▼
┌────────────� Management:

Processes

Threads

Memory

Drivers

File System

Network

That's why bypassing just one layer of User Mode doesn't mean you'll be invisible.



Direct Syscall
What does it mean?

The main idea is to approach the System Call boundary directly instead of going through the usual User-Mode API path.

Conceptual:

Normal:

Application
↓
Win32 API
↓
ntdll
↓
System Call
↓
Kernel

Direct Syscall:

Application
↓
System Call
↓
Kernel

But here's the important point:

Direct Syscall
does not mean bypassing EDR completely

Because the Kernel and other telemetry sources can still see the behavior that happened

So why are hackers interested in it?

Because if a specific defense mechanism is in User Mode, rerouting the program can affect that mechanism

But modern EDR is not just about one point:

┌── User Mode
│
Process ──────┼── Memory
│
├── ETW / Telemetry
│
├── Kernel
│
└── Network

So real evasion is a multi-layered game, not about finding a weird API

What you need to learn

If you really want to understand AV/EDR evasion, you shouldn't start by memorizing techniques

You need to understand:

Where is the defense
What can it see
What telemetry is it receiving
What visibility is the attacker trying to reduce

@reverseengine
Heap Overflow

یکی از معروف‌ترین باگ‌های Heap:

Heap Overflow
,
ایده‌اش خیلی شبیه Buffer Overflow روی Stack هست با این تفاوت که این بار سر ریز داخل Heap اتفاق میوفته

Heap Overflow
فرض کنید برنامه از Heap یک Chunk با ظرفیت مشخص میگیره:

C
char *buffer = malloc(64);


یعنی برنامه فضایی برای نگهداری داده در اختیار داره حالا اگر برنامه بیشتر از ظرفیتی که برای این Chunk در نظر گرفته شده داخلش بنویسه داده از محدوده خودش خارج میشه

به این اتفاق میگیم:
Heap Overflow
به شکل ساده:

[ Buffer A ][ Buffer B ]

████████████████████
↓
نوشتن بیش از ظرفیت A
↓
[ Buffer A ][AAAAAAA...]
↑
وارد محدوده B


چرا این اتفاق خطرناکه؟

چون Chunk ها معمولا کنار هم قرار میگیرن پس اگر برنامه از محدوده‌ ی Chunk خودش خارج بشه ممکنه داده‌ های مربوط به حافظه‌ ی مجاور رو تغییر بده
اون حافظه‌ ی مجاور میتونه متعلق به:

یک Object دیگه
یک ساختار داده
اطلاعات مدیریتی Heap
یا داده‌های مهم برنامه باشه
در نتیجه یک اشتباه ساده در اندازه‌ ی داده میتونه روی بخش دیگری از برنامه تاثیر بذاره


تفاوت Heap Overflow و Stack Overflow

هر دو از یک ایده‌ی کلی پیروی میکنن:

نوشتن بیشتر از ظرفیت حافظه‌ ای که در اختیار برنامه قرار گرفته

اما محل اتفاق فرق داره

Stack Overflow:

Stack
↓
Buffer
↓
داده بیشتر از ظرفیت
↓
اطلاعات اطراف Buffer

Heap Overflow:

Heap
↓
Chunk
↓
داده بیشتر از ظرفیت
↓
Chunk یا Metadata مجاور

پس نباید این دو رو یکی بدونیم

چرا Chunkهای مجاور اهمیت دارن؟

فرض کنید Heap این شکلیه:

+----------------+
| Chunk A |
+----------------+
| Chunk B |
+----------------+
| Chunk C |
+----------------+

اگر برنامه داخل Chunk A بیشتر از ظرفیتش بنویسه ممکنه نوشتن وارد محدوده‌ ی Chunk B بشه در نتیجه ممکنه داده‌ی B تغییر کنه بدون اینکه برنامه مستقیما قصد تغییر B رو داشته باشه این دقیقا همون چیزی هست که Heap Overflow رو خطرناک میکنه

Metadata
هم میتونه مهم باشه یادتون هست گفتیم Chunk علاوه بر User Data اطلاعات مدیریتی هم داره اگر یک Overflow از محدوده‌ ی خودش خارج بشه در بعضی شرایط ممکنه به Metadata مربوط به بخش‌های مجاور هم برسه اینجاست که موضوع پیچیده‌ تر میشه چون دیگه فقط داده‌ی برنامه تغییر نکرده ممکنه اطلاعاتی که Allocator برای مدیریت Heap استفاده میکنه هم تحت تاثیر قرار گرفته باشه البته Allocator های مدرن بررسی‌ های مختلفی دارن و خیلی از روش‌ های قدیمی Heap Exploitation دیگه به سادگی گذشته قابل استفاده نیستن

یک نکته خیلی مهم

هر Heap Overflow حتما قابل تبدیل شدن به Exploit نیست
ممکنه فقط باعث بشه:
برنامه Crash کنه
یک مقدار اشتباه تغییر کنه
داده‌ای خراب بشه
یا رفتار برنامه غیرقابل‌ پیش‌بینی بشه

برای اینکه یک Memory Corruption واقعا قابل سواستفاده بشه باید بررسی کنیم چه چیزی قابل تغییره و این تغییر چه اثری روی برنامه داره

پس:
Bug ≠ Exploit


وجود باگ فقط نقطه‌ی شروع تحلیلمونه


Heap Overflow
زمانی اتفاق میوفته که برنامه بیشتر از ظرفیت یک بخش اختصاص‌ یافته روی Heap بنویسه و از محدوده‌ی خودش خارج بشه چون Chunk ها در Heap کنار هم قرار میگیرن این Overflow ممکنه داده‌ها یا ساختارهای مجاور رو تحت تاثیر قرار بده.برای تحلیل Heap Exploitation باید بفهمیم این Overflow دقیقا چه چیزی رو میتونه تغییر بده و اون تغییر چه اثری روی رفتار برنامه داره


@reverseengine
ReverseEngineering
Heap Overflow یکی از معروف‌ترین باگ‌های Heap: Heap Overflow , ایده‌اش خیلی شبیه Buffer Overflow روی Stack هست با این تفاوت که این بار سر ریز داخل Heap اتفاق میوفته Heap Overflow فرض کنید برنامه از Heap یک Chunk با ظرفیت مشخص میگیره: C char *buffer = malloc(64);…
Heap Overflow

One of the most famous Heap bugs:

Heap Overflow

Its idea is very similar to Buffer Overflow on Stack except that this time the overflow occurs inside the Heap

Heap Overflow
Suppose the program gets a Chunk with a certain capacity from the Heap:

C
char *buffer = malloc(64);

That is, the program has space to store data. Now if the program writes more than the capacity intended for this Chunk, the data will go out of its bounds

We call this event:
Heap Overflow
Simply put:

[ Buffer A ][ Buffer B ]

████████████████████
↓
Writing more than the capacity of A
↓
[ Buffer A ][AAAAAAAA...]
↑
Entering B range

Why is this dangerous?

Because chunks are usually placed next to each other, if the program goes out of its own chunk, it may change the data in the adjacent memory. That adjacent memory could belong to:

Another object

A data structure

Heap management information

Or important program data

As a result, a simple mistake in the size of the data can affect another part of the program

Difference between Heap Overflow and Stack Overflow

Both follow the same general idea:

Writing more than the memory capacity provided to the program

But the location of the event is different

Stack Overflow:

Stack
↓
Buffer
↓
Data more than capacity
↓
Information around Buffer

Heap Overflow:

Heap
↓
Chunk
↓
Data more than capacity
↓
Adjacent Chunk or Metadata

So we should not consider these two as the same

Why are adjacent chunks important?

Suppose the Heap looks like this:

+----------------+
| Chunk A |
+----------------+
| Chunk B |
+----------------+
| Chunk C |
+----------------+

If the program writes more than its capacity into Chunk A, the write may enter the Chunk B area, as a result, the data in B may change without the program directly intending to change B. This is exactly what makes Heap Overflow dangerous

Metadata
can also be important. Remember that we said that Chunk has management information in addition to User Data. If an Overflow goes out of its own scope, in some situations it may also reach the Metadata related to adjacent sections. This is where the matter becomes more complicated because not only the program data has changed, the information that the Allocator uses to manage the Heap may also be affected. Of course, modern Allocators have different checks and many of the old Heap Exploitation methods are no longer as simple as they used to be.

A very important point

Not every Heap Overflow can be turned into an Exploit. It may only cause:
The program to crash
A wrong value to change
Data to be corrupted
Or the program to behave unpredictable

For a Memory Corruption to really To be exploitable, we need to examine what can be changed and what effect this change has on the program

So:

Bug ≠ Exploit


The existence of a bug is only the starting point for our analysis

Heap Overflow

Occurs when a program writes more than the capacity of an allocated section on the Heap and goes out of its bounds because the Chunks in the Heap are placed next to each other. This Overflow may affect adjacent data or structures. To analyze Heap Exploitation, we need to understand what exactly this Overflow can change and what effect that change has on the behavior of the program.

@reverseengine
Forwarded from club1337
Devman-ArticleXakep.txt
17.9 KB
Вымогатель-болтун. Как Devman прошел путь от новичка до преступника в розыске Интерпола

👑 Статья для подписчиков

31 июля 2025 года Джон Ди Маджо открыл сообщение в зашифрованном мессенджере. Преступники обычно не любят, когда их деятельность расследуют, но этот написал сам. К сообщению была приложена фотография: дорогие часы, спортивные автомобили. Отправителя Ди Маджо знал.

https://xakep.ru/2026/08/13/devman/

Telegram ✉️ @club1337
X (Twitter) 🕊 @club31337
Please open Telegram to view this post
VIEW IN TELEGRAM
String Obfuscation
وقتی رشته‌ها هم مخفی میشن

تا اینجا درباره Control Flow Flattening و Opaque Predicate صحبت کردیم

حالا بریم سراغ یکی از چیزهایی که توی تحلیل استاتیک خیلی زود باهاش برخورد میکنید

رشته‌ها

وقتی یک برنامه رو با IDA یا Ghidra باز میکنیم معمولا یکی از اولین کارها اینه که Strings رو بررسی کنیم

چون رشته‌ هایی مثل این خیلی اطلاعات میدن:

Login failed
Access denied
https://example.com
config.json
username

اما Obfuscation میتونه همین رشته‌ها رو هم از حالت واضح خارج کنه

مثلا به جای اینکه داخل باینری داشته باشیم:

Access denied

ممکنه فقط یک سری بایت ببینیم:

12 37 21 04 55 19 ...

و برنامه موقع اجرا اونها رو به رشته اصلی تبدیل کنه

یک روش ساده برای این کار XOR هست

مثلا رشته اصلی:

HELLO

با یک کلید مشخص XOR میشه و نتیجه داخل فایل قرار میگیره

وقتی برنامه اجرا میشه دوباره همون عملیات انجام میشه و رشته اصلی برمیگرده

در نتیجه اگر فقط Strings رو روی فایل اجرا کنیم ممکنه اصلا HELLO رو نبینیم

اما اینجا یک نکته خیلی مهم وجود داره

هدف ما این نیست که فقط دنبال رشته خوندن بگردیم

باید بفهمهیم رشته کجا ساخته میشه

مثلا ممکنه داخل دیس‌اسمبل ببینیم:

داده رمزگذاری‌شده
↓
Decode
↓
Buffer
↓
استفاده توسط برنامه

اگر تابع Decode رو پیدا کنیم میتونیم بفهمیم برنامه چطور رشته‌ها رو در زمان اجرا بازسازی میکنه

بعضی برنامه‌ها حتی رشته‌ها رو از اول به صورت کامل در حافظه نگه نمیدارن

ممکنه فقط زمانی که لازم دارن رشته رو Decode کنن و بعد دوباره پاکش کنن

برای همین تحلیل داینامیک اینجا خیلی کمک میکنه

می‌تونیم ببینیم چه زمانی Buffer ساخته میشه و چه زمانی محتوای قابل خوندن داخلش قرار میگیره

یک نکته مهم دیگه هم اینه که هر داده‌ای که شبیه متن نیست لزوما رمزگذاری نشده

ممکنه فشرده شده باشه ممکنه یک ساختار باینری باشه یا حتی فقط داده‌ای باشه که با یک Encoding خاص ذخیره شده

پس قبل از اینکه بگیم این رشته رمزگذاری شده باید مسیر استفاده از اون داده رو بررسی کنیم

تمرین:

یک برنامه ساده بسازید که یک رشته مشخص داشته باشه

بعد رشته رو با یک XOR ساده قبل از ذخیره شدن تغییر بدید و در زمان اجرا دوباره Decode کنید

حالا برنامه رو داخل Ghidra باز کنید

اول Strings رو بررسی کنید

بعد تابعی که داده رو Decode میکنه پیدا کنید

در اخر سعی کنید بدون اجرای برنامه الگوریتم Decode رو از روی اسمبلی بازسازی کنید

اینجا یک چیز مهم به دست میارید:

به جای اینکه فقط دنبال چیزی که برنامه نشون میده بگردید یاد میگیرید بفهمید اون داده چطور ساخته شده

@reverseengine
ReverseEngineering
String Obfuscation وقتی رشته‌ها هم مخفی میشن تا اینجا درباره Control Flow Flattening و Opaque Predicate صحبت کردیم حالا بریم سراغ یکی از چیزهایی که توی تحلیل استاتیک خیلی زود باهاش برخورد میکنید رشته‌ها وقتی یک برنامه رو با IDA یا Ghidra باز میکنیم معمولا…
String Obfuscation
When strings are also hidden

So far we have talked about Control Flow Flattening and Opaque Predicate

Now let's move on to one of the things you will encounter very soon in static analysis

Strings

When we open a program with IDA or Ghidra, one of the first things we usually do is check the Strings

Because strings like this give a lot of information:

Login failed

Access denied

https://example.com

config.json

username

But Obfuscation can also remove these strings from the clear state

For example, instead of having inside the binary:

Access denied

We may only see a series of bytes:

12 37 21 04 55 19 ...

And the program converts them into the original string when it runs

A simple way to do this is XOR

For example, the original string:

HELLO

It is XORed with a specific key and the result is placed in the file takes

When the program is run, the same operation is performed again and the original string is returned

As a result, if we just run Strings on the file, we may not see HELLO at all

But there is a very important point here

Our goal is not to just look for the string to read

We need to understand where the string is created

For example, we may see in the disassembler:

Encrypted data
↓
Decode
↓
Buffer
↓
Used by the program

If we find the Decode function, we can understand how the program reconstructs the strings at runtime

Some programs do not even keep the strings completely in memory from the beginning

They may only need to decode the string and then clear it again

That is why dynamic analysis is very helpful here

We can see when the Buffer is created and when the readable content is placed in it

Another important point is that any data that does not look like text is necessarily encrypted Not

It could be compressed, it could be a binary structure, or even just data stored with a specific encoding

So before we say this string is encrypted, we need to look at how that data is used

Exercise:

Write a simple program that takes a given string

Then modify the string with a simple XOR before saving it and decode it again at runtime

Now open the program in Ghidra

First examine the Strings

Then find the function that decodes the data

Finally, try to recreate the Decode algorithm from assembly without running the program

Here you will learn something important:

Instead of just looking for what the program shows, you will learn to understand how that data is constructed

@reverseengine
Constant Obfuscation وقتی حتی عدد ها هم مخفی میشن

تا اینجا دیدیم که چطور رشته‌ها رو میشه مخفی کردولی فقط رشته‌ها نیستن که Obfuscate میشن عددها و ثابت‌ های برنامه هم میتونن مخفی بشن
فرض کنید برنامه باید مقدار 100 رو استفاده کنه در حالت عادی ممکنه توی اسمبلی چیزی شبیه این ببینید:

mov eax, 100


حالا برنامه‌ نویس یا ابزار Obfuscation میتونه همون مقدار رو به شکل پیچیده‌ تری تولید کنه

مثلا:

mov eax, 73
add eax, 27


در نهایت مقدار eax میشه 100

اما یک قدم جلوتر:

mov eax, 0x12345678
xor eax, 0x12345614


نتیجه XOR دوباره میتونه یک مقدار مشخص باشه در این حالت وقتی فقط یک دستور رو میبینید مقدار واقعی ثابت فورا مشخص نمیشه بعصی وقتا حتی محاسبات بیشتری استفاده میشه:

XOR
ADD
SUB
ROL
ROR


و همه اینها فقط برای تولید یک مقدار ثابت انجام میشن مثلا ممکنه برنامه برای ساختن یک عدد ساده چند تا دستور اجرا کنه
اینجا کاری که ما انجام میدیم اینه که به جای نگاه کردن به تک‌ تک دستورها مسیر تولید مقدار رو دنبال میکنیم

یعنی میپرسیم:

این مقدار از کجا اومد؟
چه عملیاتی روش انجام شده؟
در نهایت کجا استفاده شده؟

مثلا:

مقدار اولیه

↓
XOR
↓
ADD
↓
SUB
↓
مقدار نهایی

اگر بتونید این زنجیره رو ساده کنید مقدار واقعی ثابت دوباره مشخص میشه
این کار مخصوصا وقتی مهم میشه که ثابت‌ها بخشی از یک الگوریتم باشن
مثلا یک برنامه ممکنه یک مقدار ثابت رو برای مقایسه محاسبه یا ساختن یک جدول استفاده کنه اگر مقدار اصلی مخفی شده باشه فهمیدن الگوریتم هم سخت‌ تر میشه
یک نکته جالب اینه که بعضی وقتا Decompiler خودش میتونه این محاسبات رو ساده کنه

مثلا چند دستور اسمبلی رو تبدیل کنه به:


x = 100;


اما همیشه نباید به خروجی Decompiler اعتماد کرد گاهی Obfuscation باعث میشه خروجی چیزی کاملا پیچیده و گمراه‌کننده باشه

پس یکی از مهارت‌های مهم Reverse Engineer اینه که بتونه بین سه چیز حرکت کنه:

Assembly
Decompiler
منطق واقعی برنامه


تمرین:

این عبارت رو بدون اجرای برنامه ساده کنید:

int x = 73;
x = x + 27;
x = x ^ 0;


بعد یک مثال پیچیده‌ تر برای خودتون بسازید که در نهایت به یک عدد مشخص برسه بعد همون برنامه رو Compile کنید و داخل Ghidra ببینید Compiler چه شکلی از اون ساخته

هدف تمرین این نیست که فقط جواب عدد رو پیدا کنید هدف اینه که یاد بگیرید یک مقدار رو از مسیر محاسباتش دنبال کنید

@reverseengine
ReverseEngineering
Constant Obfuscation وقتی حتی عدد ها هم مخفی میشن تا اینجا دیدیم که چطور رشته‌ها رو میشه مخفی کردولی فقط رشته‌ها نیستن که Obfuscate میشن عددها و ثابت‌ های برنامه هم میتونن مخفی بشن فرض کنید برنامه باید مقدار 100 رو استفاده کنه در حالت عادی ممکنه توی اسمبلی…
Constant Obfuscation When Even Numbers Are Hiding

So far we have seen how strings can be hidden, but it is not only strings that are obfuscated, numbers and program constants can also be hidden
Suppose the program needs to use the value 100, in normal case you might see something like this in the assembly:

mov eax, 100


Now the programmer or the Obfuscation tool can generate the same value in a more complex way

For example:

mov eax, 73

add eax, 27


Finally the value of eax becomes 100

But one step further:

mov eax, 0x12345678

xor eax, 0x12345614


The result of XOR can again be a specific value. In this case, when you see only one instruction, the actual value of the constant is not immediately clear. Sometimes even more calculations are used:

XOR
ADD
SUB
ROL
ROR


And all this
They are only used to produce a constant value. For example, a program might execute several instructions to generate a simple number. What we do here is that instead of looking at each instruction individually, we follow the path of the value. That is, we ask:

Where did this value come from?

What operation did the method perform?

Where was it finally used?

For example:

Initial value

↓
XOR
↓
ADD
↓
SUB
↓
Final value


If you can simplify this chain, the actual value of the constant is revealed again. This is especially important when constants are part of an algorithm. For example, a program might use a constant value to compare calculations or build a table. If the original value is hidden, it becomes harder to understand the algorithm. One interesting thing is that sometimes the decompiler itself can simplify these calculations.

For example, it can convert a few assembly instructions to:

x = 100;


But you shouldn't always trust the output of the Decompiler. Sometimes Obfuscation can make the output of something quite complex and misleading.

So one of the important skills of a Reverse Engineer is to be able to move between three things:

Assembly
Decompiler
The actual logic of the program

Exercise:

Simplify this expression without running the program:

int x = 73;

x = x + 27;

x = x ^ 0;


Then create a more complex example for yourself that will eventually reach a specific number. Then compile the same program and see what the Compiler makes of it in Ghidra.

The goal of the exercise is not to just find the answer to the number. The goal is to learn to follow a value through its calculation path.

@reverseengine