ReverseEngineering
1.32K subscribers
50 photos
11 videos
106 files
888 links
Download Telegram
ReverseEngineering
Instruction Substitution وقتی یک دستور رو با چند دستور دیگه عوض میکنیم تا اینجا درباره روش‌های مختلفی صحبت کردیم که یک باینری میتونه طوری ساخته بشه که تحلیلش سخت‌ تر بشه حالا بریم سراغ یکی از تکنیک‌های ساده ولی خیلی کاربردی: Instruction Substitution ایده‌اش…
Instruction Substitution

When we replace an instruction with several other instructions

So far we have talked about different ways that a binary can be made to be more difficult to analyze

Now let's move on to one of the simple but very useful techniques:

Instruction Substitution

The idea is very simple:

Instead of using a specific instruction, we make several other instructions produce exactly the same result

That is, for example, if we can perform an operation with one instruction, we deliberately replace it with several instructions

A simple example:

Suppose we want to zero a register in Assembly

The usual method:

xor eax, eax


After execution:

EAX = 0


But we can create the same result in another way:

sub eax, eax


Again:

EAX = 0


In terms of the final result, both are the same

But in terms of the code format, they are different

Another example

Suppose:

mov eax, 0


Can also be zeroed in other ways:

xor eax, eax


or:

sub eax, eax


So a single logic can have multiple different representations

This is exactly what Instruction Substitution is counting on

Why is this important?

Here is an important point
If an Analyst or Analysis Tool expects to see an operation in a specific form, Instruction Substitution can change that pattern
For example, the Analyst might look for:

xor eax, eax


in a binary

but the Binary Builder instead writes:

sub eax, eax


The result is the same but the pattern has changed

A slightly more interesting example

Suppose we want to:

add eax, 1


We can use:

inc eax


Both increase the value of EAX by one

That is:

Before:

EAX = 10


After:

EAX = 11


But Assembly is different

Now a more important point:

Instruction Substitution

is not just for a simple instruction

Sometimes we can implement an operation with multiple instructions

For example:

add eax, 5


We can do it in a different way Let's write it:

push 5
pop ecx
add eax, ecx


They can be equivalent in terms of the result, although this example is intentionally simple and ineffective.

The goal here is not speed.

The goal is to:

Have several different representations of an operation.

Why is it important?

Because you shouldn't just look at the form of the instruction.

You should ask:

What do these instructions do in the end?

For example, if we see:

xor eax, eax


we say very quickly:

EAX = 0


But if we see:

sub eax, eax


we still have to recognize the same concept.

So in Reverse Engineering, we should gradually move from:

Instruction Thinking

to:

Semantic Thinking

That is, instead of just saying:

What is this instruction?

we should ask:

What is the logical result of this part of the program?

This is also important in Decompiler

Decompilers try to convert Assembly back to something like high-level code

For example:

xor eax, eax


may be converted to something like:

eax = 0


If the same thing is done with several different instructions, the Decompiler needs to understand what this set means in the end

So:

Assembly
↓
Instructions
↓
Semantic Analysis
↓


High-Level Representation
One of the important parts of Reverse Engineering is this semantic recognition

But one very important point

Instruction Substitution
is not always completely equivalent
Two instructions may be the same in terms of final value but have different effects on:

Flags
Registers
Memory
Exception behavior
Timing


For example:

inc eax


and:

add eax, 1


are similar in terms of EAX value but their Flag behavior is not exactly the same

So when we say:

two instructions are equivalent Are

We need to specify:

In what sense?

Only the result of the register?

Or the entire processor state?

This difference is very important in Reverse Engineering

Instruction Substitution
In Obfuscation

Now let's return to the main discussion
Suppose a program performs this operation:

xor eax, eax


In the normal case, it is very easy to recognize
But if the program creator uses different methods to produce the same result, the Assembly pattern changes

For example:

sub eax, eax
ReverseEngineering
Instruction Substitution وقتی یک دستور رو با چند دستور دیگه عوض میکنیم تا اینجا درباره روش‌های مختلفی صحبت کردیم که یک باینری میتونه طوری ساخته بشه که تحلیلش سخت‌ تر بشه حالا بریم سراغ یکی از تکنیک‌های ساده ولی خیلی کاربردی: Instruction Substitution ایده‌اش…
Or in appropriate conditions, more complex combinations
Now, if the same thing is repeated in different parts of the program with different forms, pattern-based analysis becomes more difficult

For example:

Operation A
↓
Representation 1

Operation A
↓
Representation 2

Operation A
↓


Representation 3
While:

Semantic = is one

@reverseengine
ReverseEngineering
Instruction Substitution وقتی یک دستور رو با چند دستور دیگه عوض میکنیم تا اینجا درباره روش‌های مختلفی صحبت کردیم که یک باینری میتونه طوری ساخته بشه که تحلیلش سخت‌ تر بشه حالا بریم سراغ یکی از تکنیک‌های ساده ولی خیلی کاربردی: Instruction Substitution ایده‌اش…
نکته طلایی

وقتی با کد Obfuscated مواجه شدید دنبال شکل ظاهری instruction ها نباشید
دنبال اثر اونها روی Machine State باشید

یعنی بررسی کنید:

Registerها چی شدن؟
Memory چه تغییری کردن؟
Flagها چی شدن؟
Control Flow کجا رفت؟

اگر این‌ها رو بفهمید حتی اگر شکل Assembly کاملا عجیب باشه میتونید منطق واقعی برنامه رو استخراج کنید

یک تمرین خیلی خوب

این چهار قطعه رو ببینید:
xor eax, eax

sub eax, eax

mov eax, 0

push 0

pop eax

همه رو فقط از یک زاویه نگاه نکنید

برای هر کدوم بررسی کنید:

EAX چی میشه؟
Flags چی میشن؟
Stack چه تغییری میکنه؟
Memory چه تغییری میکنه؟




Instruction Substitution:

یک عملیات رو با instruction یا instruction های دیگه ای پیاده سازی کنیم بدون اینکه semantic مورد نظر تغییر کنه

هدف میتونه:

تغییر شکل کد
سخت‌تر کردن Pattern Matching

افزایش تنوع instruction ها
Obfuscation
سخت‌تر کردن Static Analysis
باشده

اما برای ما یک درس خیلی مهم داره:

Assembly
رو نباید فقط بر اساس اسم instruction ها تحلیل کنیم باید semantic و اثر واقعی اونها رو روی CPU
بفهمیم

و این دقیقا ساختار ماست:

Assembly Basics
        ↓
Disassembly
        ↓
Instruction Semantics
        ↓
Obfuscation
        ↓
Deobfuscation

@reverseengine
ReverseEngineering
Instruction Substitution When we replace an instruction with several other instructions So far we have talked about different ways that a binary can be made to be more difficult to analyze Now let's move on to one of the simple but very useful techniques:…
Golden Tip

When you encounter Obfuscated code, don't look for the appearance of the instructions
Look for their effect on the Machine State

That is, check:

What happened to the Registers?

What changed in Memory?

What happened to the Flags?

Where did the Control Flow go?

If you understand these, even if the Assembly form is completely strange, you can extract the real logic of the program

A very good exercise

Look at these four pieces:

xor eax, eax


sub eax, eax


mov eax, 0


push 0


pop eax


Don't look at everything from just one angle

For each, check:

What happens to EAX?

What happens to Flags?

What changes to the Stack?

What changes to Memory?

Instruction Substitution:

Implement an operation with another instruction or instructions without changing the intended semantics

The goal can be:

Change the shape of the code

Make Pattern Matching harder

Increase the variety of instructions

Obfuscation

Make Static Analysis harder

But it has a very important lesson for us:

We should not analyze assembly based only on the name of the instructions, we should understand their semantics and their real effect on the CPU

And this is exactly our structure:

Assembly Basics
        ↓
Disassembly
         ↓
Instruction Semantics
         ↓
Obfuscation
         ↓
Deobfuscation


@reverseengine
Parent
چطور منتظر Child میمونه؟

تا اینجا دو تا از مهم‌ ترین قسمت‌ های Process API رو یاد گرفتیم:

fork() → ساخت Child

exec() → اجرای یک برنامه جدید داخل Child

حالا یک سؤال:

Parent
از کجا بفهمد Child کارش تموم شده؟

اینجاست که wait() وارد میشه




wait() چیکار میکنه؟

خیلی ساده:

wait()
باعث میشه Parent منتظر بمونه تا یکی از Child هاش تموم شه

مثلا:

Parent
│
├── fork()
│
▼
Child
│


├── کار خودش رو انجام میده
│
└── تموم میشه
│
▼
Parent ادامه میده

یعنی Parent نمیگه:

من دیگه کاری باهات ندارم



پس میگه:

کارت که تموم شد من ادامه میدم






یک مثال خیلی ساده:

فرض کنید یک برنامه داریم که میخاد یک Child بسازه:

pid = fork();

if (pid == 0) {
// Child
exec(...);
}
else {
// Parent
wait(...);
}


اینجا:

Child

برنامه جدید رو اجرا میکنه

Parent

با wait() منتظر میمونه

وقتی Child تموم شد Parent از wait() برمیگرده و اجرای خودش رو ادامه میده




چرا اصلا باید منتظر بمونیم؟

فرض کنید Shell رو باز کردید و مینویسید:

python test.py

اگر Shell بدون هیچ هماهنگی فورا کارهای بعدی رو انجام بده ممکنه خروجی و ترتیب اجرای برنامه‌ها چیزی نباشه که انتظار داریم

برای اجرای معمولی یک دستور Shell میتونه:

Child بسازه

2 برنامه رو داخل Child اجرا کنه
3 منتظر پایان Child بمونه
4 دوباره Prompt رو نمایش بده

تقریبا:

Shell
│
├── fork()
│
▼
Child
│
└── exec()
│
▼
Program
│
exit()
│
▼
Shell
│
▼
Prompt

یک نکته مهم:
wait() Process رو متوقف نمیکنه

اینجا یک اشتباه رایج وجود داره

وقتی Parent
wait()
میکنه به این معنی نیست که کل سیستم‌ عامل متوقف شده

فقط همون Process منتظر میمونه

CPU
میتونه در همین مدت
Process
های دیگه ای رو اجرا کنه

مثلا:

Parent → Waiting
Child → Running
Chrome → Running
Browser → Running

سیستم‌ عامل همچنان به بقیه Process ها CPU میده


اگر Child زودتر تموم بشه چی؟

اگر Child قبل از اینکه Parent به wait() برسه تموم شده باشه سیستم‌ عامل اطلاعات لازم مربوط به پایان Child رو نگه میداره تا Parent بتونه وضعیت پایان اون رو بگیره

اینجا به مفهوم مهمی به نام Zombie Process میرسیم

Zombie
یعنی Process ی که اجرای خودش رو تموم کرده اما هنوز Parent وضعیت پایان اون رو نگرفته

بعدا درباره Zombie و wait() دقیق‌ تر صحبت میکنیم


حالا سه‌ تایی اصلی رو کنار هم بذاریم

تا اینجا داریم:

fork()


ساختن Child:

Parent
↓
Child

exec()


جایگزین کردن برنامه داخل Child:

Child
↓
New Program

wait()

Parent

منتظر پایان Child میمونه:

Parent
↓
wait()
↓
Child finishes
↓
Parent continues

پس الگوی کلاسیک میشه:

Parent
│
fork()
│
▼
Child
│
exec()
│
▼
New Program
│
exit()
│
▼
Parent
│
wait()
│
▼
Continue

البته ترتیب دقیق wait() نسبت به اجرای Child به نحوه پیاده‌ سازی برنامه بستگی دارده این دیاگرام فقط الگوی مفهومی رو نشون میده




چرا این قسمت مهمه؟

وقتی یک برنامه رو روی Linux تحلیل میکنید و داخل کد به:

fork
exec
wait
exit

میخورید حالا میتونید حدس بزنید چه اتفاقی در حال رخ دادنه

مثلا:

fork()
↓
یک Process جدید
↓
exec()
↓
اجرای برنامه دیگه
↓
exit()
↓
wait()
↓
Parent ادامه میده

دیگه فقط اسم چند تابع رو نمیبینید منطق پشت اونها رو هم میفهمید



تا اینجا Process API رو با سه مفهوم اصلی یاد گرفتیم:

fork() → ساخت Child

exec() →
جایگزین کردن برنامه داخل Process

wait() →
انتظار Parent برای پایان Child

@reverseengine
ReverseEngineering
Parent چطور منتظر Child میمونه؟ تا اینجا دو تا از مهم‌ ترین قسمت‌ های Process API رو یاد گرفتیم: fork() → ساخت Child exec() → اجرای یک برنامه جدید داخل Child حالا یک سؤال: Parent از کجا بفهمد Child کارش تموم شده؟ اینجاست که wait() وارد میشه wait()…
How does Parent
wait for Child?

So far we have learned two of the most important parts of the Process API:

fork() → create Child

exec() → execute a new program inside Child

Now a question:

How does Parent
know when Child is finished?

This is where wait() comes in

What does wait() do?

Very simple:

wait()
makes Parent wait until one of its Children is finished

For example:

Parent
│
├── fork()
│
▼
Child
│

├── does its job
│
└── finishes
│
▼
Parent continues


That is, Parent does not say:

I have nothing more to do with you

So it says:

When the job is done, I will continue

A very simple example:

Suppose we have a program that wants to create a Child:

pid = fork();

if (pid == 0) {
// Child
exec(...);
}
else {
// Parent
wait(...);
}


Here:

Child

executes the new program

Parent

waits with wait()

When Child finishes, Parent returns from wait() and continues its execution

Why do we need to wait at all?

Suppose you open Shell and type:

python test.py

If Shell immediately executes the following tasks without any coordination, the output and execution order of the programs may not be what we expect

For a typical execution of a command Shell can:

Create a Child

2 Run the program inside the Child

3 Wait for the Child to finish

4 Display the Prompt again

Approximately:

Shell
│
├── fork()
│
▼
Child
│
└── exec()
│
▼
Program
│
exit()
│
▼
Shell
│
▼
Prompt


An important point:
wait() does not stop the Process

Here is a common mistake

When Parent
wait()
it does not mean that the entire operating system is stopped

Only that Process waits

The CPU can
execute other Processes

during the same time

For example:

Parent → Waiting
Child → Running
Chrome → Running
Browser → Running


The operating system still gives CPU to other processes

What if the Child terminates early?

If the Child has finished before the Parent reaches wait(), the operating system keeps the necessary information about the end of the Child so that the Parent can get its end status.

Here we come to an important concept called Zombie Process.

Zombie
is a Process that has finished executing but the Parent has not yet received its end status.

We will talk about Zombie and wait() in more detail later.

Now let's put the main three together.

So far we have:

fork()


Create Child:

Parent
↓
Child

exec()


Replace the program inside Child:

Child
↓
New Program

wait()


Parent Waits for the Child to finish:

Parent
↓
wait()
↓
Child finishes
↓
Parent continues
So the classic pattern becomes:

Parent
│
fork()
│
▼
Child
│
exec()
│
▼
New Program
│
exit()
│
▼
Parent
│
wait()
│
▼
Continue


Of course, the exact order of wait() in relation to Child execution depends on how the program is implemented. This diagram only shows the conceptual pattern.

Why is this part important?

When you analyze a program on Linux and you come across:

fork
exec
wait
exit


in the code, you can now guess what is happening.

For example:

fork()
↓
A new Process
↓
exec()
↓
Executing another program
↓
exit()
↓
wait()
↓
Parent continues.


You no longer just see the names of a few functions, you also understand the logic behind them.

So far, we have learned the Process API with three main concepts:

fork() → Creating a Child

exec() → Replacing a program inside a Process

wait() → Waiting for Parent to terminate Child

@reverseengine
The classic and official reference for ELF structure

https://refspecs.linuxfoundation.org/elf/gabi4+/contents.html

@reverseengine
Modern iOS Security Features – A Deep Dive into SPTM, TXM, and Exclaves

https://arxiv.org/pdf/2510.09272

@reverseengine
این مدیوم منه اگه اونجا هم منو فالو کنید خوشحال میشم اونجا هم همین چیزای کانال رو میزارم و ی سری چیزای اضافه🖤

This is my Medium. I would be happy if you followed me there too. I will post the same things from the channel there and a few extra things🩶

https://medium.com/@addcss012
❤8
Dead Code و Junk Code
کدی که هست ولی قرار نیست کاری انجام بده

یکی از روش‌ های رایج Obfuscation اینه که داخل برنامه مقدار زیادی کد اضافه قرار بدن که یا اصلا اجرا نمیشه یا اجرا میشه ولی هیچ تاثیری روی نتیجه نهایی برنامه نداره

هدف اینه که وقتی ما فایل رو باز میکنیم حجم زیادی از دستورهای اضافی ببینیم و پیدا کردن منطق واقعی سخت‌تر بشه

مثلا این کد ساده رو ببینید:

int result = a + b;

int x = 50;
x = x * 2;
x = x - 30;

return result;
اینجا محاسبات مربوط به x هیچ تاثیری روی result نداره پس از نظر منطق برنامه این بخش Dead Code محسوب میشه

حالا یک مثال در اسمبلی ببینیم:

mov eax, 10
add eax, 20

mov ecx, 500
xor ecx, ecx
add ecx, 100

ret



اگر مقدار ecx هیچ جا بعدا استفاده نشه بخش مربوط به ecx عملا تاثیری روی خروجی تابع نداره این همون چیزیه که ما باید یاد بگیریم تشخیص بدیم
اما Junk Code همیشه به این سادگی نیست ممکنه یک Obfuscator دستورهایی اضافه کنه که ظاهرشون مهم به نظر میرسه:

push rax
xor rcx, rcx
inc rcx
dec rcx
pop rax


در ظاهر چند عملیات انجام شده
ولی در اخر وضعیت مهم برنامه تقریبا همون چیزیه که قبل از این بلاک بوده
در تحلیل واقعی یکی از بهترین سوال‌ها اینه:

این بلاک چه چیزی رو تغییر داد که بعدا واقعا استفاده میشه؟

اگر جواب هیچ‌ چیز باشه احتمال داره با Junk Code طرف باشیم یک روش خوب برای تحلیل اینه که فقط مقدار هایی رو دنبال کنید که به خروجی یا مرحله های بعدی برنامه میرسن مثلا اگر یک مقدار داخل RAX ساخته بشه ولی قبل از استفاده دوباره overwrite بشه احتمالا محاسبه قبلی اهمیت نداشته ولی اینجا باید حواستون جمع باشه هر کدی که خروجی واضحی نداره Junk Code نیست ممکنه روی Flagها تاثیر بذاره حافظه رو تغییر بده یا اثر جانبی داشته باشه پس قبل از حذف ذهنی یک بلاک رو باید بررسی کنید:

مقدارهای خروجی کجا میرن؟
حافظه تغییر کرده؟
Flag مهمی تغییر کرده؟
تابع دیگه ای صدا زده شده؟
نتیجه این عملیات بعدا استفاده میشه؟

Deobfuscation
یعنی همین کم‌ کم چیزهایی که تاثیری روی منطق اصلی ندارن کنار میرن و ساختار واقعی برنامه مشخص میشه

تمرین:

این کد رو بررسی کنید:
C++
int calculate(int a, int b)
{
int x = a + b;

int temp = 500;
temp ^= 123;
temp += 20;
temp -= 20;

return x;
}



مشخص کنید کدوم قسمت روی خروجی تابع تاثیر داره و کدوم قسمت فقط باعث شلوغ شدن تحلیل میشه بعد همین مثال رو Compile کنید و داخل Ghidra باز کنید ببینید Compiler با بخش اضافی چه کاری میکنه ممکنه حتی قبل از اینکه تو فایل خروجی رو ببینید خودش کل بخش بی‌ استفاده رو حذف کرده باشه چون کامپایلر ها هم بعضی وقتا برخلاف انتظارمون کار مفید انجام میدن

@reverseengine
❤1
ReverseEngineering
Dead Code و Junk Code کدی که هست ولی قرار نیست کاری انجام بده یکی از روش‌ های رایج Obfuscation اینه که داخل برنامه مقدار زیادی کد اضافه قرار بدن که یا اصلا اجرا نمیشه یا اجرا میشه ولی هیچ تاثیری روی نتیجه نهایی برنامه نداره هدف اینه که وقتی ما فایل رو باز…
Dead Code and Junk Code

Code that exists but is not supposed to do anything

One of the common methods of Obfuscation is to put a lot of extra code into the program that either does not run at all or runs but has no effect on the final result of the program

The goal is that when we open the file, we see a lot of extra instructions and it becomes harder to find the real logic

For example, look at this simple code:

int result = a + b;

int x = 50;
x = x * 2;
x = x - 30;

return result;


Here, the calculations related to x have no effect on the result, so from the logic of the program this section is considered Dead Code

Now let's see an example in assembly:

mov eax, 10
add eax, 20

mov ecx, 500
xor ecx, ecx
add ecx, 100

ret


If the value of ecx is not used anywhere later, the part related to ecx has practically no effect on the output of the function. This is what we need to learn to recognize.

But Junk Code is not always this simple. An Obfuscator may add instructions that appear to be important:

push rax
xor rcx, rcx
inc rcx
dec rcx
pop rax


In appearance, several operations have been performed

But at the end, the important state of the program is almost the same as before this block.

In real analysis, one of the best questions is:

What did this block change that will actually be used later?

If the answer is nothing, we are probably dealing with Junk Code. A good way to analyze is to only follow the values ​​that reach the output or subsequent stages of the program.

For example, if a value is created in RAX but is overwritten before being used again, the previous calculation probably does not matter. But you have to be careful here. Any code that does not have a clear output is not Junk Code. It may affect flags, change memory, or have side effects. So before mentally deleting a block, you should check:

Where do the output values ​​go?

Has memory changed?

Has an important flag changed?

Has another function been called?

Will the result of this operation be used later?

Deobfuscation
That is, gradually things that do not affect the main logic are removed and the real structure of the program is revealed.

Exercise:

Examine this code:

C++
int calculate(int a, int b)
{
int x = a + b;

int temp = 500;

temp ^= 123;

temp += 20;
temp -= 20;

return x;
}


Determine which part affects the output of the function and which part just makes the analysis busy. Then compile this example and open it in Ghidra and see what the compiler does with the extra part. It may have removed the entire useless part even before you see it in the output file, because compilers sometimes do useful things against our expectations.

@reverseengine
❤1