ReverseEngineering
1.32K subscribers
50 photos
11 videos
106 files
888 links
Download Telegram
Chunk
دقیقا چیسه و چه ساختاری داره؟
در پست قبل گفتیم Heap از بخش‌ های کوچیکی به اسم Chunk تشکیل شده
هر بار که برنامه از malloc() استفاده میکنه Allocator یک Chunk در اختیارش قرار میده اما این Chunk فقط یه تیکه حافظه ساده نیست اطلاعات مهمی هم داخلش وجود داره

Chunk
فقط داده‌ ی برنامه نیست
فرض کنید از Heap 100 بایت حافظه درخواست کردید شاید فکر کنید Allocator دقیقا همون 100 بایت رو بهتون میده ولی در عمل این اتفاق نمیوفته
قبل از داده‌ ای که برنامه استفاده میکنه Allocator چند بایت برای خودش کنار میذاره این قسمت همون Metadata هست

پس ساختار یک Chunk تقریبا این شکلیه:

+------------------+
| Metadata |
+------------------+
| User Data |
| |
| |
+------------------+

برنامه فقط به بخش User Data دسترسی داره اما Allocator از Metadata برای مدیریت Heap استفاده میکنه

داخل Metadata چه اطلاعاتی وجود داره؟
بسته به نوع Allocator ممکنه فرق داشته باشه اما معمولا اطلاعاتی مثل این‌ ها نگهداری میشن:

اندازه‌ی Chunk
وضعیت آزاد یا اشغال بودن
اطلاعات لازم برای مدیریت حافظه
ارتباط با Chunk های کناری در بعضی Allocator ها
این اطلاعات باعث میشن Allocator بدونه هر قسمت از Heap چه وضعیتی داره

چرا اندازه‌ی واقعی Chunk با چیزی که درخواست کرده فرق داره؟

فرض کنی این درخواست رو نوشتید:

malloc(100);

این به این معنی نیست که دقیقا 100 بایت از Heap اشغال میشه.

چون Allocator باید:

Metadata
رو ذخیره کنه
حافظه رو تراز (Alignment) کنه
اندازه‌ ها رو به واحد های مشخص گرد کنه
برای همین ممکنه در عمل بیشتر از 100 بایت مصرف بشه
این موضوع یکی از چیزهاییه که خیلی از برنامه‌نویس‌ها در ابتدا بهش توجه نمیکنن


Alignment
پردازنده دوست داره داده‌ ها روی مرز های مشخصی از حافظه قرار بگیرن
مثلا روی سیستم‌های 64 بیتی معمولا داده‌ ها روی مضرب‌های 8 یا 16 بایت تراز میشن این کار باعث میشه دسترسی به حافظه سریع‌ تر و بهینه‌ تر باشه
به همین دلیل Allocator بعضی وقتا اندازه‌ ی درخواستی رو کمی بزرگ‌ تر در نظر میگیره

وقتی free() صدا زده میشه چه اتفاقی میوفته؟

خیلی‌ها فکر میکنن با free() حافظه بلافاصله از بین میره
در واقع معمولا این‌طور نیست
بیشتر Allocator ها حافظه رو فقط آزاد علامت‌ گذاری میکنن تا بعدا دوباره از همون Chunk استفاده کنن
یعنی داده‌هایی که داخل اون Chunk بودن ممکنه هنوز در حافظه باقی مونده باشن فقط برنامه دیگه نباید از اونها استفاده کنه
به همین خاطر باگ‌ هایی مثل Use-After-Free به وجود میان یعنی برنامه بعد از آزاد شدن حافظه اشتباها دوباره به همون بخش دسترسی پیدا میکنه


Heap
از بخش‌هایی به نام Chunk تشکیل شده که هر کدام علاوه بر فضای مورد استفاده‌ ی برنامه اطلاعات مدیریتی هم دارن این اطلاعات به Allocator کمک میکنه تا حافظه رو مدیریت کنه همچنین حافظه‌ ای که با free() آزاد میشه معمولا بلافاصله پاک نمیشه بلکه برای استفاده‌ ی مجدد آماده نگه داشته میشه

@reverseengine
❤1
ReverseEngineering
Chunk دقیقا چیسه و چه ساختاری داره؟ در پست قبل گفتیم Heap از بخش‌ های کوچیکی به اسم Chunk تشکیل شده هر بار که برنامه از malloc() استفاده میکنه Allocator یک Chunk در اختیارش قرار میده اما این Chunk فقط یه تیکه حافظه ساده نیست اطلاعات مهمی هم داخلش وجود داره…
Chunk

What exactly is it and what is its structure?

In the previous post, we said that the Heap is made up of small sections called Chunk
Every time the program uses malloc(), the Allocator provides it with a Chunk, but this Chunk is not just a simple piece of memory, it also contains important information

Chunk
is not just the program's data
Suppose you request 100 bytes of memory from the Heap, you might think that the Allocator will give you exactly those 100 bytes, but in practice this does not happen
Before the data that the program uses, the Allocator sets aside a few bytes for itself. This part is called Metadata

So the structure of a Chunk is approximately like this:

+----+
| Metadata |
+------------------+
| User Data |
| |
| |
+------------------+

The program only has access to the User Data section, but the Allocator uses Metadata to manage the Heap

What information is inside Metadata?
It may vary depending on the type of Allocator, but usually information like this is kept:

Chunk size

Free or busy status

Information needed for memory management

Relationship with neighboring Chunks in some Allocators

This information lets the Allocator know what state each part of the Heap is in

Why is the actual Chunk size different from what it requested?

Suppose you wrote this request:

malloc(100);

This does not mean that exactly 100 bytes of the Heap will be occupied.

Because the Allocator must:

Store metadata

Align the memory

Round the sizes to specific units

This may actually use more than 100 bytes

This is something that many programmers don't pay attention to at first

Alignment

The processor likes data to be on specific memory boundaries

For example, on 64-bit systems, data is usually aligned on multiples of 8 or 16 bytes, which makes memory access faster and more efficient

For this reason, the Allocator sometimes considers the requested size to be slightly larger

What happens when free() is called?

Many people think that free() immediately destroys the memory
In fact, this is usually not the case
Most allocators only mark the memory as free so that they can use the same chunk again later
That means that the data inside that chunk may still be in memory, but the program should not use it anymore
That is why bugs like Use-After-Free occur, meaning that the program mistakenly accesses the same section again after the memory has been freed
Heap
Consists of sections called Chunks, each of which, in addition to the space used by the program, also has management information. This information helps the allocator to manage the memory. Also, the memory that is freed with free() is usually not immediately deleted, but is kept ready for reuse

@reverseengine
بخش بیست و سوم بافر اورفلو


Information Leak

یعنی برنامه ناخواسته اطلاعاتی از حافظه رو نمایش بده یا برگردونه که این اطلاعات میتونه شامل
آدرس‌ های حافظه
داده‌ های حساس
رشته‌ های محرمانه
محتوای متغیرها
باشه
یک مثال ساده:

#include <stdio.h>

int main() {

int numbers[5] = {1,2,3,4,5};

printf("%d\n", numbers[10]);

return 0;
}


مشکل این کجاست؟
برنامه داره مقداری خارج از آرایه رو میخونه ممکنه چیزی که چاپ میشه مربوط به یک متغیر دیگه یا بخشی از حافظه باشه
این یعنی اطلاعاتی که نباید دیده بشن نمایش داده شدن
یک مثال دیگه:

char secret[] = "password123";

printf("%s\n", secret);


اگر برنامه به اشتباه آدرس این رشته رو در اختیار کاربر قرار بده
یا مسیر اجرای برنامه طوری باشه که این داده نمایش داده بشه
یک Information Leak رخ داده

چرا برای مهندسی معکوس مهمه؟

فرض کنید یک برنامه ASLR داره
یعنی آدرس‌ های حافظه هر بار تغییر میکنن
اگر یک Information Leak پیدا کنید
ممکنه آدرس یکی از توابع یا کتابخانه‌ها رو به دست بیارید حالا میتونید تحلیل دقیق‌ تری انجام بدید و بفهمید برنامه چطور در حافظه قرار گرفته
به همین دلیل Information Leak خیلی وقت‌ها اولین قدم برای تحلیل آسیب‌پذیری‌ های پیچیده‌ تره

موقع تحلیل باینری دنبال چی بگردیم؟
اگر دیدید برنامه آدرس اشاره‌ گرها رو چاپ میکنه داده‌ ای خارج از محدوده میخونه
پیام‌ های خطای بیش از حد دقیق نمایش میده اطلاعات حافظه رو بدون بررسی برمیگردونه باید بیشتر بررسیش کنید

یک مثال اسمبلی:

lea rdi,[rip+message]
call puts


این خودش مشکلی نداره ولی اگر قبلا puts آدرس یا داده‌ ای از حافظه بدون کنترل آماده شده باشه باید بررسی کنید که آیا اطلاعات حساسی ممکنه نمایش داده بشه یا نه؟


Information Leak
برنامه اطلاعاتی رو که نباید در اختیار کاربر قرار بده نمایش میده این اطلاعات ممکنه برای تحلیل باینری یا پیدا کردن مسیرهای آسیب‌پذیر خیلی ارزشمند باشن برای یک Reverse Engineer پیدا کردن این نشت‌ های اطلاعاتی یکی از مهارت‌ های مهمه

@reverseengine
❤1
ReverseEngineering
بخش بیست و سوم بافر اورفلو Information Leak یعنی برنامه ناخواسته اطلاعاتی از حافظه رو نمایش بده یا برگردونه که این اطلاعات میتونه شامل آدرس‌ های حافظه داده‌ های حساس رشته‌ های محرمانه محتوای متغیرها باشه یک مثال ساده: #include <stdio.h> int main() { …
Part 23 Buffer Overflow


Information Leak

This means that the program unintentionally displays or returns information from memory, which can include
memory addresses
sensitive data
secret strings
variable contents
A simple example:

#include <stdio.h>

int main() {

int numbers[5] = {1,2,3,4,5};

printf("%d\n", numbers[10]);

return 0;
}


What is the problem?

The program is reading a value outside the array. What is printed may be related to another variable or part of memory.

This means that information that should not be seen is displayed.

Another example:

char secret[] = "password123";

printf("%s\n", secret);


If the program mistakenly provides the address of this string to the user
or the program execution path is such that this data is displayed
an Information Leak has occurred

Why is it important for reverse engineering?

Suppose a program has ASLR
that is, the memory addresses change every time
If you find an Information Leak
you may get the address of one of the functions or libraries. Now you can do a more detailed analysis and understand how the program is located in memory
That is why Information Leak is often the first step in analyzing more complex vulnerabilities

What should we look for when analyzing binary?

If you see that the program prints pointer addresses, reads data out of bounds, displays overly detailed error messages, returns memory information without checking, you should investigate further. Here is an assembly example: lea rdi,[rip+message] call puts This is not a problem, but if the address or data from memory has been prepared without checking, you should check whether sensitive information may be displayed. Information Leak The program displays information that should not be made available to the user. This information may be very valuable for binary analysis or finding vulnerable paths. Finding these information leaks is one of the important skills for a reverse engineer.


@reverseengine
Process API
برنامه‌ ها چطور با سیستم‌ عامل صحبت میکنن
تا اینجا فهمیدیم سیستم‌ عامل میتونه Process بسازه اجرا کنه و بین اونا جا به‌ جا بشه

اما یک سوال مهم:

برنامه‌ ها چطور به سیستم‌ عامل میگن که یک Process جدید بساز یا یک برنامه رو اجرا کن؟

آیا مستقیما با کرنل صحبت میکنن؟
نه
برنامه‌ها از چیزی به نام API (Application Programming Interface) استفاده میکنند

API
رو میتونیم مثل یک واسطه بین برنامه و سیستم‌ عامل در نظر بگیریم

API
یعنی چی؟
فرض کنید به یک رستوران رفتید
شما مستقیم وارد آشپزخانه نمیشید تا غذا درست کنید
سفارش رک به گارسون میدید گارسون اون رو به آشپز میرسونه و بعد غذا رو براتون میارن

در این مثال:

آشپز = Kernel

گارسون = API

مشتری = برنامه

برنامه هم به همین شکل درخواستش رو از طریق API به سیستم‌ عامل میفرسته

Process API
چیه؟
Process API
مجموعه‌ای از توابع که به برنامه اجازه میده Processها رو مدیریت کنه

مثلا:

ساختن یک Process جدید
اجرای یک برنامه
منتظر موندن تا پایان یک Process
تموم شدن به یک Process

مهم‌ترین توابع در لینوکس

در لینوکس بیشتر با این سه تابع آشنا میشید:
fork()
وظیفه‌اش ساختن یک Process جدیده
وقتی fork() اجرا میشه سیستم‌عامل از Process فعلی یک کپی میسازه

بعد از اون دو Process وجود دارند:

Parent (والد)

Child (فرزند)

هر دو از همون نقطه به اجرای خودشون ادامه میدن

exec()
بعضی وقتا نمیخایم فقط یک کپی از Process داشته باشیم میخایم برنامه دیگه ای اجرا بشه

اینجاست که exec() وارد عمل میشه


exec()
برنامه فعلی رو با یک برنامه جدید جایگزین میکنه

مثلا یک Shell بعد از گرفتن دستور کاربر با exec() برنامه‌ ای مثل ls یا cat رو اجرا میکنه

wait()
اگر Process والد بخاد صبر کنه تا Process فرزند کارش تموم بشه از wait() استفاده میکنه
بدون wait() ممکنه والد زودتر ادامه پیدا کنه و ترتیب اجرای برنامه به‌ هم بخوره


یک مثال واقعی:

فرض کنید در ترمینال مینویسید:
python script.py

پشت صحنه تقریبا این اتفاق‌ ها میوفته:

Shell
یک Process جدید با fork() ایجاد میکنه

Process
فرزند با exec() برنامه Python رو اجرا میکنه

Process
والد با wait() منتظر میمونه تا اجرای اسکریپت تموم بشه

همه این مراحل در کسری از ثانیه انجام میشن

چرا این موضوع مهمه؟

اگر روی لینوکس برنامه‌ ها یا بدافزار ها رو تحلیل کنید به مرتب با این توابع رو به‌ رو میشید همچنین در ویندوز هم مفاهیم مشابه ای وجود دارن فقط اسم توابع فرق میکنه برای مثال ایجاد Process در ویندوز معمولا با توابعی مانند

CreateProcess
انجام میشه اما ایده اصلی همونه:

درخواست از سیستم‌ عامل برای ساخت و مدیریت یک Process


Process API
راه ارتباط برنامه با سیستم‌ عامل برای مدیریت Process هاست

سه تابع مهم:

fork() → ساختن Process جدید

exec() → اجرای یک برنامه جدید

wait()
منتظر موندن برای تموم شدن Process فرزند

اگر این سه تابع رو خوب بفهمید درک نحوه اجرای برنامه‌ ها داخل لینوکس و حتی مفاهیم مشابه در ویندوز براتون بسیار راحت‌ تره
👍1
ReverseEngineering
Process API برنامه‌ ها چطور با سیستم‌ عامل صحبت میکنن تا اینجا فهمیدیم سیستم‌ عامل میتونه Process بسازه اجرا کنه و بین اونا جا به‌ جا بشه اما یک سوال مهم: برنامه‌ ها چطور به سیستم‌ عامل میگن که یک Process جدید بساز یا یک برنامه رو اجرا کن؟ آیا مستقیما با…
Process API How do programs talk to the operating system
So far we have understood that the operating system can create, run, and switch between processes

But one important question:

How do programs tell the operating system to create a new process or run a program?

Do they talk directly to the kernel?

No
Programs use something called API (Application Programming Interface)

We can think of API
as an intermediary between the program and the operating system

What does API
mean?
Suppose you went to a restaurant
You didn't go directly into the kitchen to prepare the food
You would give the order to the waiter, the waiter would give it to the chef, and then they would bring the food to you

In this example:

Chef = Kernel

Waiter = API

Customer = Program

The program also sends its request to the operating system via API

What is Process API?

Process API A set of functions that allow a program to manage processes

For example:

Creating a new process

Executing a program

Waiting for a process to finish

Exiting a process

Most important functions in Linux

In Linux, you will be most familiar with these three functions:

fork()
Its function is to create a new process

When fork() is executed, the operating system makes a copy of the current process

After that, there are two processes:

Parent

Child

Both continue their execution from the same point

exec()
Sometimes we don't want to have just a copy of the process, we want another program to run

This is where exec() comes in

exec()
Replaces the current program with a new program

For example, a shell executes a program like ls or cat after receiving a user command with exec()

wait()
If the parent process wants to wait until the process The child uses wait() when it finishes its work

Without wait(), the parent might continue earlier and the execution order of the program might be disrupted

A real-world example:

Suppose you type in the terminal:
python script.py

Behind the scenes, this is what happens:

Shell
creates a new Process with fork()

Process
the child executes the Python program with exec()

Process
the parent waits with wait() until the script is finished

All of these steps are done in a fraction of a second

Why is this important?

If you analyze programs or malware on Linux, you will encounter these functions regularly. There are also similar concepts in Windows, only the names of the functions are different. For example, creating a Process in Windows is usually done with functions like

CreateProcess
, but the main idea is the same:

Requesting the operating system to create and manage a Process

Process API

The way a program communicates with the operating system to manage Processes

Three important functions:

fork() Create a new Process

exec() Run a new program

wait() Wait for a child Process to finish

If you understand these three functions well, it will be much easier for you to understand how to run programs in Linux and even similar concepts in Windows
بخش بیست و چهارم بافر اورفلو


Fuzzing
شکار باگ بدون اینکه خط به خط کد رو بخونیم


تا اینجا خودمان با تحلیل کد و اسمبلی دنبال باگ می‌گشتیم
ولی اگر برنامه چند میلیون خط کد داشته باشه چی

اینجاست که Fuzzing وارد میشه
Fuzzing
به جای اینکه ما دنبال باگ بگردیم خودش هزار بار یا حتی میلیون‌ ها ورودی مختلف به برنامه میده تا ببیند برنامه کرش میکنه یا نه
Fuzzing یعنی چی
به زبان ساده
یک ابزار به صورت خودکار ورودی‌ های مختلف تولید میکنه و به برنامه میده
اگر برنامه
کرش کنه
هنگ کنه
رفتار غیرعادی داشته باشه
ابزار اون ورودی رو ذخیره میکنه تا بعدا بررسی کنیم

یک مثال ساده:

فرض کنید برنامه فقط یک رشته از کاربر بگیره

#include <stdio.h>

int main() {

char input[64];

fgets(input,sizeof(input),stdin);

printf("%s",input);

return 0;
}


یک Fuzzer ممکنه این ورودی‌ ها رو امتحان کنه

AAAA

AAAAAAAAAAAAAAAAAAAAAAAAAAAA

123456789

!@#$%^&*

AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA



چرا Fuzzing مهمه؟

چون انسان نمیتونه میلیون‌ ها ورودی رو امتحان کنه
ولی Fuzzer این کار رو در مدت کوتاهی انجام میده
به همین دلیل خیلی از آسیب‌پذیری‌ های معروف دنیا اولین بار با Fuzzing پیدا شدن

انواع Fuzzing:

Dumb Fuzzing
ساده‌ترین حالت
فقط داده‌ های تصادفی به برنامه میده
هیچ اطلاعی از ساختار برنامه نداره

Smart Fuzzing
ساختار ورودی رو میشناسه

مثلا اگر برنامه فایل PNG باز میکنه
ورودی‌هایی شبیه PNG تولید میکنه
در نتیجه شانس پیدا کردن باگ بیشتر میشه

Coverage Guided Fuzzing
این روش خیلی محبوبه
ابزار بررسی میکنه هر ورودی باعث اجرای کدوم قسمت‌ های برنامه شده
اگر ورودی جدید مسیر جدیدی از کد رو اجرا کنه
همون مسیر رو بیشتر بررسی میکنه
به همین دلیل خیلی سریع‌ تر از روش‌ های ساده باگ پیدا میکنه

ابزارهای معروف

چند ابزار معروف که تقریباً هر Reverse Engineer باید اسمشون رو بدونه

AFL++
libFuzzer
Honggfuzz


این ابزار ها سال‌ هاست برای پیدا کردن باگ‌های حافظه استفاده میشن

موقع مهندسی معکوس چرا مهمه؟

فرض کنید یک باینری دارید و هیچ سورسی از اون موجود نیست

میتونید اون رو Fuzz کنید
اگر کرش کرد
همن ورودی رو داخل GDB یا IDA بررسی کنید

و قدم به قدم علت کرش رو پیدا میکنید
به همین دلیل Fuzzing و Reverse Engineering مکمل هم هستن


Fuzzing
یعنی به جای اینکه خودتون حدس بزنید چه ورودی باعث باگ میشه یک ابزار هزار بار یا میلیون‌ها ورودی مختلف رو امتحان میکنه هر جا برنامه رفتار غیرعادی داشت
همان نقطه تبدیل به هدف تحلیل مهندسی معکوس میشه

تمرین:

یک برنامه ساده که از ورودی کاربر استفاده میکنه بنویسید بعد فکر کنید اگر قرار بود یک Fuzzer برای اون بنویسید چه نوع ورودی‌ هایی رو امتحان میکردید

@reverseengine
ReverseEngineering
بخش بیست و چهارم بافر اورفلو Fuzzing شکار باگ بدون اینکه خط به خط کد رو بخونیم تا اینجا خودمان با تحلیل کد و اسمبلی دنبال باگ می‌گشتیم ولی اگر برنامه چند میلیون خط کد داشته باشه چی اینجاست که Fuzzing وارد میشه Fuzzing به جای اینکه ما دنبال باگ بگردیم…
Part 24 Buffer Overflow


Fuzzing
Bug hunting without reading the code line by line

So far we have been looking for bugs ourselves by analyzing the code and assembly

But what if the program has several million lines of code

This is where Fuzzing comes in

Fuzzing
Instead of us looking for bugs, it gives the program thousands or even millions of different inputs to see if the program crashes or not

What does Fuzzing mean

In simple terms
A tool automatically generates different inputs and gives them to the program

If the program

Crash
Hangs

Or behaves abnormally
The tool saves that input for later review

A simple example:

Suppose the program only takes a string from the user

#include <stdio.h>

int main() {

char input[64];

fgets(input,sizeof(input),stdin);

printf("%s",input);

return 0;

}


A Fuzzer might try these inputs

AAAA

AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA

123456789

!@#$%^&*

AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA


Why is Fuzzing Important?

Because humans cannot try millions of inputs
But Fuzzer does this in a short time
That is why many of the world's famous vulnerabilities were first found with Fuzzing

Types of Fuzzing:

Dumb Fuzzing
The simplest
Only gives random data to the program
It has no information about the structure of the program

Smart Fuzzing
It knows the structure of the input

For example, if the program opens a PNG file
It produces inputs similar to PNG
As a result, the chances of finding a bug increase

Coverage Guided Fuzzing
This method is very popular
The tool checks which parts of the program each input causes to be executed
If the new input executes a new path of code
It checks the same path more
That is why it finds bugs much faster than simple methods

Famous tools

A few famous tools that almost every Reverse Engineer should know their names

AFL++
libFuzzer
Honggfuzz


These tools have been used for years to find Memory bugs are used

Why is it important when reverse engineering?

Suppose you have a binary and no source is available

You can fuzz it

If it crashes

Examine the same input in GDB or IDA

And you will find the cause of the crash step by step

That is why Fuzzing and Reverse Engineering are complementary

Fuzzing

Instead of guessing what input causes the bug, a tool tries thousands or millions of different inputs. Wherever the program behaves abnormally

That point becomes the target of reverse engineering analysis

Exercise:

Write a simple program that uses user input. Then think about what kind of inputs you would try if you were to write a fuzzer for it

@reverseengine
Dissecting_the_Dark_Web_Reverse_Engineering_the_Tools_of_the_Underground.pdf
21.7 MB
D I S S E C T I N G T H E
DARK WEB

R e v e r s e E n g i n e e r i n g t h e To o l s
of the Underground Economy

@reverseengine