ReverseEngineering
1.32K subscribers
50 photos
11 videos
108 files
895 links
Download Telegram
Virtualization مجازی‌ سازی

فرض کنید سیستم شما:

8 گیگ رم داره

8 هسته CPU داره

اما همزمان:

مرورگر بازه

تلگرام بازه

موزیک پخش میشه

VS Coode بازه

هر برنامه فکر میکنه:

این CPU مال اونه

این حافظه مال اونه

در حالی که واقعیت این نیست

سیستم‌ عامل منابع واقعی رو بین برنامه‌ ها

تقسیم میکنه و به هر برنامه یک تصویر

مجازی از منابع میده

مثال CPU:

فرض کنید فقط یک CPU دارید

اما همزمان:
Chrome

Telegram

Discord
در حال اجران

سیستم‌ عامل خیلی سریع بین اون جا به‌ جا میشه اینقدر سریع که شما فکر میکنید همه با هم اجرا میشن در حالی که CPU در هر لحظه فقط یک کار انجام میده

مثال حافظه:

هر برنامه فکر میکنه:
Plain text

من از آدرس 0x00000000 شروع میشم
اما در واقعیت همه برنامه‌ها در حافظه واقعی کنار هم قرار دارن
سیستم‌عامل با Virtual Memory این موضوع رو مدیریت میکنه

چرا برای مهندسی معکوس مهمه؟

چون موضوعاتی مثل:
Stack
Heap
ASLR
Paging
Virtual Memory

همه از همین مفهوم Virtualization میان



Virtualization

Suppose your system:

has 8 GB of RAM

has 8 CPU cores

but at the same time:

browser is open

telegram is open

music is playing

vs code is open

each program thinks:

this is its CPU

this is its memory

while in reality this is not

the operating system divides real resources between programs

and gives each program a virtual image

of resources

CPU example:

Suppose you only have one CPU

but at the same time:

Chrome

Telegram

Discord

are running

the operating system switches between them very quickly so fast that you think they are all running together while the CPU is only doing one thing at a time

Memory example:

each program thinks:

plain text

I start at address 0x00000000

but in reality all programs are located together in real memory

Operating system with Virtual Memory handles this

Why is it important for reverse engineering?

Because topics like:

Stack
Heap
ASLR
Paging
Virtual Memory

all come from the same concept of Virtualization

@reverseengine
Devirtualization مجازی‌ سازی‌ زدایی

بعضی محافظ‌ ها مثل VMProtect کد اصلی برنامه رو به بایت‌ کد تبدیل میکنن و یک ماشین مجازی داخل برنامه قرار میدن تا اون بایت‌ کد رو اجرا کنه مشکل از جایی شروع میشه که دیگه با یک تابع معمولی طرف نیستیم به جای چند دستور اسمبلی ساده با صدها یا هزار دستور مربوط به ماشین مجازی طرف میشیم
اینجاست که Devirtualization به دردمون میخوره

Devirtualization
یعنی تلاش برای تبدیل منطق مخفی‌ شده داخل ماشین مجازی به شکلی که دوباره قابل فهم باشه

هدف Devirtualization این نیست که محافظ رو حذف کنیم
هدف اینه که بفهمیم برنامه واقعا چه کاری انجام میده

فرض کنید کد اصلی این بوده:

C

if(password == "1234") { success(); }


بعد از Virtualization ممکنه این منطق به صد ها دستور بایت‌ کد تبدیل بشه
در ظاهر دیگر هیچ strcmp یا if مشخصی وجود نداره فقط یک Dispatcher و تعداد زیادی Handler دیده میشه
کاری که تحلیلگر انجام میده اینه که به تدریج معنی هر Opcode رو کشف کنه

مثلا متوجه میشه:

Opcode 0x01 داده‌ای رو داخل رجیستر مجازی اپلود میکنه


Opcode 0x02 دو مقدار رو مقایسه میکنه


Opcode 0x03 یک پرش شرطی انجام میده


بعد از شناسایی Opcode ها میتونیم رفتار ماشین مجازی رو روی کاغذ بازسازی کنیم
به این فرایند ساختن نقشه Opcode ها
معمولا فرایند تحلیل به این شکل پیش میره:

اول Dispatcher پیدا میشه

بعد Handler های مختلف شناسایی میشن

بعد Opcode ها دسته‌ بندی میشن

در نهایت منطق اصلی برنامه بازسازی میشه

یکی از اشتباهات رایج افراد تازه‌کار اینه که مستقیم سراغ Handler ها میرن
در حالی که اول باید ساختار کلی ماشین مجازی رو بفهمید

اگر Dispatcher رو نفهمید Handler ها تقریبا بی‌ معنی هستن

یک قانون مهم در تحلیل ماشین‌ های مجازی:

اول جریان اجرای بایت‌ کد رو درک کنید

بعد سراغ معنی دستور ها برید

تمرین

یک ماشین مجازی ساده طراحی کنید که فقط سه تا Opcode داشته باشه:

LOAD

ADD

PRINT

بعد سعی کنید فقط با نگاه کردن به اجرای برنامه بفهمید هر Opcode چه کاری انجام میده



Devirtualization

Some protectors, such as VMProtect, convert the original program code into bytecode and place a virtual machine inside the program to execute that bytecode. The problem starts when we are no longer dealing with a normal function, but instead of a few simple assembly instructions, we are dealing with hundreds or thousands of instructions related to the virtual machine.

This is where Devirtualization comes in.

Devirtualization
is an attempt to convert the hidden logic inside the virtual machine into a form that is understandable again.

The goal of Devirtualization is not to remove the protector. The goal is to understand what the program is actually doing.

Suppose the original code was:

C

if(password == "1234") { success(); }


After Virtualization, this logic may be converted into hundreds of bytecode instructions. On the surface, there is no specific strcmp or if, only a Dispatcher and a large number of Handlers. What the analyzer does is to gradually discover the meaning of each Opcode.

For example, it finds:

Opcode 0x01 uploads data into a virtual register


Opcode 0x02 compares two values


Opcode 0x03 performs a conditional jump


After identifying the Opcodes, we can reconstruct the behavior of the virtual machine on paper. This process is called creating an Opcode map.

Usually, the analysis process goes like this:

First, the Dispatcher is found.

Then the different Handlers are identified.

Then the Opcodes are categorized.

Finally, the main logic of the program is reconstructed.

One of the common mistakes of beginners is to go straight to the Handlers.

While you should first understand the overall structure of the virtual machine.

If you don't understand the Dispatcher, the Handler are almost meaningless

An important rule in analyzing virtual machines:

First understand the flow of bytecode execution

Then move on to the meaning of the instructions

Exercise

Design a simple virtual machine that has only three Opcodes:

LOAD

ADD

PRINT

Then try to figure out what each Opcode does just by looking at the program execution

@reverseengine
بخش هجدهم بافر اورفلو


انواع Buffer Overflow
خیلی‌ها وقتی اسم Buffer Overflow میاد فقط یاد استک میفتن در صورتی که بافر اورفلو فقط Stack Overflow نیست
چندین نوع مختلف داره که هر کدوم رفتار و اثر متفاوتی دارن

نوع اول
Stack Buffer Overflow
معروف‌ترین نوع
وقتی اتفاق میفته که داده بیشتر از ظرفیت یک بافر روی استک نوشته بشه

مثال:

void vuln(char *input) { char buf[16]; strcpy(buf,input); }

اینجا اگر بیشتر از 16 بایت ارسال بشه
داده از بافر خارج میشه
و کم کم به
متغیرهای محلی
Saved RBP

Return Address

میرسه
همون چیزی که تا الان یاد گرفتیم

نوع دوم
Heap Buffer Overflow
این یکی روی Heap اتفاق میفته
نه روی Stack

مثال:

char *buf = malloc(16); strcpy(buf,input);


اگر ورودی بیشتر از 16 بایت باشه
حافظه‌های کنار این Chunk خراب میشن

چرا مهمه؟
چون روی Heap معمولا خبری از Return Address نیست پس هدف مهاجم بیشتر خراب کردن ساختار Heap یا دستکاری Pointer هاست

نوع سوم
Off By One
یکی از باگ‌های مورد علاقه مهندس‌ های معکوسه
چون ظاهرا کوچیک به نظر میاد
ولی گاهی به اکسپلویت کامل تبدیل میشه

مثال:

char buf[16]; for(int i=0;i<=16;i++) { buf[i]='A'; }

اشتباه اینجاست
i <= 16
باید میبود
i < 16

اینجا فقط یک بایت خارج از بافر نوشته میشه ولی همون یک بایت بعضی وقتا برای خراب کردن ساختار حافظه کافیه

نوع چهارم
Integer Overflow
این یکی مستقیم Buffer Overflow نیست
ولی خیلی وقت‌ها باعثش میشه

مثال:

size = count * 100; buf = malloc(size);


فرض کن count خیلی بزرگ باشه
در نتیجه ضرب از محدوده نوع داده خارج بشه و size کوچیکتر از چیزی بشه که انتظار داریم بعد برنامه فکر میکنه حافظه زیادی گرفته ولی در واقع حافظه کمی گرفته
و در ادامه Overflow رخ میده

نوع پنجم
Format String
از نظر فنی Buffer Overflow نیست
ولی همیشه کنار این مباحث آموزش داده میشه

مثال:

printf(user_input);
به جای
printf("%s",user_input);


اینجا کاربر میتونه فرمت‌های printf رو کنترل کنه

مثل:

%x %x %x %x


و اطلاعات حافظه رو بخونه
برای شما مهمه چون خیلی وقت‌ ها باگ Format String تبدیل به راهی برای دور زدن ASLR یا پیدا کردن آدرس‌ ها میشه


Stack Overflow

خراب کردن استک و Return Address

Heap Overflow
خراب کردن حافظه Heap


Off By One
نوشتن فقط یک بایت اضافه


Integer Overflow
اشتباه در محاسبه اندازه حافظه


Format String
کنترل فرمت‌های printf و نشت اطلاعات



وقتی یک باینری رو تحلیل میکنید
سعی کنید تشخیص بدید باگ از کدوم دسته است
چون روش تحلیل Stack Overflow با Heap Overflow کاملا فرق میکنه
و همین تشخیص اولیه خیلی وقتا نصف حل مسئله است

@reverseengine
Part 18 Buffer Overflow


Types of Buffer Overflow

Many people only think of stack when the name Buffer Overflow comes up, but buffer overflow is not just Stack Overflow

It has several different types, each with different behavior and effects

Type 1

Stack Buffer Overflow

The most famous type

It happens when data is written to the stack beyond the capacity of a buffer

Example:

void vuln(char *input) { char buf[16]; strcpy(buf,input); }

Here, if more than 16 bytes are sent

the data is overflowed

and gradually reaches

local variables

Saved RBP

Return Address

The same thing we have learned so far

Type 2

Heap Buffer Overflow

This one happens on the Heap

not on the Stack

Example:

char *buf = malloc(16); strcpy(buf,input);

If the input is more than 16 bytes
The memories next to this Chunk will be corrupted

Why is it important?
Since there is usually no Return Address on the Heap, the attacker's goal is more to corrupt the Heap structure or manipulate the Pointers

The third type
Off By One
It is one of the favorite bugs of reverse engineers
Because it seems small
But sometimes it turns into a full-fledged exploit

Example:

char buf[16]; for(int i=0;i<=16;i++) { buf[i]='A'; }

The error here is
i <= 16
It should be
i < 16
Here only one byte is written outside the buffer, but that one byte is sometimes enough to corrupt the memory structure

The fourth type
Integer Overflow
This one is not a direct Buffer Overflow
But it often causes it

Example:

size = count * 100; buf = malloc(size);

Suppose count is too large
As a result, the multiplication goes out of the data type range and size becomes smaller than we expect. Then the program thinks it has taken up a lot of memory, but in fact it has taken up little memory
And then Overflow occurs

The fifth type
Format String
Is not technically a Buffer Overflow
But it is always taught alongside these topics

Example:

printf(user_input);
Instead of
printf("%s",user_input);

Here the user can control printf formats

Like:

%x %x %x %x

And read memory information

This is important for you because many times the Format String bug becomes a way to bypass ASLR or find addresses

Stack Overflow

Corrupting the stack and Return Address

Heap Overflow

Corrupting the Heap memory

Off By One
Writing only one extra byte

Integer Overflow
Error in calculating the memory size

Format String
Controlling printf formats and information leaks

When you analyze a binary
Try to identify which category the bug belongs to

Because the analysis method for Stack Overflow is completely different from Heap Overflow

And this initial identification is often half the solution to the problem

@reverseengine
Understanding_the_Linux_Kernel_Daniel_P_Bovet,_Marco_Cesati.pdf
4.8 MB
Main Chapters:
Introduction
Memory Addressing
Processes
Interrupts and Exceptions
Timing Measurements
Memory Management
Process Address Space
System Calls
Signals
Process Scheduling
Kernel Synchronization
The Virtual Filesystem
I/O Device Management
Disk Caches
Accessing Regular Files
Swapping
The Ext2 Filesystem
Process Communication
Program Execution
Appendices: System Startup, Modules, Source Code Structure
Talbot.pdf
820.4 KB
Reverse-Engineering the Intel Address Translation Caches

https://github.com/0xADE1A1DE/Talbot
Post-Build PE Obfuscation

Obfusk8 includes a post-build script to further harden the compiled binary by removing forensic artifacts.

* Script Location: Obfusk8/Obfusk8/SCRIPTS/obfuscate_pe.ps1 at main · x86byte/Obfusk8
* What it does:
1. Strips the Rich Header — removes the MSVC build-environment fingerprint that reveals compiler version and toolchain details.
2. Spoofs the TimeDateStamp — replaces the PE header timestamp with a fixed value to obscure build time.
3. Clears the Debug Directory — wipes debug directory entries that could leak PDB paths or build metadata.
* Usage:
Run as a post-build step after compiling:
powershell PowerShell -NoProfile -ExecutionPolicy Bypass -File Obfusk8/SCRIPTS/obfuscate_pe.ps1 -Path "path\to\Obfusk8.exe"

The script modifies the binary in-place. No backup is created.
We’re building a small community around binary security research, focused on things like:

- Reverse Engineering
- Binary Obfuscation / Deobfuscation
- Exploit Development
- Compiler / interpreters...
- Malware Analysis
- Binary Hardening research

we also work on open source tools and experiments here:
GitHub → BinaryHardening GitHub
Discord → BinaryHardening Discord
Toolkit for Windows internals, vulnerability analysis, and reproducible security research workflows.

https://github.com/kernelstub/NTForge