Reverse Engineering a 0day used Against CrowdStrike EDR
https://medium.com/@jehadbudagga/reverse-engineering-a-0day-used-against-crowdstrike-edr-a5ea1fbe3fd4
@reverseengine
https://medium.com/@jehadbudagga/reverse-engineering-a-0day-used-against-crowdstrike-edr-a5ea1fbe3fd4
@reverseengine
Medium
Reverse Engineering a 0day used Against EDRs
Hello again…
بخش هفدهم بافر اورفلو
چطور فقط با نگاه کردن به اسمبلی بافر اورفلو پیدا کنیم
قراره چی کار کنیم
تا الان فهمیدیم توابع خطرناک چی هستن و فریم استک چطوری ساخته میشه
الان میخایم یاد بگیریم حتی اگر سورس کد نداشتیم فقط از روی اسمبلی بفهمیم احتمال بافر اورفلو وجود داره یا نه
نشونه اول
وجود بافر روی استک
مثال:
Asm
sub rsp, 0x40
این یعنی 64 بایت فضا روی استک رزرو شده
اگر چند خط پایین تر دیدید
Asm
lea rax, [rbp-0x40]
یا
Asm
lea rcx, [rsp+0x10]
معمولاً با یک بافر طرف هستید
نشونه دوم
ورودی کاربر وارد بافر میشه
مثال
Asm
mov rdi, rax
call gets
یا
Asm
call fgets
یا
Asm
call read
یعنی یک داده از بیرون وارد برنامه شده
هر وقت ورودی دیدید باید حساس بشید
نشونه سوم
کپی بدون بررسی طول
مثال:
Asm
call strcpy
یا
Asm
call strcat
یا
Asm
call sprintf
اینها زنگ خطرهای کلاسیک هستن
چون طول ورودی رو چک نمیکنن
نشونه چهارم
بافر کوچک ورودی بزرگ
مثلا اینو تو دی کامپایلر ببینید
C
char buf[16];
strcpy(buf,input);
یا تو اسمبلی ببینید
Asm
lea rdi,[rbp-0x10]
call strcpy
بافر فقط 16 بایته
ولی هیچ محدودیتی برای input وجود نداره
پس احتمال بافر اورفلو زیاده
نشونه پنجم
نبودن Canary
اگر اول تابع این چیزها رو ندیدید
Asm
mov rax, qword ptr fs:[0x28]
mov [rbp-0x8], rax
احتمالا Stack Canary وجود نداره
وجود این دستورات معمولا نشون میده کامپایلر محافظ استک فعال کرده
نشونه شیشم
تابع قبل از ret هیچ بررسی انجام نمیده
تابع آسیبپذیر معمولا آخرش این شکلیه
Asm
leave
ret
اگر قبل از ret هیچ بررسی امنیتی انجام نشه و بالاتر strcpy دیده باشید باید بیشتر دقت کنید
مثال واقعی تحلیل:
فرض کنید این اسمبلی رو دیدید
Asm
push rbp
mov rbp,rsp
sub rsp,0x20
mov rdx,rdi
lea rax,[rbp-0x10]
mov rdi,rax
call strcpy
leave
ret
سوال
آیا این مشکوکه؟
جواب
بله
چون
Asm
lea rax,[rbp-0x10]
نشون میده بافر 16 بایتی داریم
و
Asm
call strcpy
هم بدون محدودیت داده رو داخلش میریزید
پس اولین چیزی که باید تست کنیم ارسال ورودی طولانیه
تمرین:
یک باینری ساده رو داخل Ghidra یا IDA باز کنید
سه مورد زیر رو پیدا کنید
محل ساخت بافر
محل ورود داده
محل کپی شدن داده
اگر این سه مورد رو پیدا کردید عملا دارید مثل یک Reverse Engineer واقعی فکر میکنید
Part 17 Buffer Overflow
How to find buffer overflow just by looking at the assembly
What are we going to do
So far we have understood what dangerous functions are and how stack frames are created
Now we are going to learn how to find out if there is a buffer overflow even if we don't have the source code just from the assembly
First example
There is a buffer on the stack
Example:
Asm
sub rsp, 0x40
This means that 64 bytes of space on the stack are reserved
If you see a few lines below
Asm
lea rax, [rbp-0x40]
or
Asm
lea rcx, [rsp+0x10]
Usually you are dealing with a buffer
Second example
User input enters the buffer
Example
Asm
mov rdi, rax
call gets
or
Asm
call fgets
or
Asm
call read
This means that data has entered the program from outside
Whenever you see input, you should be sensitive Besh
Third example
Copy without length check
Example:
Asm
call strcpy
or
Asm
call strcat
or
Asm
call sprintf
These are classic alarms
because they don't check the length of the input
Fourth example
Small input buffer, large input
For example, see this in the decompiler
C
char buf[16];
strcpy(buf,input);
Or see in the assembly
Asm
lea rdi,[rbp-0x10]
call strcpy
The buffer is only 16 bytes
But there is no limit for input
So the probability of buffer overflow is high
Fifth example
No Canary
If you did not see these functions first
Asm
mov rax, qword ptr fs:[0x28]
mov [rbp-0x8], rax
There is probably no Stack Canary
The presence of these commands usually indicates that the compiler has enabled stack protection
Sixth example
The function does not perform any checks before ret
A vulnerable function usually ends like this
Asm
leave
ret
If no security checks are performed before ret and you have seen strcpy above, you should be more careful
Real example of analysis:
Suppose you saw this assembly
Asm
push rbp
mov rbp,rsp
sub rsp,0x20
mov rdx,rdi
lea rax,[rbp-0x10]
mov rdi,rax
call strcpy
leave
ret
Question
Is this suspicious?
Answer
Yes
Because
Asm
lea rax,[rbp-0x10]
shows that we have a 16-byte buffer
and
Asm
you are also pouring data into it without any limit
So the first thing we need to test is sending long input
Exercise:
Open a simple binary in Ghidra or IDA
Find the following three things
Where the buffer is created
Where the data is entered
Where the data is copied
If you can find these three things, you are actually thinking like a real Reverse Engineer
@reverseengine
and
Asm
call strcpy
you are also pouring data into it without any limit
So the first thing we need to test is sending long input
Exercise:
Open a simple binary in Ghidra or IDA
Find the following three things
Where the buffer is created
Where the data is entered
Where the data is copied
If you can find these three things, you are actually thinking like a real Reverse Engineer
@reverseengine
سیستم عامل چجوری یک برنامه رو راهاندازی و اجرا میکنه؟
اولین کاری که سیستم عامل برای اجرای برنامه انجام میده آپلود کردن کد اون و هرگونه دیتای استاتیک مثل متغیرهای مقدار دهی اولیه در حافظه در فضای آدرس فراینده برنامه در ابتدا روی دیسک یا در برخی سیستمهای مدرن ssd های مبتنی بر فلش با نوعی فرمت اجرایی قرار داره
How does an operating system launch and execute a program?
The first thing the operating system does to execute a program is to upload its code and any static data, such as initialized variables, into memory in the program's process address space, initially on disk or, in some modern systems, flash-based SSDs in some kind of executable format.
@reverseengine
اولین کاری که سیستم عامل برای اجرای برنامه انجام میده آپلود کردن کد اون و هرگونه دیتای استاتیک مثل متغیرهای مقدار دهی اولیه در حافظه در فضای آدرس فراینده برنامه در ابتدا روی دیسک یا در برخی سیستمهای مدرن ssd های مبتنی بر فلش با نوعی فرمت اجرایی قرار داره
How does an operating system launch and execute a program?
The first thing the operating system does to execute a program is to upload its code and any static data, such as initialized variables, into memory in the program's process address space, initially on disk or, in some modern systems, flash-based SSDs in some kind of executable format.
@reverseengine
توصیف گرها:
به برنامه ها اجازه میده که به راحتی ورودی رو از ترمینال بخونن و خروجی رو روی صفحه نمایش چاپ کنن
Descriptors:
allow programs to easily read input from the terminal and print output to the screen
@reverseengine
به برنامه ها اجازه میده که به راحتی ورودی رو از ترمینال بخونن و خروجی رو روی صفحه نمایش چاپ کنن
Descriptors:
allow programs to easily read input from the terminal and print output to the screen
@reverseengine
فرایند در سه حالت میتونه باشه:
در حال اجرا:
یعنی اینکه یک فرایند روی یک پردازنده در حال اجراست
آماده:
یعنی فرایند آماده اجراست ولی به دلایلی سیستم عامل تصمیم میگیره اونو در این لحظه اجرا نکنه
مسدود شده:
یک فرایند نوعی عملیات انجام داده که باعث میشه تا زمان وقوع رویداد دیگهای آماده اجرا نباشه
A process can be in three states:
Running:
This means that a process is running on a processor
Ready:
This means that the process is ready to run but for some reason the operating system decides not to run it at this time
Blocked:
A process has performed some kind of operation that makes it unavailable for execution until another event occurs
@reverseengine
در حال اجرا:
یعنی اینکه یک فرایند روی یک پردازنده در حال اجراست
آماده:
یعنی فرایند آماده اجراست ولی به دلایلی سیستم عامل تصمیم میگیره اونو در این لحظه اجرا نکنه
مسدود شده:
یک فرایند نوعی عملیات انجام داده که باعث میشه تا زمان وقوع رویداد دیگهای آماده اجرا نباشه
A process can be in three states:
Running:
This means that a process is running on a processor
Ready:
This means that the process is ready to run but for some reason the operating system decides not to run it at this time
Blocked:
A process has performed some kind of operation that makes it unavailable for execution until another event occurs
@reverseengine
بلوک کنترل فرایند (PCB) چیست؟
گاهی اوقات افراد به ساختار منفردی که اطلاعات مربوط به یک فرایند رو ذخیره میکنه بلوک کنترل فرایند میگن
What is a process control block (PCB)?
Sometimes people call a single structure that stores information about a process a process control block
@reverseengine
گاهی اوقات افراد به ساختار منفردی که اطلاعات مربوط به یک فرایند رو ذخیره میکنه بلوک کنترل فرایند میگن
What is a process control block (PCB)?
Sometimes people call a single structure that stores information about a process a process control block
@reverseengine
فراخوانهای سیستمی (System Calls) در لینوکس رابطی هستن که برنامههای کاربر (User Space) از طریق اونا از هسته (Kernel Space) درخواست انجام عملیات میکننن
به زبان ساده:
برنامهها نمیتونن مستقیما به سختافزار فایلها یا حافظه سیستم دسترسی داشته باشن به همین خاطر از System Call استفاده میکنن و از کرنل میخام این کار رو براشون انجام بده
مثال:
وقتی داخل C مینویسید:
در واقع برنامه از کرنل درخواست میکنه:
از فایل بخون
100 بایت داده برگردون
مهمترین System Call های لینوکس
مدیریت فایل
مثال:
مدیریت پردازش
مثال:
یک پردازش جدید (Child Process) میسازه
مدیریت حافظه
مثال:
ارتباط بین پردازشها (IPC)
مثال:
یک سوکت TCP درست میکنه
اطلاعات سیستم
مثال:
شناسه پردازش فعلی رو برمیگردونه
پشت صحنه چه اتفاقی میوفته؟
فرض کنید برنامه:
رو اجرا میکنه
مراحل:
برنامه تابع write() رو صدا میزنه
کتابخانه libc شماره System Call مربوطه رو داخل رجیستر قرار میده
دستور syscall اجرا میشه
CPU
از User Mode به Kernel Mode میره
کرنل تابع sys_write رو اجرا میکنه
نتیجه برگردونده میشه
CPU
دوباره به User Mode برمیگرده
در معماری x86-64 معمولا رجیسترها به این شکل استفاده میشن:
بعد:
اجرا میشه
اسمبلی:
در x86-64:
روی صفحه چاپ میشه
برای مهندسی معکوس اکسپلویتنویسی و تحلیل بدافزار مهمترین System Call هایی که باید خوب بشناسید:
چون تقریبا در همه بدافزارها شلکدها و ابزارهای سطح پایین لینوکس با اینا سر و کار دارید
@reverseengine
به زبان ساده:
برنامهها نمیتونن مستقیما به سختافزار فایلها یا حافظه سیستم دسترسی داشته باشن به همین خاطر از System Call استفاده میکنن و از کرنل میخام این کار رو براشون انجام بده
مثال:
وقتی داخل C مینویسید:
read(fd, buffer, 100); در واقع برنامه از کرنل درخواست میکنه:
از فایل بخون
100 بایت داده برگردون
مهمترین System Call های لینوکس
مدیریت فایل
open() read() write() close() lseek() مثال:
int fd = open("test.txt", O_RDONLY); read(fd, buf, 100); close(fd);
مدیریت پردازش
fork() execve() wait() exit() kill() مثال:
pid_t pid = fork(); یک پردازش جدید (Child Process) میسازه
مدیریت حافظه
mmap() munmap() brk() mprotect() مثال:
یک صفحه حافظه جدید اختصاص میدهmmap(NULL, 4096, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0);
ارتباط بین پردازشها (IPC)
pipe() socket() connect() accept() send() recv() مثال:
socket(AF_INET, SOCK_STREAM, 0); یک سوکت TCP درست میکنه
اطلاعات سیستم
getpid() getuid() uname() time() مثال:
printf("%d\n", getpid());
شناسه پردازش فعلی رو برمیگردونه
پشت صحنه چه اتفاقی میوفته؟
فرض کنید برنامه:
write(1, "Hello", 5); رو اجرا میکنه
مراحل:
برنامه تابع write() رو صدا میزنه
کتابخانه libc شماره System Call مربوطه رو داخل رجیستر قرار میده
دستور syscall اجرا میشه
CPU
از User Mode به Kernel Mode میره
کرنل تابع sys_write رو اجرا میکنه
نتیجه برگردونده میشه
CPU
دوباره به User Mode برمیگرده
در معماری x86-64 معمولا رجیسترها به این شکل استفاده میشن:
RAX = syscall number RDI = arg1 RSI = arg2 RDX = arg3 R10 = arg4 R8 = arg5 R9 = arg6 بعد:
syscall
اجرا میشه
اسمبلی:
mov rax, 1 mov rdi, 1 mov rsi, message mov rdx, 5 syscall
در x86-64:
1 = syscall شماره writeدر نهایت:
rdi=1 یعنی stdout
rsi آدرس رشته
rdx=5 طول رشته
Hello
روی صفحه چاپ میشه
برای مهندسی معکوس اکسپلویتنویسی و تحلیل بدافزار مهمترین System Call هایی که باید خوب بشناسید:
open
read
write
mmap
mprotect
fork
execve
socket
connect
accept
clone
ptrace
kill
چون تقریبا در همه بدافزارها شلکدها و ابزارهای سطح پایین لینوکس با اینا سر و کار دارید
@reverseengine
👍1
System Calls in Linux are the interface through which user space programs request operations from the kernel space
In simple terms:
Programs cannot directly access the hardware files or system memory, so they use System Calls and ask the kernel to do this for them
Example:
When you write in C:
In fact, the program asks the kernel:
Read from file
Return 100 bytes of data
Most important Linux System Calls
File management
Example:
Process Management
Example:
Creates a new process (Child Process)
Memory Management
Example:
Allocates a new memory page
Interprocess Communication (IPC)
Example:
Creates a TCP socket
System Information
Example:
Returns the current process ID
What happens behind the scenes?
Suppose the program:
executes
Steps:
The program calls the write() function
The libc library places the corresponding System Call number into the register
The syscall instruction is executed
The CPU
goes from User Mode to Kernel Mode
The kernel executes the sys_write function
The result is returned
The CPU
returns to User Mode
In the x86-64 architecture, registers are usually used in this way:
Next:
syscall
is executed
Assembly:
In x86-64:
In Finally:
Printed on the screen
For reverse engineering, exploit writing and malware analysis, the most important system calls you should know well are:
Because in almost all malware, shellcodes and low-level Linux tools are involved
@reverseengine
In simple terms:
Programs cannot directly access the hardware files or system memory, so they use System Calls and ask the kernel to do this for them
Example:
When you write in C:
read(fd, buffer, 100);
In fact, the program asks the kernel:
Read from file
Return 100 bytes of data
Most important Linux System Calls
File management
open() read() write() close() lseek()
Example:
int fd = open("test.txt", O_RDONLY); read(fd, buf, 100); close(fd);
Process Management
fork() execve() wait() exit() kill()
Example:
pid_t pid = fork();
Creates a new process (Child Process)
Memory Management
mmap() munmap() brk() mprotect()
Example:
mmap(NULL, 4096, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0);
Allocates a new memory page
Interprocess Communication (IPC)
pipe() socket() connect() accept() send() recv()
Example:
socket(AF_INET, SOCK_STREAM, 0);
Creates a TCP socket
System Information
getpid() getuid() uname() time()
Example:
printf("%d\n", getpid());
Returns the current process ID
What happens behind the scenes?
Suppose the program:
write(1, "Hello", 5);
executes
Steps:
The program calls the write() function
The libc library places the corresponding System Call number into the register
The syscall instruction is executed
The CPU
goes from User Mode to Kernel Mode
The kernel executes the sys_write function
The result is returned
The CPU
returns to User Mode
In the x86-64 architecture, registers are usually used in this way:
RAX = syscall number RDI = arg1 RSI = arg2 RDX = arg3 R10 = arg4 R8 = arg5 R9 = arg6
Next:
syscall
is executed
Assembly:
mov rax, 1 mov rdi, 1 mov rsi, message mov rdx, 5 syscall
In x86-64:
1 = syscall number write
rdi=1 means stdout
rsi is the address of the string
rdx=5 is the length of the string
In Finally:
Hello
Printed on the screen
For reverse engineering, exploit writing and malware analysis, the most important system calls you should know well are:
open
read
write
mmap
mprotect
fork
execve
socket
connect
accept
clone
ptrace
kill
Because in almost all malware, shellcodes and low-level Linux tools are involved
@reverseengine
Fingerprinting Modern C2 Implants Through Runtime Telemetry
https://github.com/threathunters-io/tracebit_x33fcon_2026
Идея конечно интересная, но зачем выкладывать EXE, а не исходник
https://github.com/threathunters-io/tracebit_x33fcon_2026
Идея конечно интересная, но зачем выкладывать EXE, а не исходник
GitHub
GitHub - threathunters-io/kassandra_x33fcon_2026: This repository contains the research tool presented at x33fcon 2026, along with…
This repository contains the research tool presented at x33fcon 2026, along with the associated presentation slides. The content is made available for research and educational purposes. - threathun...
از اینکه تا اینجا از من حمایت کردید خیلی ممنونم من سعی میکنم تمام مطالبی که ارائه میدم بهترین باشن و امیدوارم که از اونا خوشتون بیاد و احساس خوبی داشته باشید وقتی میخونیدشون همه تون رو دوست دارم.
الون 🩶
Thank you so much for supporting me so far. I try to make all the content I post the best and I hope you like it and feel good when you read it. I love you all.
Alone 🖤
الون 🩶
Thank you so much for supporting me so far. I try to make all the content I post the best and I hope you like it and feel good when you read it. I love you all.
Alone 🖤
❤20❤🔥1
Virtualization مجازی سازی
فرض کنید سیستم شما:
8 گیگ رم داره
8 هسته CPU داره
اما همزمان:
مرورگر بازه
تلگرام بازه
موزیک پخش میشه
VS Coode بازه
هر برنامه فکر میکنه:
این CPU مال اونه
این حافظه مال اونه
در حالی که واقعیت این نیست
سیستم عامل منابع واقعی رو بین برنامه ها
تقسیم میکنه و به هر برنامه یک تصویر
مجازی از منابع میده
مثال CPU:
فرض کنید فقط یک CPU دارید
اما همزمان:
Chrome
Telegram
Discord
در حال اجران
سیستم عامل خیلی سریع بین اون جا به جا میشه اینقدر سریع که شما فکر میکنید همه با هم اجرا میشن در حالی که CPU در هر لحظه فقط یک کار انجام میده
مثال حافظه:
هر برنامه فکر میکنه:
Plain text
من از آدرس 0x00000000 شروع میشم
اما در واقعیت همه برنامهها در حافظه واقعی کنار هم قرار دارن
سیستمعامل با Virtual Memory این موضوع رو مدیریت میکنه
چرا برای مهندسی معکوس مهمه؟
چون موضوعاتی مثل:
همه از همین مفهوم Virtualization میان
Virtualization
Suppose your system:
has 8 GB of RAM
has 8 CPU cores
but at the same time:
browser is open
telegram is open
music is playing
vs code is open
each program thinks:
this is its CPU
this is its memory
while in reality this is not
the operating system divides real resources between programs
and gives each program a virtual image
of resources
CPU example:
Suppose you only have one CPU
but at the same time:
Chrome
Telegram
Discord
are running
the operating system switches between them very quickly so fast that you think they are all running together while the CPU is only doing one thing at a time
Memory example:
each program thinks:
plain text
I start at address 0x00000000
but in reality all programs are located together in real memory
Operating system with Virtual Memory handles this
Why is it important for reverse engineering?
Because topics like:
all come from the same concept of Virtualization
@reverseengine
فرض کنید سیستم شما:
8 گیگ رم داره
8 هسته CPU داره
اما همزمان:
مرورگر بازه
تلگرام بازه
موزیک پخش میشه
VS Coode بازه
هر برنامه فکر میکنه:
این CPU مال اونه
این حافظه مال اونه
در حالی که واقعیت این نیست
سیستم عامل منابع واقعی رو بین برنامه ها
تقسیم میکنه و به هر برنامه یک تصویر
مجازی از منابع میده
مثال CPU:
فرض کنید فقط یک CPU دارید
اما همزمان:
Chrome
Telegram
Discord
در حال اجران
سیستم عامل خیلی سریع بین اون جا به جا میشه اینقدر سریع که شما فکر میکنید همه با هم اجرا میشن در حالی که CPU در هر لحظه فقط یک کار انجام میده
مثال حافظه:
هر برنامه فکر میکنه:
Plain text
من از آدرس 0x00000000 شروع میشم
اما در واقعیت همه برنامهها در حافظه واقعی کنار هم قرار دارن
سیستمعامل با Virtual Memory این موضوع رو مدیریت میکنه
چرا برای مهندسی معکوس مهمه؟
چون موضوعاتی مثل:
Stack
Heap
ASLR
Paging
Virtual Memory
همه از همین مفهوم Virtualization میان
Virtualization
Suppose your system:
has 8 GB of RAM
has 8 CPU cores
but at the same time:
browser is open
telegram is open
music is playing
vs code is open
each program thinks:
this is its CPU
this is its memory
while in reality this is not
the operating system divides real resources between programs
and gives each program a virtual image
of resources
CPU example:
Suppose you only have one CPU
but at the same time:
Chrome
Telegram
Discord
are running
the operating system switches between them very quickly so fast that you think they are all running together while the CPU is only doing one thing at a time
Memory example:
each program thinks:
plain text
I start at address 0x00000000
but in reality all programs are located together in real memory
Operating system with Virtual Memory handles this
Why is it important for reverse engineering?
Because topics like:
Stack
Heap
ASLR
Paging
Virtual Memory
all come from the same concept of Virtualization
@reverseengine
Devirtualization مجازی سازی زدایی
بعضی محافظ ها مثل VMProtect کد اصلی برنامه رو به بایت کد تبدیل میکنن و یک ماشین مجازی داخل برنامه قرار میدن تا اون بایت کد رو اجرا کنه مشکل از جایی شروع میشه که دیگه با یک تابع معمولی طرف نیستیم به جای چند دستور اسمبلی ساده با صدها یا هزار دستور مربوط به ماشین مجازی طرف میشیم
اینجاست که Devirtualization به دردمون میخوره
Devirtualization
یعنی تلاش برای تبدیل منطق مخفی شده داخل ماشین مجازی به شکلی که دوباره قابل فهم باشه
هدف Devirtualization این نیست که محافظ رو حذف کنیم
هدف اینه که بفهمیم برنامه واقعا چه کاری انجام میده
فرض کنید کد اصلی این بوده:
بعد از Virtualization ممکنه این منطق به صد ها دستور بایت کد تبدیل بشه
در ظاهر دیگر هیچ strcmp یا if مشخصی وجود نداره فقط یک Dispatcher و تعداد زیادی Handler دیده میشه
کاری که تحلیلگر انجام میده اینه که به تدریج معنی هر Opcode رو کشف کنه
مثلا متوجه میشه:
بعد از شناسایی Opcode ها میتونیم رفتار ماشین مجازی رو روی کاغذ بازسازی کنیم
به این فرایند ساختن نقشه Opcode ها
معمولا فرایند تحلیل به این شکل پیش میره:
اول Dispatcher پیدا میشه
بعد Handler های مختلف شناسایی میشن
بعد Opcode ها دسته بندی میشن
در نهایت منطق اصلی برنامه بازسازی میشه
یکی از اشتباهات رایج افراد تازهکار اینه که مستقیم سراغ Handler ها میرن
در حالی که اول باید ساختار کلی ماشین مجازی رو بفهمید
اگر Dispatcher رو نفهمید Handler ها تقریبا بی معنی هستن
یک قانون مهم در تحلیل ماشین های مجازی:
اول جریان اجرای بایت کد رو درک کنید
بعد سراغ معنی دستور ها برید
تمرین
یک ماشین مجازی ساده طراحی کنید که فقط سه تا Opcode داشته باشه:
بعد سعی کنید فقط با نگاه کردن به اجرای برنامه بفهمید هر Opcode چه کاری انجام میده
Devirtualization
Some protectors, such as VMProtect, convert the original program code into bytecode and place a virtual machine inside the program to execute that bytecode. The problem starts when we are no longer dealing with a normal function, but instead of a few simple assembly instructions, we are dealing with hundreds or thousands of instructions related to the virtual machine.
This is where Devirtualization comes in.
Devirtualization
is an attempt to convert the hidden logic inside the virtual machine into a form that is understandable again.
The goal of Devirtualization is not to remove the protector. The goal is to understand what the program is actually doing.
Suppose the original code was:
After Virtualization, this logic may be converted into hundreds of bytecode instructions. On the surface, there is no specific strcmp or if, only a Dispatcher and a large number of Handlers. What the analyzer does is to gradually discover the meaning of each Opcode.
For example, it finds:
After identifying the Opcodes, we can reconstruct the behavior of the virtual machine on paper. This process is called creating an Opcode map.
Usually, the analysis process goes like this:
First, the Dispatcher is found.
Then the different Handlers are identified.
Then the Opcodes are categorized.
Finally, the main logic of the program is reconstructed.
One of the common mistakes of beginners is to go straight to the Handlers.
While you should first understand the overall structure of the virtual machine.
If you don't understand the Dispatcher, the Handler are almost meaningless
An important rule in analyzing virtual machines:
First understand the flow of bytecode execution
Then move on to the meaning of the instructions
Exercise
Design a simple virtual machine that has only three Opcodes:
Then try to figure out what each Opcode does just by looking at the program execution
@reverseengine
بعضی محافظ ها مثل VMProtect کد اصلی برنامه رو به بایت کد تبدیل میکنن و یک ماشین مجازی داخل برنامه قرار میدن تا اون بایت کد رو اجرا کنه مشکل از جایی شروع میشه که دیگه با یک تابع معمولی طرف نیستیم به جای چند دستور اسمبلی ساده با صدها یا هزار دستور مربوط به ماشین مجازی طرف میشیم
اینجاست که Devirtualization به دردمون میخوره
Devirtualization
یعنی تلاش برای تبدیل منطق مخفی شده داخل ماشین مجازی به شکلی که دوباره قابل فهم باشه
هدف Devirtualization این نیست که محافظ رو حذف کنیم
هدف اینه که بفهمیم برنامه واقعا چه کاری انجام میده
فرض کنید کد اصلی این بوده:
C
if(password == "1234") { success(); }
بعد از Virtualization ممکنه این منطق به صد ها دستور بایت کد تبدیل بشه
در ظاهر دیگر هیچ strcmp یا if مشخصی وجود نداره فقط یک Dispatcher و تعداد زیادی Handler دیده میشه
کاری که تحلیلگر انجام میده اینه که به تدریج معنی هر Opcode رو کشف کنه
مثلا متوجه میشه:
Opcode 0x01 دادهای رو داخل رجیستر مجازی اپلود میکنه
Opcode 0x02 دو مقدار رو مقایسه میکنه
Opcode 0x03 یک پرش شرطی انجام میده
بعد از شناسایی Opcode ها میتونیم رفتار ماشین مجازی رو روی کاغذ بازسازی کنیم
به این فرایند ساختن نقشه Opcode ها
معمولا فرایند تحلیل به این شکل پیش میره:
اول Dispatcher پیدا میشه
بعد Handler های مختلف شناسایی میشن
بعد Opcode ها دسته بندی میشن
در نهایت منطق اصلی برنامه بازسازی میشه
یکی از اشتباهات رایج افراد تازهکار اینه که مستقیم سراغ Handler ها میرن
در حالی که اول باید ساختار کلی ماشین مجازی رو بفهمید
اگر Dispatcher رو نفهمید Handler ها تقریبا بی معنی هستن
یک قانون مهم در تحلیل ماشین های مجازی:
اول جریان اجرای بایت کد رو درک کنید
بعد سراغ معنی دستور ها برید
تمرین
یک ماشین مجازی ساده طراحی کنید که فقط سه تا Opcode داشته باشه:
LOAD
ADD
بعد سعی کنید فقط با نگاه کردن به اجرای برنامه بفهمید هر Opcode چه کاری انجام میده
Devirtualization
Some protectors, such as VMProtect, convert the original program code into bytecode and place a virtual machine inside the program to execute that bytecode. The problem starts when we are no longer dealing with a normal function, but instead of a few simple assembly instructions, we are dealing with hundreds or thousands of instructions related to the virtual machine.
This is where Devirtualization comes in.
Devirtualization
is an attempt to convert the hidden logic inside the virtual machine into a form that is understandable again.
The goal of Devirtualization is not to remove the protector. The goal is to understand what the program is actually doing.
Suppose the original code was:
C
if(password == "1234") { success(); }
After Virtualization, this logic may be converted into hundreds of bytecode instructions. On the surface, there is no specific strcmp or if, only a Dispatcher and a large number of Handlers. What the analyzer does is to gradually discover the meaning of each Opcode.
For example, it finds:
Opcode 0x01 uploads data into a virtual register
Opcode 0x02 compares two values
Opcode 0x03 performs a conditional jump
After identifying the Opcodes, we can reconstruct the behavior of the virtual machine on paper. This process is called creating an Opcode map.
Usually, the analysis process goes like this:
First, the Dispatcher is found.
Then the different Handlers are identified.
Then the Opcodes are categorized.
Finally, the main logic of the program is reconstructed.
One of the common mistakes of beginners is to go straight to the Handlers.
While you should first understand the overall structure of the virtual machine.
If you don't understand the Dispatcher, the Handler are almost meaningless
An important rule in analyzing virtual machines:
First understand the flow of bytecode execution
Then move on to the meaning of the instructions
Exercise
Design a simple virtual machine that has only three Opcodes:
LOAD
ADD
Then try to figure out what each Opcode does just by looking at the program execution
@reverseengine
بخش هجدهم بافر اورفلو
انواع Buffer Overflow
خیلیها وقتی اسم Buffer Overflow میاد فقط یاد استک میفتن در صورتی که بافر اورفلو فقط Stack Overflow نیست
چندین نوع مختلف داره که هر کدوم رفتار و اثر متفاوتی دارن
نوع اول
Stack Buffer Overflow
معروفترین نوع
وقتی اتفاق میفته که داده بیشتر از ظرفیت یک بافر روی استک نوشته بشه
مثال:
void vuln(char *input) { char buf[16]; strcpy(buf,input); }
اینجا اگر بیشتر از 16 بایت ارسال بشه
داده از بافر خارج میشه
و کم کم به
متغیرهای محلی
Saved RBP
Return Address
میرسه
همون چیزی که تا الان یاد گرفتیم
نوع دوم
Heap Buffer Overflow
این یکی روی Heap اتفاق میفته
نه روی Stack
مثال:
char *buf = malloc(16); strcpy(buf,input);
اگر ورودی بیشتر از 16 بایت باشه
حافظههای کنار این Chunk خراب میشن
چرا مهمه؟
چون روی Heap معمولا خبری از Return Address نیست پس هدف مهاجم بیشتر خراب کردن ساختار Heap یا دستکاری Pointer هاست
نوع سوم
Off By One
یکی از باگهای مورد علاقه مهندس های معکوسه
چون ظاهرا کوچیک به نظر میاد
ولی گاهی به اکسپلویت کامل تبدیل میشه
مثال:
char buf[16]; for(int i=0;i<=16;i++) { buf[i]='A'; }
اشتباه اینجاست
i <= 16 باید میبود
i < 16 اینجا فقط یک بایت خارج از بافر نوشته میشه ولی همون یک بایت بعضی وقتا برای خراب کردن ساختار حافظه کافیه
نوع چهارم
Integer Overflow
این یکی مستقیم Buffer Overflow نیست
ولی خیلی وقتها باعثش میشه
مثال:
size = count * 100; buf = malloc(size); فرض کن count خیلی بزرگ باشه
در نتیجه ضرب از محدوده نوع داده خارج بشه و size کوچیکتر از چیزی بشه که انتظار داریم بعد برنامه فکر میکنه حافظه زیادی گرفته ولی در واقع حافظه کمی گرفته
و در ادامه Overflow رخ میده
نوع پنجم
Format String
از نظر فنی Buffer Overflow نیست
ولی همیشه کنار این مباحث آموزش داده میشه
مثال:
printf(user_input);
به جای
printf("%s",user_input);
اینجا کاربر میتونه فرمتهای printf رو کنترل کنه
مثل:
%x %x %x %x
و اطلاعات حافظه رو بخونه
برای شما مهمه چون خیلی وقت ها باگ Format String تبدیل به راهی برای دور زدن ASLR یا پیدا کردن آدرس ها میشه
Stack Overflow
خراب کردن استک و Return Address
Heap Overflow
خراب کردن حافظه Heap
Off By One
نوشتن فقط یک بایت اضافه
Integer Overflow
اشتباه در محاسبه اندازه حافظه
Format String
کنترل فرمتهای printf و نشت اطلاعات
وقتی یک باینری رو تحلیل میکنید
سعی کنید تشخیص بدید باگ از کدوم دسته است
چون روش تحلیل Stack Overflow با Heap Overflow کاملا فرق میکنه
و همین تشخیص اولیه خیلی وقتا نصف حل مسئله است
@reverseengine
Part 18 Buffer Overflow
Types of Buffer Overflow
Many people only think of stack when the name Buffer Overflow comes up, but buffer overflow is not just Stack Overflow
It has several different types, each with different behavior and effects
Type 1
Stack Buffer Overflow
The most famous type
It happens when data is written to the stack beyond the capacity of a buffer
Example:
void vuln(char *input) { char buf[16]; strcpy(buf,input); }
Here, if more than 16 bytes are sent
the data is overflowed
and gradually reaches
local variables
Saved RBP
Return Address
The same thing we have learned so far
Type 2
Heap Buffer Overflow
This one happens on the Heap
not on the Stack
Example:
char *buf = malloc(16); strcpy(buf,input);
If the input is more than 16 bytes
The memories next to this Chunk will be corrupted
Why is it important?
Since there is usually no Return Address on the Heap, the attacker's goal is more to corrupt the Heap structure or manipulate the Pointers
The third type
Off By One
It is one of the favorite bugs of reverse engineers
Because it seems small
But sometimes it turns into a full-fledged exploit
Example:
char buf[16]; for(int i=0;i<=16;i++) { buf[i]='A'; }
The error here is
i <= 16
It should be
i < 16
Here only one byte is written outside the buffer, but that one byte is sometimes enough to corrupt the memory structure
The fourth type
Integer Overflow
This one is not a direct Buffer Overflow
But it often causes it
Example:
size = count * 100; buf = malloc(size);
Suppose count is too large
As a result, the multiplication goes out of the data type range and size becomes smaller than we expect. Then the program thinks it has taken up a lot of memory, but in fact it has taken up little memory
And then Overflow occurs
The fifth type
Format String
Is not technically a Buffer Overflow
But it is always taught alongside these topics
Example:
printf(user_input);
Instead of
printf("%s",user_input);
Here the user can control printf formats
Like:
%x %x %x %x
And read memory information
This is important for you because many times the Format String bug becomes a way to bypass ASLR or find addresses
Stack Overflow
Corrupting the stack and Return Address
Heap Overflow
Corrupting the Heap memory
Off By One
Writing only one extra byte
Integer Overflow
Error in calculating the memory size
Format String
Controlling printf formats and information leaks
When you analyze a binary
Try to identify which category the bug belongs to
Because the analysis method for Stack Overflow is completely different from Heap Overflow
And this initial identification is often half the solution to the problem
@reverseengine
I Jailbroke DeepSeek… And It Got SCARY
https://youtu.be/HDkI_z3IQ9s?si=2ou0gw8vJYkVlB7T
https://youtu.be/HDkI_z3IQ9s?si=2ou0gw8vJYkVlB7T
YouTube
I Jailbroke DeepSeek… And It Got SCARY
All demonstrations are intended solely for lawful, ethical, and defensive use. The creator assumes no liability for actions viewers take; attempting to replicate any activity on systems without authorization is illegal and may result in criminal or civil…