Defeating AI-Assisted Reverse Engineering (or at Least Trying To)
https://blog.quarkslab.com/defeating-ai-assisted-reverse-engineering-or-at-least-trying-to.html
@reverseengine
https://blog.quarkslab.com/defeating-ai-assisted-reverse-engineering-or-at-least-trying-to.html
@reverseengine
Quarkslab
Defeating AI-Assisted Reverse Engineering (or at Least Trying To) - Quarkslab's blog
Is LLM-assisted reverse engineering making obfuscation pointless? We spent a couple of weeks trying to find out, by handing sandboxed agents a series of progressively hardened AArch64 binaries and one prompt: recover the hidden strings inside. This post walks…
👍1
این مدیوم منه اگه اونجا هم منو فالو کنید خوشحال میشم اونجا هم همین چیزای کانال رو میزارم و ی سری چیزای اضافه🖤
This is my Medium. I would be happy if you followed me there too. I will post the same things from the channel there and a few extra things🩶
https://medium.com/@addcss012
This is my Medium. I would be happy if you followed me there too. I will post the same things from the channel there and a few extra things🩶
https://medium.com/@addcss012
❤8
Dead Code و Junk Code
کدی که هست ولی قرار نیست کاری انجام بده
یکی از روش های رایج Obfuscation اینه که داخل برنامه مقدار زیادی کد اضافه قرار بدن که یا اصلا اجرا نمیشه یا اجرا میشه ولی هیچ تاثیری روی نتیجه نهایی برنامه نداره
هدف اینه که وقتی ما فایل رو باز میکنیم حجم زیادی از دستورهای اضافی ببینیم و پیدا کردن منطق واقعی سختتر بشه
مثلا این کد ساده رو ببینید:
حالا یک مثال در اسمبلی ببینیم:
اگر مقدار
اما Junk Code همیشه به این سادگی نیست ممکنه یک Obfuscator دستورهایی اضافه کنه که ظاهرشون مهم به نظر میرسه:
در ظاهر چند عملیات انجام شده
ولی در اخر وضعیت مهم برنامه تقریبا همون چیزیه که قبل از این بلاک بوده
در تحلیل واقعی یکی از بهترین سوالها اینه:
این بلاک چه چیزی رو تغییر داد که بعدا واقعا استفاده میشه؟
اگر جواب هیچ چیز باشه احتمال داره با Junk Code طرف باشیم یک روش خوب برای تحلیل اینه که فقط مقدار هایی رو دنبال کنید که به خروجی یا مرحله های بعدی برنامه میرسن مثلا اگر یک مقدار داخل
مقدارهای خروجی کجا میرن؟
حافظه تغییر کرده؟
Flag مهمی تغییر کرده؟
تابع دیگه ای صدا زده شده؟
نتیجه این عملیات بعدا استفاده میشه؟
Deobfuscation
یعنی همین کم کم چیزهایی که تاثیری روی منطق اصلی ندارن کنار میرن و ساختار واقعی برنامه مشخص میشه
تمرین:
این کد رو بررسی کنید:
C++
مشخص کنید کدوم قسمت روی خروجی تابع تاثیر داره و کدوم قسمت فقط باعث شلوغ شدن تحلیل میشه بعد همین مثال رو Compile کنید و داخل Ghidra باز کنید ببینید Compiler با بخش اضافی چه کاری میکنه ممکنه حتی قبل از اینکه تو فایل خروجی رو ببینید خودش کل بخش بی استفاده رو حذف کرده باشه چون کامپایلر ها هم بعضی وقتا برخلاف انتظارمون کار مفید انجام میدن
@reverseengine
کدی که هست ولی قرار نیست کاری انجام بده
یکی از روش های رایج Obfuscation اینه که داخل برنامه مقدار زیادی کد اضافه قرار بدن که یا اصلا اجرا نمیشه یا اجرا میشه ولی هیچ تاثیری روی نتیجه نهایی برنامه نداره
هدف اینه که وقتی ما فایل رو باز میکنیم حجم زیادی از دستورهای اضافی ببینیم و پیدا کردن منطق واقعی سختتر بشه
مثلا این کد ساده رو ببینید:
int result = a + b;اینجا محاسبات مربوط به
int x = 50;
x = x * 2;
x = x - 30;
return result;
x هیچ تاثیری روی result نداره پس از نظر منطق برنامه این بخش Dead Code محسوب میشهحالا یک مثال در اسمبلی ببینیم:
mov eax, 10
add eax, 20
mov ecx, 500
xor ecx, ecx
add ecx, 100
ret
اگر مقدار
ecx هیچ جا بعدا استفاده نشه بخش مربوط به ecx عملا تاثیری روی خروجی تابع نداره این همون چیزیه که ما باید یاد بگیریم تشخیص بدیماما Junk Code همیشه به این سادگی نیست ممکنه یک Obfuscator دستورهایی اضافه کنه که ظاهرشون مهم به نظر میرسه:
push rax
xor rcx, rcx
inc rcx
dec rcx
pop rax
در ظاهر چند عملیات انجام شده
ولی در اخر وضعیت مهم برنامه تقریبا همون چیزیه که قبل از این بلاک بوده
در تحلیل واقعی یکی از بهترین سوالها اینه:
این بلاک چه چیزی رو تغییر داد که بعدا واقعا استفاده میشه؟
اگر جواب هیچ چیز باشه احتمال داره با Junk Code طرف باشیم یک روش خوب برای تحلیل اینه که فقط مقدار هایی رو دنبال کنید که به خروجی یا مرحله های بعدی برنامه میرسن مثلا اگر یک مقدار داخل
RAX ساخته بشه ولی قبل از استفاده دوباره overwrite بشه احتمالا محاسبه قبلی اهمیت نداشته ولی اینجا باید حواستون جمع باشه هر کدی که خروجی واضحی نداره Junk Code نیست ممکنه روی Flagها تاثیر بذاره حافظه رو تغییر بده یا اثر جانبی داشته باشه پس قبل از حذف ذهنی یک بلاک رو باید بررسی کنید:مقدارهای خروجی کجا میرن؟
حافظه تغییر کرده؟
Flag مهمی تغییر کرده؟
تابع دیگه ای صدا زده شده؟
نتیجه این عملیات بعدا استفاده میشه؟
Deobfuscation
یعنی همین کم کم چیزهایی که تاثیری روی منطق اصلی ندارن کنار میرن و ساختار واقعی برنامه مشخص میشه
تمرین:
این کد رو بررسی کنید:
C++
int calculate(int a, int b)
{
int x = a + b;
int temp = 500;
temp ^= 123;
temp += 20;
temp -= 20;
return x;
}
مشخص کنید کدوم قسمت روی خروجی تابع تاثیر داره و کدوم قسمت فقط باعث شلوغ شدن تحلیل میشه بعد همین مثال رو Compile کنید و داخل Ghidra باز کنید ببینید Compiler با بخش اضافی چه کاری میکنه ممکنه حتی قبل از اینکه تو فایل خروجی رو ببینید خودش کل بخش بی استفاده رو حذف کرده باشه چون کامپایلر ها هم بعضی وقتا برخلاف انتظارمون کار مفید انجام میدن
@reverseengine
❤1
ReverseEngineering
Dead Code و Junk Code کدی که هست ولی قرار نیست کاری انجام بده یکی از روش های رایج Obfuscation اینه که داخل برنامه مقدار زیادی کد اضافه قرار بدن که یا اصلا اجرا نمیشه یا اجرا میشه ولی هیچ تاثیری روی نتیجه نهایی برنامه نداره هدف اینه که وقتی ما فایل رو باز…
Dead Code and Junk Code
Code that exists but is not supposed to do anything
One of the common methods of Obfuscation is to put a lot of extra code into the program that either does not run at all or runs but has no effect on the final result of the program
The goal is that when we open the file, we see a lot of extra instructions and it becomes harder to find the real logic
For example, look at this simple code:
Here, the calculations related to x have no effect on the result, so from the logic of the program this section is considered Dead Code
Now let's see an example in assembly:
If the value of ecx is not used anywhere later, the part related to ecx has practically no effect on the output of the function. This is what we need to learn to recognize.
But Junk Code is not always this simple. An Obfuscator may add instructions that appear to be important:
In appearance, several operations have been performed
But at the end, the important state of the program is almost the same as before this block.
In real analysis, one of the best questions is:
What did this block change that will actually be used later?
If the answer is nothing, we are probably dealing with Junk Code. A good way to analyze is to only follow the values that reach the output or subsequent stages of the program.
For example, if a value is created in RAX but is overwritten before being used again, the previous calculation probably does not matter. But you have to be careful here. Any code that does not have a clear output is not Junk Code. It may affect flags, change memory, or have side effects. So before mentally deleting a block, you should check:
Where do the output values go?
Has memory changed?
Has an important flag changed?
Has another function been called?
Will the result of this operation be used later?
Deobfuscation
That is, gradually things that do not affect the main logic are removed and the real structure of the program is revealed.
Exercise:
Examine this code:
C++
Determine which part affects the output of the function and which part just makes the analysis busy. Then compile this example and open it in Ghidra and see what the compiler does with the extra part. It may have removed the entire useless part even before you see it in the output file, because compilers sometimes do useful things against our expectations.
@reverseengine
Code that exists but is not supposed to do anything
One of the common methods of Obfuscation is to put a lot of extra code into the program that either does not run at all or runs but has no effect on the final result of the program
The goal is that when we open the file, we see a lot of extra instructions and it becomes harder to find the real logic
For example, look at this simple code:
int result = a + b;
int x = 50;
x = x * 2;
x = x - 30;
return result;
Here, the calculations related to x have no effect on the result, so from the logic of the program this section is considered Dead Code
Now let's see an example in assembly:
mov eax, 10
add eax, 20
mov ecx, 500
xor ecx, ecx
add ecx, 100
ret
If the value of ecx is not used anywhere later, the part related to ecx has practically no effect on the output of the function. This is what we need to learn to recognize.
But Junk Code is not always this simple. An Obfuscator may add instructions that appear to be important:
push rax
xor rcx, rcx
inc rcx
dec rcx
pop rax
In appearance, several operations have been performed
But at the end, the important state of the program is almost the same as before this block.
In real analysis, one of the best questions is:
What did this block change that will actually be used later?
If the answer is nothing, we are probably dealing with Junk Code. A good way to analyze is to only follow the values that reach the output or subsequent stages of the program.
For example, if a value is created in RAX but is overwritten before being used again, the previous calculation probably does not matter. But you have to be careful here. Any code that does not have a clear output is not Junk Code. It may affect flags, change memory, or have side effects. So before mentally deleting a block, you should check:
Where do the output values go?
Has memory changed?
Has an important flag changed?
Has another function been called?
Will the result of this operation be used later?
Deobfuscation
That is, gradually things that do not affect the main logic are removed and the real structure of the program is revealed.
Exercise:
Examine this code:
C++
int calculate(int a, int b)
{
int x = a + b;
int temp = 500;
temp ^= 123;
temp += 20;
temp -= 20;
return x;
}
Determine which part affects the output of the function and which part just makes the analysis busy. Then compile this example and open it in Ghidra and see what the compiler does with the extra part. It may have removed the entire useless part even before you see it in the output file, because compilers sometimes do useful things against our expectations.
@reverseengine
❤1
بخش بیست و هفتم بافر اورفلو
libFuzzer و Coverage Guided
Fuzzing
تا اینجا فهمیدیم Fuzzing یعنی دادن تعداد زیادی ورودی مختلف به برنامه و منتظر موندن تا یک جایی خرابکاری کنه😁
ولی Fuzzer های جدید فقط ورودی رندوم تولید نمیکنن بعضی از اونها بررسی میکنن هر ورودی برنامه رو از چه مسیر هایی عبور داده و همین باعث میشه کم کم ورودی های جالب تر تولید کنه
Coverage Guided یعنی چی
فرض کن یک برنامه این شکلیه
Inputاگر ورودی اول فقط به Check 1 برسه
↓
Check 1
↓
Check 2
↓
Hidden Function
سعی میکنه ورودی مسیر رو تغییر بده تا Fuzzer جدیدی باز بشه
مثلا به Check 2 برسه
بعد دوباره از همون ورودی استفاده میکنه و تغییرات بیشتری میده
هدف اینه که قسمت های بیشتری از برنامه اجرا بشن چون خب ظاهرا ما تصمیم گرفتیم برای پیدا کردن باگ باید به همه جای برنامه سرک بکشیم 😅
libFuzzer
چیکار میکنه
libFuzzer
یک موتور Fuzzing برای برنامههای C و ++C است که با LLVM و Clang کار میکنه ما یک تابع مشخص به اون میدیم
بعد خودش بار ها و بار ها اون تابع رو با ورودی های مختلف اجرا میکنه
هر ورودی که باعث رسیدن به مسیر جدیدی بشه ارزشمند تر میشه
تابع اصلی Fuzzing
معمولا چیزی شبیه این داریم:
C
#include <stdint.h>
#include <stddef.h>
int LLVMFuzzerTestOneInput(
const uint8_t *data,
size_t size
)
{
return 0;
}
توضیح کد زیر
این تابع هدف Fuzzer هست
هر بار libFuzzer یک ورودی جدید تولید میکنه محتوای ورودی داخل
data قرار میگیره و اندازه اون داخل size قرار میگیرهیک مثال ساده:
C
#include <stdint.h>
#include <stddef.h>
#include <string.h>
int LLVMFuzzerTestOneInput(
const uint8_t *data,
size_t size
)
{
if (size >= 5)
{
if (memcmp(data, "HELLO", 5) == 0)
{
volatile int x = 1;
(void)x;
}
}
return 0;
}
اینجا چه اتفاقی میوفته
Fuzzer
ورودی های مختلف رو امتحان میکنه
مثلا
AAAAA
بعد
HELAA
بعد شاید
HELLO
وقتی ورودی به
HELLO برسهیک مسیر جدید از برنامه اجرا میشه
Coverage Guided Fuzzing
این مسیر جدید رو تشخیص میده
و اون ورودی رو نگه میداره تا از اون برای پیدا کردن مسیرهای بعدی استفاده کنه
کامپایل با Clang
shell
clang -g -fsanitize=fuzzer,address file23_fuzz.c -o file23_fuzz
اینجا دو چیز با هم فعال شده
libFuzzer
+
AddressSanitizer
libFuzzer
ورودی تولید میکنه
ASan
مراقب خطا های حافظه هست
این ترکیب برای پیدا کردن Memory Bug خیلی قدرتمنده
اجرای Fuzzer
shell
./file23_fuzz
بعد برنامه شروع میکنه به تولید و تغییر ورودی ها
اگر ورودی باعث کرش بشه معمولا همون ورودی ذخیره میشه
تا بتونیم بعدا دوباره بررسیش کنیم
چرا برای ما مهمه؟
فرض کنید یک برنامه پیچیده دارید
ولی دقیقا نمیدونید چه ورودی باعث رسیدن به یک تابع حساس میشه
Fuzzer
میتونه با امتحان کردن ورودی های مختلف مسیر های جدید رو پیدا کنه
بعد شما میتونید همون مسیر ها رو داخل Ghidra یا IDA بررسی کنید
یعنی:
Fuzzing
↓
New Code Path
↓
Crash یا Behavior
↓
Ghidra / IDA
↓
Assembly Analysis
Coverage Guided Fuzzing
فقط دنبال کرش نیست
دنبال مسیر های جدید هم هست
هر مسیر جدید یعنی بخش جدیدی از برنامه که ارزش بررسی داره و وقتی libFuzzer رو با ASan ترکیب میکنیم هم میتونیم ورودی های هوشمندانه تر تولید کنیم هم Memory Bug ها رو سریع تر تشخیص بدیم
تمرین:
تابع بالا رو کمی تغییر بدید و یک شرط جدید برای یک ورودی خاص اضافه کنید بعد فکر کنید Fuzzer چطور باید قدم به قدم ورودی رو تغییر بده تا به اون مسیر جدید برسه
@reverseengine
❤1
ReverseEngineering
بخش بیست و هفتم بافر اورفلو libFuzzer و Coverage Guided Fuzzing تا اینجا فهمیدیم Fuzzing یعنی دادن تعداد زیادی ورودی مختلف به برنامه و منتظر موندن تا یک جایی خرابکاری کنه😁 ولی Fuzzer های جدید فقط ورودی رندوم تولید نمیکنن بعضی از اونها بررسی میکنن هر ورودی…
Part 27 Buffer Overflow
libFuzzer and Coverage Guided
Fuzzing
So far, we have understood that Fuzzing means giving a lot of different inputs to the program and waiting for it to mess up somewhere😁
But new Fuzzers don't just generate random inputs. Some of them check what paths each input has taken in the program, which makes it gradually generate more interesting inputs
What does Coverage Guided mean
Suppose a program looks like this
Input
↓
Check 1
↓
Check 2
↓
Hidden Function
If the first input only reaches Check 1
It tries to change the input path so that a new Fuzzer opens
For example,
it reaches Check 2
Then it uses the same input again and makes more changes
The goal is to run more parts of the program because apparently we decided to go everywhere in the program to find the bug 😅
libFuzzer
What does libFuzzer do
A Fuzzing Engine for C Programs And it's C++ that works with LLVM and Clang. We give it a specific function.
Then it runs that function over and over again with different inputs.
Each input that leads to a new path becomes more valuable.
The main Fuzzing function
Usually we have something like this:
C
#include <stdint.h>
#include <stddef.h>
int LLVMFuzzerTestOneInput(
const uint8_t *data,
size_t size
)
{
return 0;
}
Explanation of the code below
This function is the target of the Fuzzer
Every time libFuzzer generates a new input, the content of the input is placed in data and its size is placed in size
A simple example:
C
#include <stdint.h>
#include <stddef.h>
#include <string.h>
int LLVMFuzzerTestOneInput(
const uint8_t *data,
size_t size
)
{
if (size >= 5)
{
if (memcmp(data, "HELLO", 5) == 0)
{
volatile int x = 1;
(void)x;
}
}
return 0;
}
What happens here
The Fuzzer
trys different inputs
For example
AAAA
then
HELAA
then maybe
HELLO
When the input reaches HELLO
a new path is executed from the program
Coverage Guided Fuzzing
detects this new path
and stores that input to use for finding subsequent paths
Compile with Clang:
shell
clang -g -fsanitize=fuzzer,address file23_fuzz.c -o file23_fuzz
Here two things are enabled together
libFuzzer
+
AddressSanitizer
libFuzzer
generates input
ASan
watches for memory errors
This combination is very powerful for finding memory bugs
Run Fuzzer
shell
./file23_fuzz
Then the program starts generating and modifying inputs
If the input causes a crash, usually the same input is saved
so that we can check it again later Let's do
Why is it important to us?
Suppose you have a complex program
But you don't know exactly what input will cause a critical function to be reached
Fuzzer
can find new paths by trying different inputs
Then you can examine those paths in Ghidra or IDA
That is:
Fuzzing
↓
New Code Path
↓
Crash or Behavior
↓
Ghidra / IDA
↓
Assembly Analysis
Coverage Guided Fuzzing
It doesn't just look for crashes
It also looks for new paths
Each new path is a new part of the program that is worth examining and when we combine libFuzzer with ASan we can both generate smarter inputs and detect memory bugs faster
Exercise:
Change the above function a little and add a new condition for a specific input then think about how the Fuzzer should change the input step by step to reach that new path
@reverseengine
❤1
Process Injection from the perspective of an attacker and EDR
https://medium.com/@addcss012/process-injection-from-the-perspective-of-an-attacker-and-edr-40e927f4619e
@reverseengine
https://medium.com/@addcss012/process-injection-from-the-perspective-of-an-attacker-and-edr-40e927f4619e
@reverseengine
Medium
Process Injection from the perspective of an attacker and EDR
Process Injection
The Domain Generation Algorithm of Orchard v3
https://bin.re/blog/a-dga-seeded-by-the-bitcoin-genesis-block
@reverseengine
https://bin.re/blog/a-dga-seeded-by-the-bitcoin-genesis-block
@reverseengine
Binary Reverse Engineering Blog
The Domain Generation Algorithm of Orchard v3 - A DGA Seeded by the Bitcoin Genesis Block
The Orchard malware uses a domain generation algorithm (DGA) that is seeded both by the current date, and also by the current balance of the Bitcoin genesis block.
Defeating AI-Assisted Reverse Engineering (or at Least Trying To)
https://blog.quarkslab.com/defeating-ai-assisted-reverse-engineering-or-at-least-trying-to.html
@reverseengine
https://blog.quarkslab.com/defeating-ai-assisted-reverse-engineering-or-at-least-trying-to.html
@reverseengine
Quarkslab
Defeating AI-Assisted Reverse Engineering (or at Least Trying To) - Quarkslab's blog
Is LLM-assisted reverse engineering making obfuscation pointless? We spent a couple of weeks trying to find out, by handing sandboxed agents a series of progressively hardened AArch64 binaries and one prompt: recover the hidden strings inside. This post walks…
Vtable
چیه و چرا داخل ++C اهمیت داره؟
پست قبل دربارهی Function Pointer صحبت کردیم و دیدیم که برنامه میتونه آدرس یک تابع رو نگه داره و بعدا از طریق اون تابعی رو اجرا کنه
حالا در ++C یک مفهوم مشابه ولی پیچیده تر داریم:
Vtable
یا:
Virtual Table
این موضوع برای فهم ساختار داخلی Object های C++ خیلی مهمه
Virtual Function
در ++C ممکنه یک کلاس تابعی داشته باشه که با کلمهی virtual تعریف شده:
class Animal {
public:
virtual void speak() {
std::cout << "Animal";
}
};
حالا یک کلاس دیگه از اون ارثبری میکنه:
class Dog : public Animal {
public:
void speak() override {
std::cout << "Dog";
}
};
اینجا نکته اینه که اگر با یک Pointer از نوع Animal به یک Object از نوع Dog اشاره کنیم برنامه باید بدونه در نهایت کدوم نسخه از ()speak اجرا میشه
یعنی:
Animal Pointer
│
▼
Dog Object
│
▼
Dog::speak()
برای حل این مسئله معمولا از مکانیزمی مثل Vtable استفاده میکنیم
Vtable
رو میتونیم یک جدول در نظر بگیریم که شامل آدرس توابع Virtual یک کلاس هست
به شکل ساده:
Vtable
+----------------------+
| &function_1 |
+----------------------+
| &function_2 |
+----------------------+
| &function_3 |
+----------------------+
هر Object از کلاسی که Virtual Function دارده معمولا به این جدول مرتبطه
برای این ارتباط کامپایلر معمولا چیزی شبیه یک Pointer داخلی به نام:
vptr
داخل Object قرار میده
پس یک Object ممکنه به شکل مفهومی این طوری باشه:
Object
+----------------------+
| vptr --------------+----------+
+----------------------+ |
| Data | ▼
| Data | Vtable
+----------------------+ +------------------+
| &Virtual Function|
+------------------+
| &Virtual Function|
+------------------+
vptr به Vtable
مربوط به نوع واقعی Object اشاره میکنه
وقتی Virtual Function صدا زده میشه چه اتفاقی میوفته؟
فرض کنید این کد رو داریم:
Animal *animal = new Dog();
animal->speak();
برنامه نمیتونه فقط بر اساس نوع Pointer تصمیم بگیره چون نوع واقعی Object ممکنه Dog باشه
پس بهصورت مفهومی این مسیر رو طی میکنه:
animal
│
▼
Object
│
▼
vptr
│
▼
Vtable
│
▼
آدرس speak()
│
▼
اجرای تابع
یعنی برنامه از اطلاعات داخلی Object کمک میگیره تا بفهمه کدوم تابع باید اجرا بشه
چرا Vtable برای Binary Analysis مهمه؟
وقتی یک برنامهی ++C رو Reverse میکنیم Vtable ها میتونن اطلاعات مفیدی درباره ی ساختار برنامه به ما بدن
مثلا با پیدا کردن یک Vtable ممکنه بتونیم حدس بزنیم:
یک Class چه Virtual Function هایی داره
چه Class هایی به هم مرتبطن
ساختار Object تقریبا چجوریه
کدوم توابع به یک Class مربوط میشن
برای همین Vtable ها در Reverse برنامههای ++C اهمیت زیادی دارن
بعضی وقتا اسم توابع از بین رفته و Symbol ها هم وجود ندارن چون انسان ها ظاهرا تصمیم گرفتن برنامهها رو بدون برچسب ول بکنن در این شرایط ساختار هایی مثل Vtable میتونن سرنخ مهمی باشن
ارتباط Vtable با Memory Corruption
نکتهی مهم اینجاست که vptr خودش یک Pointer هست
یعنی Object ممکنه چیزی شبیه این داشته باشه:
Object
+------------------+
| vptr | ← Pointer
+------------------+
| member_1 |
+------------------+
| member_2 |
+------------------+
اگر یک آسیب پذیری حافظه باعث خراب شدن دادههای Object بشه در تحلیل باید بررسی کنیم:
کدوم قسمتهای Object تحت تاثیر قرار گرفتن؟
اگر یک Pointer مهم در ساختار Object قرار داشته باشه تغییر دادن اون میتونه رفتار بعدی برنامه رو تغییر بده
اما باز هم باید تفاوت مهمی رو یادمون باشه:
خراب شدن یک Object لزوما به معنی کنترل اجرای برنامه نیست
باید بررسی کنیم برنامه بعدا با دادهی خراب شده چکاری انجام میده و چه مکانیزم های محافظتی وجود داره
Vtable با Function Pointer
چه فرقی داره؟
از نظر مفهوم هر دو به آدرس توابع مربوطن
اما تفاوت دارن
Function Pointer
یک متغیر مستقیما آدرس یک تابع رو نگه میداره:
Function Pointer
│
▼
Address of Code
Vtable
چیه و چرا داخل ++C اهمیت داره؟
پست قبل دربارهی Function Pointer صحبت کردیم و دیدیم که برنامه میتونه آدرس یک تابع رو نگه داره و بعدا از طریق اون تابعی رو اجرا کنه
حالا در ++C یک مفهوم مشابه ولی پیچیده تر داریم:
Vtable
یا:
Virtual Table
این موضوع برای فهم ساختار داخلی Object های C++ خیلی مهمه
Virtual Function
در ++C ممکنه یک کلاس تابعی داشته باشه که با کلمهی virtual تعریف شده:
class Animal {
public:
virtual void speak() {
std::cout << "Animal";
}
};
حالا یک کلاس دیگه از اون ارثبری میکنه:
class Dog : public Animal {
public:
void speak() override {
std::cout << "Dog";
}
};
اینجا نکته اینه که اگر با یک Pointer از نوع Animal به یک Object از نوع Dog اشاره کنیم برنامه باید بدونه در نهایت کدوم نسخه از ()speak اجرا میشه
یعنی:
Animal Pointer
│
▼
Dog Object
│
▼
Dog::speak()
برای حل این مسئله معمولا از مکانیزمی مثل Vtable استفاده میکنیم
Vtable
رو میتونیم یک جدول در نظر بگیریم که شامل آدرس توابع Virtual یک کلاس هست
به شکل ساده:
Vtable
+----------------------+
| &function_1 |
+----------------------+
| &function_2 |
+----------------------+
| &function_3 |
+----------------------+
هر Object از کلاسی که Virtual Function دارده معمولا به این جدول مرتبطه
برای این ارتباط کامپایلر معمولا چیزی شبیه یک Pointer داخلی به نام:
vptr
داخل Object قرار میده
پس یک Object ممکنه به شکل مفهومی این طوری باشه:
Object
+----------------------+
| vptr --------------+----------+
+----------------------+ |
| Data | ▼
| Data | Vtable
+----------------------+ +------------------+
| &Virtual Function|
+------------------+
| &Virtual Function|
+------------------+
vptr به Vtable
مربوط به نوع واقعی Object اشاره میکنه
وقتی Virtual Function صدا زده میشه چه اتفاقی میوفته؟
فرض کنید این کد رو داریم:
Animal *animal = new Dog();
animal->speak();
برنامه نمیتونه فقط بر اساس نوع Pointer تصمیم بگیره چون نوع واقعی Object ممکنه Dog باشه
پس بهصورت مفهومی این مسیر رو طی میکنه:
animal
│
▼
Object
│
▼
vptr
│
▼
Vtable
│
▼
آدرس speak()
│
▼
اجرای تابع
یعنی برنامه از اطلاعات داخلی Object کمک میگیره تا بفهمه کدوم تابع باید اجرا بشه
چرا Vtable برای Binary Analysis مهمه؟
وقتی یک برنامهی ++C رو Reverse میکنیم Vtable ها میتونن اطلاعات مفیدی درباره ی ساختار برنامه به ما بدن
مثلا با پیدا کردن یک Vtable ممکنه بتونیم حدس بزنیم:
یک Class چه Virtual Function هایی داره
چه Class هایی به هم مرتبطن
ساختار Object تقریبا چجوریه
کدوم توابع به یک Class مربوط میشن
برای همین Vtable ها در Reverse برنامههای ++C اهمیت زیادی دارن
بعضی وقتا اسم توابع از بین رفته و Symbol ها هم وجود ندارن چون انسان ها ظاهرا تصمیم گرفتن برنامهها رو بدون برچسب ول بکنن در این شرایط ساختار هایی مثل Vtable میتونن سرنخ مهمی باشن
ارتباط Vtable با Memory Corruption
نکتهی مهم اینجاست که vptr خودش یک Pointer هست
یعنی Object ممکنه چیزی شبیه این داشته باشه:
Object
+------------------+
| vptr | ← Pointer
+------------------+
| member_1 |
+------------------+
| member_2 |
+------------------+
اگر یک آسیب پذیری حافظه باعث خراب شدن دادههای Object بشه در تحلیل باید بررسی کنیم:
کدوم قسمتهای Object تحت تاثیر قرار گرفتن؟
اگر یک Pointer مهم در ساختار Object قرار داشته باشه تغییر دادن اون میتونه رفتار بعدی برنامه رو تغییر بده
اما باز هم باید تفاوت مهمی رو یادمون باشه:
خراب شدن یک Object لزوما به معنی کنترل اجرای برنامه نیست
باید بررسی کنیم برنامه بعدا با دادهی خراب شده چکاری انجام میده و چه مکانیزم های محافظتی وجود داره
Vtable با Function Pointer
چه فرقی داره؟
از نظر مفهوم هر دو به آدرس توابع مربوطن
اما تفاوت دارن
Function Pointer
یک متغیر مستقیما آدرس یک تابع رو نگه میداره:
Function Pointer
│
▼
Address of Code
Vtable
❤2
یک Object معمولا اول به یک جدول اشاره میکنه و اون جدول شامل آدرس توابع هست:
Object
│
▼
vptr
│
▼
Vtable
│
├── Function A
├── Function B
└── Function C
بنابراین Vtable یک سطح ساختارمند تر برای مدیریت توابع Virtual فراهم میکنه
Vtable
دقیقاً یک استاندارد رسمی ++C هست؟
یک نکته ی مهم:
خود مفهوم Vtable به شکل دقیق در استاندارد C++ تعریف نشده
Vtable
در واقع یک روش رایج برای پیاده سازی Polymorphism توسط کامپایلر هاست
یعنی ممکنه جزئیات پیادهسازی بین:
GCC
Clang
MSVC
متفاوت باشه
اما ایدهی کلی یعنی استفاده از اطلاعاتی برای پیدا کردن تابع درست در زمان اجرا در عمل بسیار رایجه
برای همین هنگام Reverse باید همیشه به ABI و کامپایلری که برنامه با اون ساخته شده توجه کنیم
Vtable
معمولا جدولی شامل آدرس Virtual Function های یک کلاس در ++C هست Object هایی که از Virtual Function استفاده میکنن اغلب یک Pointer داخلی مثل vptr دارن که اونها رو به Vtable مربوط میکنه موقع اجرای یک Virtual Function برنامه از این ساختار برای پیدا کردن تابع مناسب استفاده میکنه
شناخت Vtable هم در Reverse کردن برنامه های ++C مهمه و هم برای درک اینکه Object ها در حافظه چجوری سازماندهی میشن
@reverseengine
Object
│
▼
vptr
│
▼
Vtable
│
├── Function A
├── Function B
└── Function C
بنابراین Vtable یک سطح ساختارمند تر برای مدیریت توابع Virtual فراهم میکنه
Vtable
دقیقاً یک استاندارد رسمی ++C هست؟
یک نکته ی مهم:
خود مفهوم Vtable به شکل دقیق در استاندارد C++ تعریف نشده
Vtable
در واقع یک روش رایج برای پیاده سازی Polymorphism توسط کامپایلر هاست
یعنی ممکنه جزئیات پیادهسازی بین:
GCC
Clang
MSVC
متفاوت باشه
اما ایدهی کلی یعنی استفاده از اطلاعاتی برای پیدا کردن تابع درست در زمان اجرا در عمل بسیار رایجه
برای همین هنگام Reverse باید همیشه به ABI و کامپایلری که برنامه با اون ساخته شده توجه کنیم
Vtable
معمولا جدولی شامل آدرس Virtual Function های یک کلاس در ++C هست Object هایی که از Virtual Function استفاده میکنن اغلب یک Pointer داخلی مثل vptr دارن که اونها رو به Vtable مربوط میکنه موقع اجرای یک Virtual Function برنامه از این ساختار برای پیدا کردن تابع مناسب استفاده میکنه
شناخت Vtable هم در Reverse کردن برنامه های ++C مهمه و هم برای درک اینکه Object ها در حافظه چجوری سازماندهی میشن
@reverseengine
❤2
ReverseEngineering
Vtable چیه و چرا داخل ++C اهمیت داره؟ پست قبل دربارهی Function Pointer صحبت کردیم و دیدیم که برنامه میتونه آدرس یک تابع رو نگه داره و بعدا از طریق اون تابعی رو اجرا کنه حالا در ++C یک مفهوم مشابه ولی پیچیده تر داریم: Vtable یا: Virtual Table این…
Vtable
What is it and why is it important in C++?
In the previous post, we talked about Function Pointer and saw that the program can hold the address of a function and later execute a function through it
Now in C++ we have a similar but more complex concept:
Vtable
Or:
Virtual Table
This is very important to understand the internal structure of C++ Objects
Virtual Function
In C++, a class may have a function that is defined with the virtual keyword:
class Animal {
public:
virtual void speak() {
std::cout << "Animal";
}
};
Now another class inherits from it:
class Dog : public Animal {
public:
void speak() override {
std::cout << "Dog";
}
};
The point here is that if we point to an Object of type Dog with a Pointer of type Animal, the program needs to know which version of speak() will be executed in the end
That is:
Animal Pointer
│
▼
Dog Object
│
▼
Dog::speak()
To solve this problem, we usually use a mechanism like Vtable
We can consider Vtable
as a table that contains the addresses of Virtual functions of a class
Simply:
Vtable
+--------------------+
| &function_1 |
+-----------------------+
| &function_2 |
+-----------------------+
| &function_3 |
+-----------------------+
Every Object of a class that has a Virtual Function is usually linked to this table
For this connection, the compiler usually puts something like an internal Pointer called:
vptr
inside the Object
So an Object may conceptually look like this:
Object
+-----------------------+
| vptr --------------+------+
+----------------------+ |
| Data | ▼
| Data | Vtable
+---------+ +------------------+
| &Virtual Function|
+------------------+
| &Virtual Function|
+------------------+
vptr points to the Vtable
of the actual type of Object
What happens when a Virtual Function is called?
Suppose we have this code:
Animal *animal = new Dog();
animal->speak();
The program cannot decide based on the Pointer type alone because the actual type of Object may be Dog
So conceptually it goes like this:
animal
│
▼
Object
│
▼
vptr
│
▼
Vtable
│
▼
speak() address
│
▼
function execution
That is, the program uses the internal information of the Object to understand which function to execute
Why is the Vtable important for Binary Analysis?
When we reverse a C++ program, Vtables can give us useful information about the program structure.
For example, by finding a Vtable, we may be able to guess:
What virtual functions a class has
What classes are related to each other
What is the approximate structure of an object
Which functions are related to a class
That is why Vtables are so important in reversing C++ programs
Sometimes the function names are missing and the symbols are missing because humans apparently decided to leave programs unlabeled. In these situations, structures like Vtables can be important clues
The relationship of Vtables to Memory Corruption
The important point here is that vptr itself is a Pointer
That is, an Object might look something like this:
Object
+------------------+
| vptr | ← Pointer
+------------------+
| member_1 |
+------------------+
| member_2 |
+------------------+
If a memory vulnerability causes Object data corruption, in the analysis we need to consider:
Which parts of the Object are affected?
If there is an important Pointer in the Object structure, changing it can change the subsequent behavior of the program
But we still need to remember an important difference:
Corruption of an Object does not necessarily mean control of program execution
We need to consider what the program does with the corrupted data later and what protection mechanisms are in place
What is the difference between a Vtable and a Function Pointer?
Conceptually, both are related to the address of functions
But they are different
Function Pointer
A variable directly holds the address of a function:
Function Pointer
│
▼
Address of Code
Vtable
What is it and why is it important in C++?
In the previous post, we talked about Function Pointer and saw that the program can hold the address of a function and later execute a function through it
Now in C++ we have a similar but more complex concept:
Vtable
Or:
Virtual Table
This is very important to understand the internal structure of C++ Objects
Virtual Function
In C++, a class may have a function that is defined with the virtual keyword:
class Animal {
public:
virtual void speak() {
std::cout << "Animal";
}
};
Now another class inherits from it:
class Dog : public Animal {
public:
void speak() override {
std::cout << "Dog";
}
};
The point here is that if we point to an Object of type Dog with a Pointer of type Animal, the program needs to know which version of speak() will be executed in the end
That is:
Animal Pointer
│
▼
Dog Object
│
▼
Dog::speak()
To solve this problem, we usually use a mechanism like Vtable
We can consider Vtable
as a table that contains the addresses of Virtual functions of a class
Simply:
Vtable
+--------------------+
| &function_1 |
+-----------------------+
| &function_2 |
+-----------------------+
| &function_3 |
+-----------------------+
Every Object of a class that has a Virtual Function is usually linked to this table
For this connection, the compiler usually puts something like an internal Pointer called:
vptr
inside the Object
So an Object may conceptually look like this:
Object
+-----------------------+
| vptr --------------+------+
+----------------------+ |
| Data | ▼
| Data | Vtable
+---------+ +------------------+
| &Virtual Function|
+------------------+
| &Virtual Function|
+------------------+
vptr points to the Vtable
of the actual type of Object
What happens when a Virtual Function is called?
Suppose we have this code:
Animal *animal = new Dog();
animal->speak();
The program cannot decide based on the Pointer type alone because the actual type of Object may be Dog
So conceptually it goes like this:
animal
│
▼
Object
│
▼
vptr
│
▼
Vtable
│
▼
speak() address
│
▼
function execution
That is, the program uses the internal information of the Object to understand which function to execute
Why is the Vtable important for Binary Analysis?
When we reverse a C++ program, Vtables can give us useful information about the program structure.
For example, by finding a Vtable, we may be able to guess:
What virtual functions a class has
What classes are related to each other
What is the approximate structure of an object
Which functions are related to a class
That is why Vtables are so important in reversing C++ programs
Sometimes the function names are missing and the symbols are missing because humans apparently decided to leave programs unlabeled. In these situations, structures like Vtables can be important clues
The relationship of Vtables to Memory Corruption
The important point here is that vptr itself is a Pointer
That is, an Object might look something like this:
Object
+------------------+
| vptr | ← Pointer
+------------------+
| member_1 |
+------------------+
| member_2 |
+------------------+
If a memory vulnerability causes Object data corruption, in the analysis we need to consider:
Which parts of the Object are affected?
If there is an important Pointer in the Object structure, changing it can change the subsequent behavior of the program
But we still need to remember an important difference:
Corruption of an Object does not necessarily mean control of program execution
We need to consider what the program does with the corrupted data later and what protection mechanisms are in place
What is the difference between a Vtable and a Function Pointer?
Conceptually, both are related to the address of functions
But they are different
Function Pointer
A variable directly holds the address of a function:
Function Pointer
│
▼
Address of Code
Vtable
❤1
An Object usually first points to a table, and that table contains the addresses of functions:
Object
│
▼
vptr
│
▼
Vtable
│
├── Function A
├── Function B
└── Function C
So Vtable provides a more structured level for managing Virtual functions
Vtable
Is it really an official C++ standard?
An important note:
The concept of Vtable itself is not precisely defined in the C++ standard
Vtable
is actually a common way for compilers to implement Polymorphism
That is, the implementation details may differ between:
GCC
Clang
MSVC
But the general idea of using information to find the right function at runtime is very common in practice
That is why when reversing, we should always pay attention to the ABI and the compiler with which the program was built
Vtable
Usually a table containing the addresses of Virtual Functions of a class in C++ Objects that use Virtual Functions often have an internal Pointer such as vptr that relates them to the Vtable. When executing a Virtual Function, the program uses this structure to find the appropriate function
Understanding Vtable is important both in reversing C++ programs and in understanding how Objects are organized in memory
@reverseengine
Object
│
▼
vptr
│
▼
Vtable
│
├── Function A
├── Function B
└── Function C
So Vtable provides a more structured level for managing Virtual functions
Vtable
Is it really an official C++ standard?
An important note:
The concept of Vtable itself is not precisely defined in the C++ standard
Vtable
is actually a common way for compilers to implement Polymorphism
That is, the implementation details may differ between:
GCC
Clang
MSVC
But the general idea of using information to find the right function at runtime is very common in practice
That is why when reversing, we should always pay attention to the ABI and the compiler with which the program was built
Vtable
Usually a table containing the addresses of Virtual Functions of a class in C++ Objects that use Virtual Functions often have an internal Pointer such as vptr that relates them to the Vtable. When executing a Virtual Function, the program uses this structure to find the appropriate function
Understanding Vtable is important both in reversing C++ programs and in understanding how Objects are organized in memory
@reverseengine
❤1
Data Flow Analysis
دنبال کردن مسیر واقعی داده
تا اینجا بیشتر تمرکزمون روی این بود که برنامه چه دستور هایی اجرا میکنه
ولی از اینجا به بعد یک سؤال مهمتر میپرسیم:
داده از کجا میاد و آخرش کجا میره؟
این دقیقا همون چیزیه که بهش Data Flow Analysis میگیم
فرض کنید این کد رو داریم:
C++
اگر فقط به ترتیب دستورها نگاه کنیم میگیم:
اول جمع انجام میشه
بعد ضرب انجام میشه
بعد نتیجه برمیگرده
ولی در Data Flow Analysis این شکلی نگاه میکنیم:
یعنی مسیر خود داده رو دنبال میکنیم
حالا فرض کنید وسط برنامه این دستورها هم وجود داشته باشن:
C++
اگر
اینجا Data Flow Analysis خیلی کمک میکنه Junk Code رو از منطق واقعی جدا کنیم
وقتی داخل IDA یا Ghidra یک تابع پیچیده میبینیم لازم نیست از اول تا آخر همه چیز رو حفظ کنیم
یک مقدار مهم رو انتخاب کنید و دنبالش کنید
مثلا اگر ورودی تابع داخل یک رجیستر وارد شده ببینید:
کجا کپی میشه
کجا تغییر میکنه
داخل حافظه ذخیره میشه یا نه
به تابع دیگه ای میره یا نه
و در اخر روی چه چیزی تأثیر میذاره
مثلا در اسمبلی ممکنه چنین چیزی ببینید:
Asm
اگر فرض کنیم
نکته مهم اینه که اسم رجیستر به معنی مسیر داده نیست ممکنه یک مقدار از
اگر فقط اسم رجیسترها رو نگاه کنید خیلی زود گم میشید😅
باید خود مقدار رو دنبال کنید
به این مفهوم بعصی وقتا Data Provenance هم میگیم
یعنی بفهمید منشا یک مقدار کجاست
مثلا یک مقدار ممکنه از اینجا اومده باشه:
این مدل نگاه مخصوصا موقع تحلیل برنامه های Obfuscate شده خیلی مهمه
چون Obfuscation ممکنه مسیر اجرای برنامه رو شلوغ کنه ولی داده هنوز باید از یک جایی وارد بشه و به یک جایی برسه
تمرین:
این تابع رو بررسی کنید:
C
روی کاغذ مسیر
بعد مشخص کنید
هدف این تمرین اینه که کم کم وقتی یک تابع رو باز میکنید فقط دستورها رو نبینید
داده رو ببینید که داره بین رجیسترها حافظه و توابع حرکت میکنه
@reverseengine
دنبال کردن مسیر واقعی داده
تا اینجا بیشتر تمرکزمون روی این بود که برنامه چه دستور هایی اجرا میکنه
ولی از اینجا به بعد یک سؤال مهمتر میپرسیم:
داده از کجا میاد و آخرش کجا میره؟
این دقیقا همون چیزیه که بهش Data Flow Analysis میگیم
فرض کنید این کد رو داریم:
C++
int calculate(int a, int b)
{
int x = a + b;
int y = x * 2;
return y;
}
اگر فقط به ترتیب دستورها نگاه کنیم میگیم:
اول جمع انجام میشه
بعد ضرب انجام میشه
بعد نتیجه برمیگرده
ولی در Data Flow Analysis این شکلی نگاه میکنیم:
a ─┐
├── ADD ──> x ──> MUL ──> y ──> Return
b ─┘
یعنی مسیر خود داده رو دنبال میکنیم
حالا فرض کنید وسط برنامه این دستورها هم وجود داشته باشن:
C++
int temp = 500;
temp ^= 123;
temp += 20;
اگر
temp هیچوقت روی y یا خروجی تابع تأثیر نذاره توی مسیر اصلی داده قرار نمیگیرهاینجا Data Flow Analysis خیلی کمک میکنه Junk Code رو از منطق واقعی جدا کنیم
وقتی داخل IDA یا Ghidra یک تابع پیچیده میبینیم لازم نیست از اول تا آخر همه چیز رو حفظ کنیم
یک مقدار مهم رو انتخاب کنید و دنبالش کنید
مثلا اگر ورودی تابع داخل یک رجیستر وارد شده ببینید:
کجا کپی میشه
کجا تغییر میکنه
داخل حافظه ذخیره میشه یا نه
به تابع دیگه ای میره یا نه
و در اخر روی چه چیزی تأثیر میذاره
مثلا در اسمبلی ممکنه چنین چیزی ببینید:
Asm
mov eax, edi
add eax, esi
imul eax, 2
اگر فرض کنیم
EDI و ESI ورودی هستن مسیر داده این شکلیه:EDI ─┐
├──> EAX ──> ADD ──> IMUL
ESI ─┘
نکته مهم اینه که اسم رجیستر به معنی مسیر داده نیست ممکنه یک مقدار از
RAX وارد RBX بشه بعد داخل حافظه ذخیره بشه و دوباره داخل RCX برگردهاگر فقط اسم رجیسترها رو نگاه کنید خیلی زود گم میشید😅
باید خود مقدار رو دنبال کنید
به این مفهوم بعصی وقتا Data Provenance هم میگیم
یعنی بفهمید منشا یک مقدار کجاست
مثلا یک مقدار ممکنه از اینجا اومده باشه:
User Input
↓
Buffer
↓
Function A
↓
Function B
↓
Comparison
این مدل نگاه مخصوصا موقع تحلیل برنامه های Obfuscate شده خیلی مهمه
چون Obfuscation ممکنه مسیر اجرای برنامه رو شلوغ کنه ولی داده هنوز باید از یک جایی وارد بشه و به یک جایی برسه
تمرین:
این تابع رو بررسی کنید:
C
int process(int a)
{
int x = a;
x = x ^ 0x55;
x = x + 10;
int junk = 100;
junk *= 5;
return x;
}
روی کاغذ مسیر
a تا خروجی رو بکشیدبعد مشخص کنید
junk وارد مسیر اصلی داده میشه یا نههدف این تمرین اینه که کم کم وقتی یک تابع رو باز میکنید فقط دستورها رو نبینید
داده رو ببینید که داره بین رجیسترها حافظه و توابع حرکت میکنه
@reverseengine
❤2