Bit x Wisdom
48 subscribers
370 photos
31 videos
25 files
342 links
Brains are more complex.
Questions are deeper.
Answers are weirder.
And your mind?
It's about to be blown(⁠◠⁠‿⁠◕⁠)..
------------------------------------
(نا)گفتنی‌هایِ (نا)گاهِ (نا)موجود..
Download Telegram
Naqb
Inhumanist marxism.pdf
امر بیگانه در حال اغوا کردن است. پرسش این است:
آیا با آن خواهیم رقصید؟
Naqb
realism-esarmaye-f@bamun1.pdf
با وجود تمامِ نقداندرنقد‌هایی که به این کتاب دارم ولی خواندش بسیار چسبید.
پیشنهاد میشود.

مارک فیشر یک نابغه بود زین باب که تحلیلش از اینکه چرا افسردگی به بیماریِ همه‌گیرِ جوامعِ غربی تبدیل شده (و چرا بیوشیمیایی‌سازیِ آن به نفعِ شرکت‌های داروسازیِ تمام‌شده(SSRIها🤑))، پرده‌برداری‌اش از بوروکراسیِ عجیبی که نئولیبرال‌ها با اسمِ بازار سرمون درآوردند و مفهومِ ناتوانیِ انعکاسی (==می‌دونیم خرابه، ولی کاری ازمون برنمیاد) چنان تیزبینانه‌س که آدم شاکی میشه چرا خودش نتونست راهِ فراری پیدا کنه(نه که فیشر راه ۱۰۰٪ بگه ولی تلاش‌هایی میکند) و متاسفانه تسلیمِ همون افسردگی شد و ۲۰۱۷ رفت.

تایتل ۹ فصلش رو بخونید مجذوب میشید ولی صرفا چند نقل قول جسته گریخته از کتاب:


تصور پایان جهان از تصور پایان سرمایه‌داری آسان‌تر است.



سرمایه یک انگل مجرد است، یک خون‌آشام سپری‌ناپذیر و یک زامبی‌ساز خستگی‌ناپذیر؛ اما گوشتِ زنده‌ای که به کارِ مرده بدل می‌سازد، گوشتِ تنِ ماست، و زامبی‌هایی که می‌سازد، ما هستیم.

همه‌چیز تسلیم تغییراتِ دائمِ مُد و تصاویرِ رسانه‌ای می‌شود، و [در عینِ حال] دیگر هیچ‌چیز تغییر نمی‌کند.


پس از وضعیتی که در آن هیچ‌چیز ممکن نیست، ناگهان همه‌چیز دوباره ممکن می‌شود.


ما باور داریم که پول تنها یک ژتونِ بی‌معناست که ارزشِ ذاتی ندارد، اما چنان عمل می‌کنیم که گویی واجدِ نوعی ارزشِ قدسی‌ست. و این عملکرد دقیقاً به‌لطف نفی‌وانکار (Fetishist Disavowal) ممکن می‌شود.
گایست در مقام صانع امر مطلق
تأمّلاتی در بابِ خوانش رابرت براندوم بر معمای قصدیّت، از اسپینوزا تا هگل
براندوم در مواجهه با اسپینوزا، کلِّ معماریِ اخلاق را از سطحِ متافیزیک به سطحِ معناشناختی، معرفت‌شناسی و نظریّۀ قصدیّت منتقل می‌کند. براندوم این ادّعای اسپینوزا را به‌رسمیّت می‌شناسد که می‌گوید: یک ایده می‌تواند در ذهنِ من ناکافی باشد در حالی که همان ایده در ذهنِ خدا کافی باشد. مسئلۀ اصلی اما این است که اگر محتوا خاصّیّتِ درونیِ ایده می‌بود، این امر غیرممکن می‌گشت. زیرا یک چیز نمی‌تواند هم‌زمان دارای دو محتوای متفاوت باشد.
📰مطالعۀ متن کامل در وب‌سایت تعمق
#شهام_شریفی
Taamoq | تَعَمُّق✅️
Please open Telegram to view this post
VIEW IN TELEGRAM
Nothing human makes it out of the near-future.


مقاومت بی‌فایده‌ست؛ بهتر است با فرآیند هم‌راستا شوی یا کنار بروی.
عُمَر جَمیل با ویدیو ۱۹ ساعتِ بعد از ۱ سال برگشت
https://youtu.be/XoGvCBRnwLs?si=0SthOvsy0EHTDoXB

ربع ساعت اول بحث غیرفنی‌ست(کجا بودم چکار کردم و چی شد و چرا...و معنی یادگیری، چرا باید هنوز یاد گرفت و مسیر و تجارب خودش و....)


کد:
https://github.com/hkproj/torchfeather



Non exhaustive list of papers cited:

Attention Is All You Need - https://arxiv.org/abs/...​
GShard - https://arxiv.org/abs/...​
Scaling Laws for Fine-Grained Mixture of Experts - https://arxiv.org/abs/...​
GPipe: https://arxiv.org/abs/...​
DeepSeek V2: https://arxiv.org/abs/...​
Auxiliary-Loss-Free Load Balancing Strategy for Mixture-of-Experts - https://arxiv.org/abs/...​
Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM - https://arxiv.org/abs/...​
Fast Transformer Decoding: One Write-Head is All You Need - https://arxiv.org/abs/...​
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints - https://arxiv.org/abs/...​
Zero Bubble Pipeline Parallelism - https://arxiv.org/abs/...​
DeepSeek V3 - https://arxiv.org/abs/...​
Training Compute-Optimal Large Language Models - https://arxiv.org/abs/...​
Scaling Laws for Neural Language Models - https://arxiv.org/abs/...​



Chapters

00:00:00​ - Introduction
00:16:50​ - Model Architecture, Parameters and Training FLOPs
00:51:57​ - RoPE from First Principles
01:11:25​ - Implementing RoPE and YaRN
01:50:16​ - Building the Transformer and Weight Initialization
02:19:12​ - Attention, the KV Cache and Arithmetic Intensity
03:02:56​ - Coding Multi-head Latent Attention (MLA)
03:15:35​ - Block Matrix Multiplication and MLA Internals
03:34:58​ - Deriving MLA and Decoupled RoPE
04:10:54​ - MLA Weight Absorption
04:32:59​ - Autograd and the Mathematics of Distributed Training
05:04:19​ - Distributed Computation Graphs and DDP
05:10:47​ - Building the Training Loop
06:06:17​ - Pipeline Parallelism from First Principles
06:26:52​ - Pipeline Schedules: GPipe, 1F1B and Zero Bubble
07:10:29​ - Datasets, Tokenization and Data Parallelism
07:42:10​ - Coding Pipeline Parallelism
08:37:34​ - Device Meshes and Combining PP with DP
09:48:52​ - Distributed Communication Collectives
10:49:41​ - Implementing Device Meshes, DDP and FSDP
12:13:08​ - Tensor Parallelism from First Principles
13:44:57​ - Coding Tensor Parallelism
14:47:43​ - Context Parallelism and Ring Attention
15:41:22​ - Metrics, Optimizers, Schedulers and Checkpointing
16:00:33​ - Combining Parallelism in the Training Loop
16:49:58​ - Mixture of Experts from First Principles
18:02:52​ - Tensor Parallelism for MoE
18:36:24​ - Expert Parallelism
19:01:56​ - All-to-All Token Dispatch and Combine
19:27:38​ - Expert Tensor Parallelism
Bit x Wisdom
البته، گاهی هم برای حفظ سلامت عقل، از مسیر اصلی خارج می‌شویم.
engrossed elsewhere for the foreseeable, hence my absence. should the day come, i'll resurface to scribe a line or two..

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

MODEL_NAME = "cross-encoder/nli-deberta-v3-base"
DEVICE = torch.device("cuda" if torch.cuda.is_available() else "cpu")

tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
model = AutoModelForSequenceClassification.from_pretrained(MODEL_NAME).to(DEVICE)
model.eval()

FREE_HYPOTHESIS = "The person is free and available."
BUSY_HYPOTHESIS = "The person is busy and unavailable."

ID_TO_LABEL = {
    int(index): label.lower()
    for index, label in model.config.id2label.items()
}

ENTAILMENT_ID = next(
    index for index, label in ID_TO_LABEL.items()
    if "entail" in label
)

CONTRADICTION_ID = next(
    index for index, label in ID_TO_LABEL.items()
    if "contrad" in label
)

ABSENCE_STATE = "BUSY"


@torch.inference_mode()
def _nli_scores(text, hypothesis):
    encoded = tokenizer(
        text,
        hypothesis,
        return_tensors="pt",
        truncation=True,
        max_length=128
    )

    encoded = {
        key: value.to(DEVICE)
        for key, value in encoded.items()
    }

    probabilities = torch.softmax(
        model(**encoded).logits,
        dim=-1
    )[0]

    return (
        probabilities[ENTAILMENT_ID].item(),
        probabilities[CONTRADICTION_ID].item()
    )


def status(text):
    if not isinstance(text, str) or not text.strip():
        return ABSENCE_STATE

    free_entailment, free_contradiction = _nli_scores(
        text,
        FREE_HYPOTHESIS
    )

    busy_entailment, busy_contradiction = _nli_scores(
        text,
        BUSY_HYPOTHESIS
    )

    free_score = free_entailment + busy_contradiction
    busy_score = busy_entailment + free_contradiction

    return "FREE" if free_score > busy_score else ABSENCE_STATE