RIML Lab
3.21K subscribers
46 photos
25 videos
7 files
159 links
Robust and Interpretable Machine Learning Lab,
Prof. Mohammad Hossein Rohban,
Sharif University of Technology

https://youtube.com/@rimllab

twitter.com/MhRohban

https://www.linkedin.com/company/robust-and-interpretable-machine-learning-lab/
Download Telegram
🔐 ML Security Journal Club

This Week's Presentation:

🔹 Title: A Machine Unlearning Approach to Safety Alignment

🔸 Presenter: Arian Komaei

🌀 Abstract:
The paper identifies a fundamental limitation in current vision language model (VLM) alignment called the "safety mirage." Traditional supervised safety fine-tuning often reinforces superficial textual patterns rather than deep harm mitigation, leaving models vulnerable to simple one-word attacks and causing "over-prudence" (unnecessary rejections of benign queries). To address this, the authors propose Machine Unlearning (MU) as a superior alternative. Unlike standard fine-tuning, MU directly removes harmful knowledge and avoids biased feature-label mappings. Extensive evaluations show that MU-based alignment reduces attack success rates by up to 60.27% and cuts unnecessary rejections by over 84.20%, all while preserving the model's general capabilities.

📄 Paper: Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning

Session Details:

* 📅 Date: Thursday پنج شنبه
* 🕒 Time: 9:00 - 10:00 AM
* 🌐 Location: Online at vc.sharif.edu/ch/rohban
We look forward to your participation! ✌️
🔐 LLM Faithfulness Journal Club

This Week's Presentation:

🔹 Title: Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

🔸 Presenter: [Farzan Rahmani](https://www.linkedin.com/in/farzan-rahmani-51128b201)

🌀 Abstract:
AI systems that “think” in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed. Nevertheless, it shows promise and we recommend further research into CoT monitorability and investment in CoT monitoring alongside existing safety methods. Because CoT monitorability may be fragile, we recommend that frontier model developers consider the impact of development decisions on CoT monitorability.

📄 Paper: [Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety](https://arxiv.org/abs/2507.11473v2)

Session Details:

- 📅 Date: Wednesday (چهارشنبه)
- 🕒 Time: 10:00 - 11:00 AM
- 🌐 Location: Online at http://vc.sharif.edu/ch/rohban

We look forward to your participation! ✌️
🚀 Open Research Position: Visual Reasoning in Large Vision-Language Models (LVLMs)

We are looking for motivated students to join our research on visual reasoning in Large Vision-Language Models (LVLMs) at RIML Lab.

🔍 Project Description

Large Vision-Language Models have achieved remarkable performance across a wide range of multimodal tasks. However, their ability to perform complex visual reasoning remains an open challenge. This research focuses on understanding, evaluating, and improving the reasoning capabilities of LVLMs, including multi-step reasoning, visual grounding, and reasoning over complex visual scenes.

📄 Relevant Papers

Question Aware Vision Transformer for Multimodal Reasoning
Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning

🔹 Must-Have Requirements

Strong Python programming skills
Knowledge of deep learning and machine learning fundamentals
Hands-on experience with PyTorch
Familiarity with Vision-Language Models or Large Language Models
Strong research interest and willingness to learn
Ready to start immediately

Workload

Commitment: At least 20 hours per week

📌 Note: Filling out this form does not guarantee acceptance. Only shortlisted candidates will be contacted via email.

🔗 Apply here: Form

💬 Telegram: @Arianaghamohseni

@RIMLLab

#research_position #ML_research #VisionLanguageModels #MultimodalAI #VisualReasoning #DeepLearning
🤖 RL Journal Club

This Week's Presentation:

🔹 Title: Test Time Exploration to Achieve Generalization in Zero-Shot RL
🔸 Presenter: Alireza Farajtabrizi

🌀 Abstract:
This paper studies zero-shot generalization in reinforcement learning, where an agent is trained on a set of tasks but must perform well on unseen test environments. The authors argue that standard reward-maximizing RL agents can overfit to training tasks, especially in environments where simple invariance-based methods fail. Their key insight is that exploration behavior is harder to memorize than reward-seeking behavior and can therefore generalize better.

To build on this idea, the paper introduces Explore to Generalize (ExpGen), an algorithm that combines a maximum-entropy exploration policy with an ensemble of reward-seeking agents. At test time, when the ensemble agrees on an action, the agent exploits that decision; when the ensemble is uncertain, the agent switches to the exploration policy to reach new parts of the state space. Experiments on ProcGen show strong improvements on challenging tasks such as Maze and Heist, setting new state-of-the-art results in several zero-shot RL settings.

📄 Paper: Explore to Generalize in Zero-Shot RL (NeurIPS 2023)

Session Details:

* 📅 Date: Tuesday سه‌شنبه
* 🕒 Time: 15:30 - 16:30
* 🌐 Location: Online at vc.sharif.edu/ch/rohban (http://vc.sharif.edu/ch/rohban)

We look forward to your participation! ✌️
🤖 RL Journal Club

This Week's Presentation:

🔹 Title: Is Reinforcement Learning Really Harder Than Bandits?
🔸 Presenter: Arshia Gharooni

🌀 Abstract:
Episodic reinforcement learning, despite having longer planning horizons, presents little additional sample complexity difficulty compared to contextual bandits, with the proposed Monotonic Value Propagation (MVP) algorithm achieving near-optimal regret bounds. The MVP algorithm utilizes a simplified, variance-aware bonus to achieve superior performance, offering an exponential improvement in horizon dependency and sample efficiency over previous state-of-the-art methods.

📄 Paper: Is Reinforcement Learning More Difficult Than Bandits? A Near-optimal Algorithm Escaping the Curse of Horizon

Session Details:

- 📅 Date: Wednesday (چهارشنبه)
- 🕒 Time: 17:00 - 18:00
- 🌐 Location: Online at http://vc.sharif.edu/ch/rohban

We look forward to your participation! ✌️
🔐 LLM Faithfulness Journal Club

This Week's Presentation:

🔹 Title: Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations

🔸 Presenter: [Farzan Rahmani](https://www.linkedin.com/in/farzan-rahmani-51128b201)

🌀 Abstract:
When large language models explain their decisions, their explanations may sound convincing—but do they faithfully reflect the model's true reasoning? This paper presents a comprehensive counterfactual analysis of self-explanation faithfulness across 75 models from 13 model families. It investigates the tradeoff between concise and comprehensive explanations, introduces two new evaluation metrics (phi-CCT and F-AUROC), and studies how explanation verbosity influences faithfulness measurements. The results reveal a clear scaling trend: larger and more capable language models consistently produce more faithful self-explanations, providing valuable insights into the relationship between model scale, explanation quality, and AI safety.

📄 Paper: [Verbosity Tradeoffs and the Impact of Scale on the Faithfulness of LLM Self-Explanations](https://arxiv.org/abs/2503.13445)

Session Details:

* 📅 Date: Wednesday (چهارشنبه)
* 🕑 Time: 2:00 - 3:00 PM
* 🌐 Location: Online at http://vc.sharif.edu/ch/rohban

We look forward to your participation! ✌️
آزمایشگاه RIML تحت نظارت دکتر رهبان در حال راه‌اندازی یک ژورنال‌کلاب پیرامون یادگیری تقویتی چندعاملی (Multi-Agent RL) بر پایه‌ی کتاب Albrecht با چشم‌انداز حرکت به سمت کار پژوهشی جدی در این حوزه است. در صورت علاقه‌مندی به این مسیر، خواهشمندست این فرم را پر کنید.
*: جلسه‌های ژورنال‌کلاب از این هفته آغاز می‌شود.
در صورتی که پرسش یا ابهامی در این زمینه دارید با شناسه‌ی زیر در تلگرام ارتباط بگیرید:
@Moein_Salimi
📢 Research Collaboration in Quantitative Finance at RIML

We are seeking motivated students interested in quantitative finance, stochastic modeling, machine learning, and portfolio optimization. Selected researchers will work under the supervision of Dr. Rohban and collaborate with international professors and researchers affiliated with the University of Manchester, the Alan Turing Institute, Virginia Tech, and the Technical University of Munich

🔬 The following four research directions are available:

1️⃣ Reinforcement Learning in High-Frequency Market Making
Study the trade-off between time discretization, learning accuracy, and sample complexity in single- and multi-agent market making, including convergence to continuous-time optimal policies and Nash equilibria.
🔗 Paper: https://arxiv.org/abs/2407.21025

2️⃣ Stochastic Optimal Control for Multi-Asset Market Making
Develop scalable closed-form approximations to the Hamilton–Jacobi equations of multi-asset market-making models, enabling interpretable near-optimal quotes under correlated prices and portfolio-wide inventory risk.
🔗 Paper: https://arxiv.org/abs/1810.04383

3️⃣ Diffolio: A Diffusion Model for Multivariate Probabilistic Financial Time-Series Forecasting and Portfolio Construction Model the conditional joint distribution of future asset returns using hierarchical asset-level and market-level attention, and use the generated scenarios for risk-aware portfolio construction.
🔗 Paper: https://arxiv.org/abs/2511.07014

4️⃣ Structured Filtering for Jump-Diffusion Time Series Forecasting to infer hidden market states from partially observed jump-diffusion data and produce calibrated probabilistic forecasts of continuous movements and abrupt price shocks.
🔗 Paper: https://arxiv.org/abs/2605.24548

✉️ Interested candidates are invited to send their CV to:
alirezanourimath@gmail.com
📈 Generative Modeling — Scaling, Multimodality & End-to-End Generation

This Week's Presentation:

🔹 Title: Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation
🔸 Presenter: Amir Qeysarbeigi

🌀 Abstract:
Modern generative models typically handle multimodal distributions by factorizing the generation process into multiple steps, as in autoregressive and diffusion models. While this enables high-quality generation, it creates a mismatch between training and inference and prevents fully end-to-end generation. In this work, we introduce Explorative Modeling (XM), a new paradigm that instead factorizes the training process by exploring multiple candidate generations and training on the best-matching one. This exploration increases generative expressivity, allowing models to capture more modes of multimodal distributions without relying solely on generation factorization.

The paper demonstrates that exploration acts as a third pretraining scaling axis, alongside model parameters and data, improving efficiency across image, video, and language generation. Increasing exploration improves FLOP, sample, and parameter efficiency, with gains that become larger as models and datasets scale. Furthermore, by moving the burden of multimodality from inference-time generation steps to training-time exploration, Explorative Modeling enables end-to-end generative models that can achieve performance comparable to diffusion-based approaches with dramatically fewer inference steps.

📄 Article: *Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation*

Session Details:
We will first review the limitations of conventional reconstructive generative models and introduce the concept of generative expressivity as a fundamental bottleneck in multimodal generation. Then, we will explore how Explorative Modeling replaces generation factorization with training-time exploration, including the Forward and Reverse XM formulations. Finally, we will examine how exploration serves as a new scaling axis, improves efficiency across multiple modalities, and enables end-to-end generation with substantially fewer inference steps.

- 📅 Date: Monday (دوشنبه)
- 🕒 Time: 17:00 - 18:00
- 🌐 **Location: https://vc.sharif.edu/rohban** (Online only)

We look forward to your participation! ✌️
🔐 ML Security Journal Club

This Week's Presentation:

🔹 Title: How Jailbreaks Evade, but Do Not Erase, LLM Safety Mechanisms

🔸 Presenter: Javad Hezareh

🌀 Abstract:
This paper investigates the internal mechanisms of Large Language Models (LLMs) during successful jailbreak attacks. The authors provide mechanistic evidence that jailbreaks do not comprehensively eliminate an LLM's safety features; instead, they selectively suppress specific components to bypass refusal mechanisms, leaving other robust internal safety representations intact. To validate the utility of these mechanistic insights, the authors developed a training-free harmful-content detector. By reading the robust internal activations without any model training, this detector achieves competitive aggregate performance and strong adversarial robustness on safety-eval benchmarks.

📄 Paper: Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models

Session Details:

* 📅 Date: Tuesday, Aug 11
* 🕒 Time: 14:00 - 15:00
* 🌐 Location: Online at vc.sharif.edu/ch/rohban

We look forward to your participation! ✌️
Call for Research Assistants: A Project on Abductive Reasoning in LLMs

If you are familiar with LLMs, you are invited to join our research project as a research assistant. This project focuses on abductive reasoning in LLMs.
This project focuses on abductive reasoning in LLMs and aims at preparing submission for ICLR.
For an introduction to the topic, you can read:
Wiring the ‘Why’: A Unified Taxonomy and Survey of Abductive Reasoning in LLMs
If you are interested, please complete the following form:
Registration Form
📢 Join the IABI TA Team!

🩻 The Intelligent Analysis of Biomedical Images (IABI) course is looking for motivated Bachelor’s and Master’s students to join its Teaching Assistant team.

🎯 If you’re interested in biomedical image analysis, enjoy helping others learn, and want to gain valuable teaching experience, we encourage you to apply!

📝 Apply here
Application deadline: 15 September 2026