| Class date |
Topic |
Content |
Presenter |
| Aug 26 |
Vision Transformers (ViT)
Oral Questions |
Background Knowledge
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale
|
instructor |
| Sep 2 | Adapter-style VLMs Part 1
Oral Questions |
Visual Instruction Tuning (LLaVA) Qwen-VL: A Versatile Vision-Language Model |
instructor |
| Sep 9 | Large-Language Models (LLMs)
Oral Questions |
The Illustrated Transformer
Ćukasz Kaiser's talk from Oct 2017
|
instructor |
| Sep 16 | Adapter-style VLMs Part 2
Oral Questions |
| instructor |
| Sep 23 | Omni VLMs | Emu3.5: Native Multimodal Models are World Learners
Taming Transformers for High-Resolution Image Synthesis (CVPR 2021) MaskGIT: Masked Generative Image Transformer (CVPR 2022) |
instructor |
| Sep 30 |
Contrastive Vision-Language Models (VLMs) |
CLIP: Learning Transferable Visual Models From Natural Language Supervision
Reproducing CLIP from Scratch
|
Jin, Ying |
| Oct 7 |
Steerable Visual Representations |
Steerable Visual Representations |
Watkins, Michael |
| Oct 14 |
Grounded Vision Models and Explainability |
Grad-CAM: Visual Explanations from Deep Networks
Grounding DINO: Marrying DINO with Grounded Pre-Training; language-conditioned object queries
Evaluation and datasets: RefCOCO, Visual Genome, COCO-Captions
|
Che-Wei Hsu Lee, Ming Chun |
| Oct 21 |
VLM Evaluation, Hallucinations, and Reliability |
HallusionBench; diagnosing visual illusions, language bias, and hallucinations
PROVE: Programmatic VLM Evaluation in the Wild; evaluating truthfulness and helpfulness in open-ended responses
|
Wu, Fei |
| Oct 28 |
Explainability and Grounding in VLMs |
Why Is Spatial Reasoning Hard for VLMs?
Entropy-Gradient Grounding: Training-Free Evidence Retrieval
|
Ohene, Anim |
| Nov 4 |
Chain of Visual Thoughts |
Chain-of-Visual-Thought: Teaching VLMs to See and Think
Training Large Language Models to Reason in a Continuous Latent Space
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
|
Bohn, Mason |
| Nov 11 | Spatial Claw | SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning |
Harraq, Othmane |
| Nov 18 |
Vision Action Models |
RT-2: Vision and Language into Action
OpenVLA
|
Choi, Seunghun |
| Dec 2 | Conclusion | | instructor |
| ?? | Thinking with Visual Primitives | Thinking with Visual Primitives | |
| ?? | Explainability in VLMs | Interpreting and Editing Vision-Language Representations (ICLR 2025) Beyond Logit Lens: Contextual Embeddings | |
| Dec 9 | Final Exam | | |