CIS 5590 - Vision-Language Models - Fall 2026 - Schedule
Dr. Longin Jan Latecki
Class date
Topic
Content
Presenter
Aug 26
Vision Transformers (ViT)
Oral Questions
Background Knowledge
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale
instructor
Sep 2
Adapter-style VLMs
Part 1
Visual Instruction Tuning (LLaVA)
Qwen-VL: A Versatile Vision-Language Model
instructor
Sep 9
Large-Language Models (LLMs)
Adapter-style VLMs
Part 2
The Illustrated Transformer
Ćukasz Kaiser's talk from Oct 2017
instructor
Sep ?
Omni VLMs
Emu3.5: Native Multimodal Models are World Learners
Taming Transformers for High-Resolution Image Synthesis
(CVPR 2021)
MaskGIT: Masked Generative Image Transformer
(CVPR 2022)
Sep 16
Contrastive Vision-Language Models (VLMs)
CLIP: Learning Transferable Visual Models From Natural Language Supervision
Reproducing CLIP from Scratch
Sep 23
Steerable Visual Representations
Steerable Visual Representations
Sep 30
Grounded Vision Models and Explainability
Grad-CAM: Visual Explanations from Deep Networks
Grounding DINO: Marrying DINO with Grounded Pre-Training
; language-conditioned object queries
Evaluation and datasets: RefCOCO, Visual Genome, COCO-Captions
Oct 7
VLM Evaluation, Hallucinations, and Reliability
HallusionBench
; diagnosing visual illusions, language bias, and hallucinations
PROVE: Programmatic VLM Evaluation in the Wild
; evaluating truthfulness and helpfulness in open-ended responses
Oct 14
Explainability in VLMs
Interpreting and Editing Vision-Language Representations
(ICLR 2025)
Beyond Logit Lens: Contextual Embeddings
Oct 21
Explainability and Grounding in VLMs
Why Is Spatial Reasoning Hard for VLMs?
Entropy-Gradient Grounding: Training-Free Evidence Retrieval
Oct 28
Chain of Visual Thoughts
Chain-of-Visual-Thought: Teaching VLMs to See and Think
Training Large Language Models to Reason in a Continuous Latent Space
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
Nov 4
Thinking with Visual Primitives
Thinking with Visual Primitives
Nov 11
Spatial Claw
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Nov 18
Vision Action Models
RT-2: Vision and Language into Action
OpenVLA
Dec 2
Conclusion
Dec 9
Final Exam