CIS 5590 - Vision-Language Models - Fall 2026 - Schedule

Dr. Longin Jan Latecki

Class date Topic Content Presenter
Aug 26 Vision Transformers (ViT)
Oral Questions
Background Knowledge
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale
instructor
Sep 2Adapter-style VLMs Part 1
Oral Questions
Visual Instruction Tuning (LLaVA)
Qwen-VL: A Versatile Vision-Language Model
instructor
Sep 9Large-Language Models (LLMs)
Oral Questions
The Illustrated Transformer
Ɓukasz Kaiser's talk from Oct 2017
instructor
Sep 16Adapter-style VLMs Part 2
Oral Questions
instructor
Sep 23Omni VLMsEmu3.5: Native Multimodal Models are World Learners
Taming Transformers for High-Resolution Image Synthesis (CVPR 2021)
MaskGIT: Masked Generative Image Transformer (CVPR 2022)
instructor
Sep 30 Contrastive Vision-Language Models (VLMs) CLIP: Learning Transferable Visual Models From Natural Language Supervision
Reproducing CLIP from Scratch
Jin, Ying
Oct 7 Steerable Visual Representations Steerable Visual Representations Watkins, Michael
Oct 14 Grounded Vision Models and Explainability Grad-CAM: Visual Explanations from Deep Networks
Grounding DINO: Marrying DINO with Grounded Pre-Training; language-conditioned object queries
Evaluation and datasets: RefCOCO, Visual Genome, COCO-Captions
Che-Wei Hsu
Lee, Ming Chun
Oct 21 VLM Evaluation, Hallucinations, and Reliability HallusionBench; diagnosing visual illusions, language bias, and hallucinations
PROVE: Programmatic VLM Evaluation in the Wild; evaluating truthfulness and helpfulness in open-ended responses
Wu, Fei
Oct 28 Explainability and Grounding in VLMs Why Is Spatial Reasoning Hard for VLMs?
Entropy-Gradient Grounding: Training-Free Evidence Retrieval
Ohene, Anim
Nov 4 Chain of Visual Thoughts Chain-of-Visual-Thought: Teaching VLMs to See and Think
Training Large Language Models to Reason in a Continuous Latent Space
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
Bohn, Mason
Nov 11Spatial ClawSpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Harraq, Othmane
Nov 18 Vision Action Models RT-2: Vision and Language into Action
OpenVLA
Choi, Seunghun
Dec 2Conclusioninstructor
??Thinking with Visual PrimitivesThinking with Visual Primitives
??Explainability in VLMsInterpreting and Editing Vision-Language Representations (ICLR 2025)
Beyond Logit Lens: Contextual Embeddings
Dec 9Final Exam