CIS 5590 - Vision-Language Models - Fall 2026 - Schedule

Dr. Longin Jan Latecki

Class date Topic Content Presenter
Aug 26 Vision Transformers (ViT)
Oral Questions
Background Knowledge
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale
instructor
Sep 2Adapter-style VLMs Part 1Visual Instruction Tuning (LLaVA)
Qwen-VL: A Versatile Vision-Language Model
instructor
Sep 9Large-Language Models (LLMs)
Adapter-style VLMs Part 2
The Illustrated Transformer
Ɓukasz Kaiser's talk from Oct 2017
instructor
Sep ?Omni VLMsEmu3.5: Native Multimodal Models are World Learners
Taming Transformers for High-Resolution Image Synthesis (CVPR 2021)
MaskGIT: Masked Generative Image Transformer (CVPR 2022)
Sep 16 Contrastive Vision-Language Models (VLMs) CLIP: Learning Transferable Visual Models From Natural Language Supervision
Reproducing CLIP from Scratch
Sep 23 Steerable Visual Representations Steerable Visual Representations
Sep 30 Grounded Vision Models and Explainability Grad-CAM: Visual Explanations from Deep Networks
Grounding DINO: Marrying DINO with Grounded Pre-Training; language-conditioned object queries
Evaluation and datasets: RefCOCO, Visual Genome, COCO-Captions
Oct 7 VLM Evaluation, Hallucinations, and Reliability HallusionBench; diagnosing visual illusions, language bias, and hallucinations
PROVE: Programmatic VLM Evaluation in the Wild; evaluating truthfulness and helpfulness in open-ended responses
Oct 14Explainability in VLMsInterpreting and Editing Vision-Language Representations (ICLR 2025)
Beyond Logit Lens: Contextual Embeddings
Oct 21 Explainability and Grounding in VLMs Why Is Spatial Reasoning Hard for VLMs?
Entropy-Gradient Grounding: Training-Free Evidence Retrieval
Oct 28 Chain of Visual Thoughts Chain-of-Visual-Thought: Teaching VLMs to See and Think
Training Large Language Models to Reason in a Continuous Latent Space
CrystaL: Spontaneous Emergence of Visual Latents in MLLMs
Nov 4Thinking with Visual PrimitivesThinking with Visual Primitives
Nov 11Spatial ClawSpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning
Nov 18 Vision Action Models RT-2: Vision and Language into Action
OpenVLA
Dec 2Conclusion
Dec 9Final Exam