Recent studies have shown that visual recognition networks can be fooled by placing objects in inconsistent contexts (e.g., a pig floating in the sky). This lecture covers two representative works modeling the role of contextual information in visual recognition. We systematically investigated critical properties of where, when, and how context modulates recognition.
In the first work, we focused on the study of the amount of context, context and object resolution, geometrical structure of context, context congruence, and temporal dynamics of contextual modulation on real-world images.
In the second work, we explored more challenging properties of contextual modulation including gravity, object co-occurrences and relative sizes in synthetic environments.
In both works, we conducted a series of experiments to gain insights into the impact of contextual cues on both human and machine vision:
Psycho-physics experiments to establish a human benchmark for out-of-context recognition and then compare it with state-of-the-art computer vision models to quantify the gap between the two.
We proposed new context-aware recognition models. The models captured useful information for contextual reasoning, enabling human-level performance and significantly better robustness in out-of-context conditions compared to baseline models across both synthetic and other existing out-of-context natural image datasets.
Lecture slides:
Part 1 - https://drive.google.com/file/d/1l-cP...
Part 2 - https://drive.google.com/file/d/1jziR...
00:00 Part 1 - Intro
02:30 Context-Aware Two-stream Attention Net
05:46 Human psychophysics experiment Experiment
07:48 The object target sizes matter
17:23 The amount of context matters
21:47 Blurred objects led to larger accuracy drop
24:39 Low-level contextual properties do not facilitate
28:32 Incongruent context impairs recognition
30:47 Visualization of predicted attention maps at 8th time step of LSTM
35:26 Ablation study
36:34 Conclusions
45:30 Network Architecture and Q&A
01:00:52 Part 2 - When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes - Intro
01:09:22 What is context?
01:12:53 How can we study context?
01:15:19 Unity and Virtual Home
01:20:48 Context-aware Recognition Transformer Network
01:23:19 Transformer Decoder
01:25:19 Context-aware Recognition Transformer Network
01:32:06 Confidence Modulation
01:33:34 When Is Contextual Information Important?
01:34:49 Context-aware Recognition Transformer Network
01:38:11 Context Matters!
01:43:33 Models Rely on Different Context Aspects
01:45:00 Heterogeneous Impact
01:46:32 CRTNet Beats Baselines in Normal Context
01:49:31 Conclusions and Discussion
[Chapters were auto-generated using our proprietary software - contact us if you are interested in access to the software]
Talk is based on the speakers' papers:
Putting visual object recognition in context (CVPR2020)
Paper: https://arxiv.org/abs/1911.07349
Git: https://github.com/kreimanlab/Put-In-...
When Pigs Fly: Contextual Reasoning in Synthetic and Natural Scenes
Paper: http://arxiv.org/abs/2104.02215
Git: https://github.com/kreimanlab/WhenPig...
*Reference to paper about adversarial attacks using context: https://arxiv.org/abs/1910.00068
Presenter BIO:
Philipp Bomatter is a master student for Computational Science and Engineering at ETH Zurich.
He is interested in artificial intelligence and neuroscience and currently works on a project concerning contextual reasoning in vision at the Kreiman Lab at Harvard University.
Mengmi Zhang completed her PhD in the Graduate School for Integrative Sciences and Engineering, NUS in 2019. She is now a postdoc in KreimanLab in Children's Hospital, Harvard Medical School.
Her research interests include computer vision, machine learning, and cognitive neuroscience. In particular, she studies high-level cognitive functions in humans including attention, memory, learning and reasoning from psychophysics experiments, machine learning approaches and neuroscience.
-------------------------
Find us at:
Newsletter for updates about more events ➜ http://eepurl.com/gJ1t-D
Sub-reddit for discussions ➜ / 2d3dai
Discord server for, well, discord ➜ / discord
Blog ➜ https://2d3d.ai
AI consultancy Abelians ➜ https://abelians.com/