Logo Logo
FAQ
Contact
Switch language to German
Transformers for efficient and high-level image and video understanding
Transformers for efficient and high-level image and video understanding
The foundation of current state-of-the-art deep learning and artificial intelligence architectures is primarily based on the Transformer and its attention mechanism. Unlike traditional convolutional neural network (CNN) approaches, Transformers offer greater scalability, reduced reliance on inductive biases, and excel at processing high-dimensional and long-sequence data such as text, images, and videos. By learning pairwise interactions and correlations among tokens, Transformers capture global context, enabling them to move beyond low-level object classification and detection toward more complex scene understanding. However, the quadratic complexity of the attention mechanism presents significant challenges for computational efficiency, particularly for long sequences, limiting the direct application of standard Transformers to high-level downstream tasks. Importantly, adapting Transformers to these downstream tasks often requires the incorporation of task-specific inductive biases and architectural modifications to capture relevant contextual information effectively. This dissertation addresses both the efficiency challenges and the need for task-specific adaptation by exploring the design and application of efficient and specialized Transformer-based architectures for high-level image and video understanding. In the image domain, we first focus on scene graph generation (SGG), a task that requires comprehensive object and relationship modeling. Relation Transformer introduces a two-stage framework that captures both local and global context for scene graph detection. To further improve efficiency, RelationFormer integrates object and relationship detection into a unified single-stage pipeline, leveraging a novel "[rel]-token" to simultaneously enhance efficiency and accuracy across diverse image domains. Building upon scene graph reasoning, we explore its application to Visual Question Answering (VQA), employing a reinforcement learning-based agent for multi-hop navigation over scene graphs. This approach effectively bridges vision and language, achieving superior performance and enhanced interpretability compared to contemporary methods. We further extend our investigation of Transformer-based models in the image domain to robust out-of-distribution (OOD) detection with OODFormer. By modeling global object context and co-occurrence, OODFormer demonstrates superior performance, especially on challenging near-OOD samples, highlighting its potential for real-world deployment. In the video domain, we propose InstanceFormer, a Transformer-based architecture for Video Instance Segmentation (VIS) that simplifies VIS pipelines by propagating representations from the current frame to the next. The framework introduces specialized components like spatial and semantic prior, memory to capture short-term and long-term temporal dependencies while ensuring temporal coherence. These novelties enable InstanceFormer to achieve state-of-the-art results on challenging datasets such as YouTube-VIS and OVIS. This dissertation explores a series of efficient Transformer paradigms that advance image and video understanding, demonstrating how task-specific adaptations can unlock the potential of Transformers for a range of complex vision tasks, offering scalable, interpretable, and high-performing solutions for future research and applications.
Not available
Koner, Rajat
2025
English
Universitätsbibliothek der Ludwig-Maximilians-Universität München
Koner, Rajat (2025): Transformers for efficient and high-level image and video understanding. Dissertation, LMU München: Faculty of Mathematics, Computer Science and Statistics
[thumbnail of Koner_Rajat.pdf]
Preview
PDF
Koner_Rajat.pdf

28MB

Abstract

The foundation of current state-of-the-art deep learning and artificial intelligence architectures is primarily based on the Transformer and its attention mechanism. Unlike traditional convolutional neural network (CNN) approaches, Transformers offer greater scalability, reduced reliance on inductive biases, and excel at processing high-dimensional and long-sequence data such as text, images, and videos. By learning pairwise interactions and correlations among tokens, Transformers capture global context, enabling them to move beyond low-level object classification and detection toward more complex scene understanding. However, the quadratic complexity of the attention mechanism presents significant challenges for computational efficiency, particularly for long sequences, limiting the direct application of standard Transformers to high-level downstream tasks. Importantly, adapting Transformers to these downstream tasks often requires the incorporation of task-specific inductive biases and architectural modifications to capture relevant contextual information effectively. This dissertation addresses both the efficiency challenges and the need for task-specific adaptation by exploring the design and application of efficient and specialized Transformer-based architectures for high-level image and video understanding. In the image domain, we first focus on scene graph generation (SGG), a task that requires comprehensive object and relationship modeling. Relation Transformer introduces a two-stage framework that captures both local and global context for scene graph detection. To further improve efficiency, RelationFormer integrates object and relationship detection into a unified single-stage pipeline, leveraging a novel "[rel]-token" to simultaneously enhance efficiency and accuracy across diverse image domains. Building upon scene graph reasoning, we explore its application to Visual Question Answering (VQA), employing a reinforcement learning-based agent for multi-hop navigation over scene graphs. This approach effectively bridges vision and language, achieving superior performance and enhanced interpretability compared to contemporary methods. We further extend our investigation of Transformer-based models in the image domain to robust out-of-distribution (OOD) detection with OODFormer. By modeling global object context and co-occurrence, OODFormer demonstrates superior performance, especially on challenging near-OOD samples, highlighting its potential for real-world deployment. In the video domain, we propose InstanceFormer, a Transformer-based architecture for Video Instance Segmentation (VIS) that simplifies VIS pipelines by propagating representations from the current frame to the next. The framework introduces specialized components like spatial and semantic prior, memory to capture short-term and long-term temporal dependencies while ensuring temporal coherence. These novelties enable InstanceFormer to achieve state-of-the-art results on challenging datasets such as YouTube-VIS and OVIS. This dissertation explores a series of efficient Transformer paradigms that advance image and video understanding, demonstrating how task-specific adaptations can unlock the potential of Transformers for a range of complex vision tasks, offering scalable, interpretable, and high-performing solutions for future research and applications.