Koner, Rajat (2025): Transformers for efficient and high-level image and video understanding. Dissertation, LMU München: Faculty of Mathematics, Computer Science and Statistics