Portrait of Adrian Bulat

Adrian Bulat

Adrian Bulat

Research ScientistSamsung AI Center Cambridge

Adrian Bulat is a Research Scientist at Samsung AI Cambridge. Previously, he received his PhD from the University of Nottingham where he worked with Dr. Georgios Tzimiropoulos as part of the Computer Vision Laboratory. His current research interests lie at the intersection of Computer Vision and Machine Learning.

Selected publications

Full publication list
  • 2026

    VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions New

    Adrian Bulat*, Alberto Baldrati*, Ioannis Maniadis Metaxas*, Yassine Ouali, Georgios Tzimiropoulos

    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

    [abstract]

    Existing approaches for improving the efficiency of Large Vision-Language Models (LVLMs) are largely based on the concept of visual token reduction. This approach, however, creates an information bottleneck that impairs performance, especially on challenging tasks that require fine-grained understanding and reasoning. In this work, we challenge this paradigm by introducing VISion On Request (VISOR), a method that reduces inference cost without discarding visual information. Instead of compressing the image, VISOR improves efficiency by sparsifying the interaction between image and text tokens. Specifically, the language model attends to the full set of high-resolution visual tokens through a small, strategically placed set of attention layers: general visual context is provided by efficient cross-attention between text-image, while a few well-placed and dynamically selected self-attention layers refine the visual representations themselves, enabling complex, high-resolution reasoning when needed. Based on this principle, we first train a single universal network on a range of computational budgets by varying the number of self-attention layers, and then introduce a lightweight policy mechanism that dynamically allocates visual computation based on per-sample complexity. Extensive experiments show that VISOR drastically reduces computational cost while matching or exceeding state-of-the-art results across a diverse suite of benchmarks, and excels in challenging tasks that require detailed visual understanding.

    [cite]
    @inproceedings{bulat2026visor,
      title={VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions},
      author={Bulat, Adrian and Baldrati, Alberto and Metaxas, Ioannis Maniadis and Ouali, Yassine and Tzimiropoulos, Georgios},
      booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2026}
    }
  • 2026

    Restore, Assess, Repeat: A Unified Framework for Iterative Image Restoration New

    I-Hsiang Chen, Isma Hadji, Enrique Sanchez, Adrian Bulat, Sy-Yen Kuo, Radu Timofte, Georgios Tzimiropoulos, Brais Martinez

    IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026

    [abstract]

    Image restoration aims to recover high quality images from inputs degraded by various factors, such as adverse weather, blur, or low light. While recent studies have shown remarkable progress across individual or unified restoration tasks, they still suffer from limited generalization and inefficiency when handling unknown or composite degradations. To address these limitations, we propose RAR, a Restore, Assess and Repeat process, that integrates Image Quality Assessment (IQA) and Image Restoration (IR) into a unified framework to iteratively and efficiently achieve high quality image restoration. Specifically, we introduce a restoration process that operates entirely in the latent domain to jointly perform degradation identification, image restoration, and quality verification. The resulting model is fully trainable end to end and allows for an all-in-one assess and restore approach that dynamically adapts the restoration process. Also, the tight integration of IQA and IR into a unified model minimizes the latency and information loss that typically arises from keeping the two modules disjoint, (e.g. during image and/or text decoding). Extensive experiments show that our approach consistent improvements under single, unknown and composite degradations, thereby establishing a new state-of-the-art.

    [cite]
    @inproceedings{chen2026restore,
      title={Restore, Assess, Repeat: A Unified Framework for Iterative Image Restoration},
      author={Chen, I-Hsiang and Hadji, Isma and Sanchez, Enrique and Bulat, Adrian and Kuo, Sy-Yen and Timofte, Radu and Tzimiropoulos, Georgios and Martinez, Brais},
      booktitle={IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)},
      year={2026}
    }
  • 2025

    Compress & Cache: Vision token compression for efficient generation and retrieval New

    Adrian Bulat*, Yassine Ouali*, Georgios Tzimiropoulos

    Advances in Neural Information Processing Systems (NeurIPS), 2025

    [abstract]

    This work aims to compress the vision tokens of an LVLM into a representation that is simultaneously suitable for (a) generative and (b) discriminative tasks, (c) is nearly lossless, and (d) storage-efficient. To this end, we propose C&C, a novel compression method that leverages the LVLM itself for task-agnostic visual token compression. Unlike prior methods that perform token reduction on-the-fly, our approach offloads computation to a dedicated, upfront indexing stage, effectively decoupling compression from generation. This enables learning more powerful representations for generation during inference. At the core of C&C is a "double-forward pass" training strategy. During the first forward pass, the LLM (of the LVLM) creates a bottleneck by compressing the dense visual tokens into a few summary tokens. Subsequently, the second forward pass processes the language instruction(s) alongside the summary tokens, used as a direct replacement for the image ones. The training of C&C is guided by two key losses: an autoregressive loss applied after the second pass that provides a direct optimization objective for reconstructing the original information flow, and a contrastive loss applied after the first pass to bolster the representational strength of the summary tokens, particularly for discriminative tasks. Moreover, we propose stage-specific adapters for further enhancing performance. C&C produces highly informative compressed representations. An in-depth ablation study confirms the efficacy of our approach. For generative tasks, we achieve a 2x higher compression rate without compromising capabilities, setting a new state-of-the-art. For discriminative tasks, we establish new state-of-the-art results on image retrieval and compositionality benchmarks.

    [cite]
    @inproceedings{bulat2025compresscache,
      title={Compress \& Cache: Vision token compression for efficient generation and retrieval},
      author={Bulat, Adrian and Ouali, Yassine and Tzimiropoulos, Georgios},
      booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
      year={2025}
    }
  • 2025

    Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions New

    Ioanna Ntinou, Alexandros Xenos, Yassine Ouali, Adrian Bulat, Georgios Tzimiropoulos

    Conference on Empirical Methods in Natural Language Processing (EMNLP), 2025

    [abstract]

    Contrastively-trained Vision-Language Models (VLMs), such as CLIP, have become the standard approach for learning discriminative vision-language representations. However, these models often exhibit shallow language understanding, manifesting bag-of-words behaviour. These limitations are reinforced by their dual-encoder design, which induces a modality gap. Additionally, the reliance on vast web-collected data corpora for training makes the process computationally expensive and introduces significant privacy concerns. To address these limitations, in this work, we challenge the necessity of vision encoders for retrieval tasks by introducing a vision-free, single-encoder retrieval pipeline. Departing from the traditional text-to-image retrieval paradigm, we migrate to a text-to-text paradigm with the assistance of VLLM-generated structured image descriptions. We demonstrate that this paradigm shift has significant advantages, including a substantial reduction of the modality gap, improved compositionality, and better performance on short and long caption queries, all attainable with only a few hours of calibration on two GPUs. Additionally, substituting raw images with textual descriptions introduces a more privacy-friendly alternative for retrieval. To further assess generalisation and address some of the shortcomings of prior compositionality benchmarks, we release two benchmarks derived from Flickr30k and COCO, containing diverse compositional queries made of short captions, which we coin subFlickr and subCOCO. Our vision-free retriever matches and often surpasses traditional multimodal models. Importantly, our approach achieves state-of-the-art zero-shot performance on multiple retrieval and compositionality benchmarks, with models as small as 0.3B parameters.

    [cite]
    @inproceedings{ntinou2025visionfree,
      title={Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions},
      author={Ntinou, Ioanna and Xenos, Alexandros and Ouali, Yassine and Bulat, Adrian and Tzimiropoulos, Georgios},
      booktitle={Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
      year={2025}
    }

Research

His current research interests lie at the intersection of Computer Vision and Machine Learning, with work conducted on topics such as efficient neural networks (via bit quantization, network binarization and compression) and human analysis (face alignment/recognition/super-resolution and human pose estimation).

Teaching

  • 2016–2018
    Teaching Assistant · University of Nottingham

    G52CPP: 2nd-year C++ Programming. G53SEC: 3rd-year Network Security. G53VIS: 3rd-year Computer Vision.