Logo: to the web site of Uppsala University

uu.sePublications from Uppsala University
Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow Matching
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Division of Scientific Computing. Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Computational Science.ORCID iD: 0000-0001-9500-1791
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Division of Scientific Computing.ORCID iD: 0009-0000-8703-5764
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Division of Scientific Computing. Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Numerical Analysis. Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Computational Science.ORCID iD: 0000-0001-7273-7923
Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Division of Systems and Control. Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Division Vi3. Uppsala University, Disciplinary Domain of Science and Technology, Mathematics and Computer Science, Department of Information Technology, Artificial Intelligence.ORCID iD: 0000-0003-4480-3158
Show others and affiliations
(English)Manuscript (preprint) (Other academic)
Abstract [en]

Vision-Language Models (VLMs) are typically deterministic in nature and lack intrinsic mechanisms to quantify epistemic uncertainty, which reflects the model's lack of knowledge or ignorance of its own representations. We theoretically motivate negative log-density of an embedding as a proxy for the epistemic uncertainty, where low-density regions signify model ignorance. The proposed method REPVLM computes the probability density on the hyperspherical manifold of the VLM embeddings using Riemannian Flow Matching. We empirically demonstrate that REPVLM achieves near-perfect correlation between uncertainty and prediction error, significantly outperforming existing baselines. Beyond classification, we also demonstrate that the model also provides a scalable metric for out-of-distribution detection and automated data curation.

National Category
Computer Vision and Learning Systems
Identifiers
URN: urn:nbn:se:uu:diva-582516OAI: oai:DiVA.org:uu-582516DiVA, id: diva2:2046714
Available from: 2026-03-17 Created: 2026-03-17 Last updated: 2026-03-31Bibliographically approved
In thesis
1. Robust Learning from Distributed and Heterogeneous Data
Open this publication in new window or tab >>Robust Learning from Distributed and Heterogeneous Data
2026 (English)Doctoral thesis, comprehensive summary (Other academic)
Abstract [en]

Modern machine learning is increasingly expanding beyond centralized, mono-modal training toward systems that must also learn from data across distributed edge devices and heterogeneous data modalities. This transition breaks the foundational identical and independent distribution (i.i.d.) assumptions of traditional models, making robustness a first-class requirement for real-world applications. This thesis studies the mechanisms and methodologies necessary to achieve algorithmic robustness across three intersecting dimensions: distributed optimization, geometry-aware uncertainty quantification, and simulation-based inference.

The first dimension addresses statistical heterogeneity in Federated Learning (FL), a distributed training framework in which multiple participants collaboratively train a shared model without exchanging their local data. In FL, the non-i.i.d. nature of distributed data often induces performance degradation, convergence issues and fairness problems. Through an empirical study on drug discovery and the development of new algorithms, this work demonstrates that adaptive optimization and dynamic hyperparameter adjustment can mitigate training instabilities. These methods ensure equitable performance across diverse data silos, preventing the global model from favoring specific participants.

The second dimension explores the structural challenges of multi-modal language models, which map data of heterogeneous modalities onto complex, non-Euclidean manifolds. This research models aleatoric and epistemic uncertainty with directional distributions via parametric models and Riemannian Flow Matching. This geometry-aware approach allows models to respect the intrinsic geometric structure of the embedding space, providing a mathematically grounded framework for models to quantify their ignorance when confronted with ambiguous or out-of-distribution inputs.

The final dimension addresses the robustness of a unified framework which supports both forward and inverse processes for Bayesian inference. The proposed framework utilizes a unified Flow Matching model to learn the joint distribution of parameters and observations. By employing randomized masking, this architecture robustly handles partially observed or noisy data, integrating forward and inverse processes into a single cohesive neural network without the need for specialized retraining.  Collectively, this thesis contributes theoretical analyses, novel algorithms, and empirical validations that advance the robustness of machine learning across federated optimization, multi-modal uncertainty quantification, and simulation-based inference, bridging the gap between idealized training assumptions and the demands of real-world applications.

Place, publisher, year, edition, pages
Uppsala: Acta Universitatis Upsaliensis, 2026. p. 62
Series
Digital Comprehensive Summaries of Uppsala Dissertations from the Faculty of Science and Technology, ISSN 1651-6214 ; 2665
Keywords
Machine Learning, Distributed Optimization, Federated Learning, Probabilistic Modeling, Multi-modal Learning
National Category
Artificial Intelligence
Research subject
Computer Science
Identifiers
urn:nbn:se:uu:diva-583490 (URN)978-91-513-2816-4 (ISBN)
Public defence
2026-06-04, 101195, Heinz-Otto Kreiss, Regementsvägen 10, Uppsala, 09:15 (English)
Opponent
Supervisors
Available from: 2026-05-07 Created: 2026-03-31 Last updated: 2026-05-13

Open Access in DiVA

fulltext(2471 kB)28 downloads
File information
File name FULLTEXT01.pdfFile size 2471 kBChecksum SHA-512
4dc0ec64e41674933b763f8d6d7b226cd6daaf881cd85b9b6924a0ac7d77fc6b9248a21d8f7af6722fc42c8e018018cf788839d6f098c185ebbf32a9c5caf08b
Type fulltextMimetype application/pdf

Other links

Pre-print in full-text

Authority records

Ju, LiNautiyal, MayankHellander, AndreasVats, EktaSingh, Prashant

Search in DiVA

By author/editor
Ju, LiNautiyal, MayankHellander, AndreasVats, EktaSingh, Prashant
By organisation
Division of Scientific ComputingComputational ScienceNumerical AnalysisDivision of Systems and ControlDivision Vi3Artificial IntelligenceScience for Life Laboratory, SciLifeLab
Computer Vision and Learning Systems

Search outside of DiVA

GoogleGoogle Scholar
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 2157 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf