ABSTRACT:
This chapter explores Vision Transformers (ViTs) and foundation models as advanced approaches for intelligent quality assessment. It explains how Vision Transformers use image patches and self-attention to capture global relationships between visual regions, overcoming some limitations of conventional CNN-based approaches. The chapter discusses ViT architecture, self-attention, quality classification, defect detection, grading, segmentation, and regression-based quality prediction. It further examines foundation models, vision-language models, and multimodal foundation models for transferable representations and task adaptation. Applications in agriculture, food quality, manufacturing, healthcare, and textiles are presented. Challenges related to computational requirements, dataset availability, domain adaptation, bias, and explainability are also considered.
Cite this article:
Kranti Kumar Dewangan. Vision Transformers and Foundation Models for Quality Assessment.Multimodal Artificial Intelligence for Intelligent Quality Assessment. 2026; 1(1): 16-22. DOI: 10.52711/book.anv.9788199786417-03DOI: https://doi.org/10.52711/book.anv.9788199786417-03
References not available.