Unifying Image Quality Assessment Datasets: MOSAIQ-500K and MOSAIQ-Bench
Image quality assessment (IQA) datasets use different subjective protocols and rating scales, so their scores are not directly comparable. The lack of a common perceptual scale hinders multi-dataset training and precludes direct inter-dataset evaluation. We address this by conducting a new subjective experiment and using its ratings as perceptual anchors to fit monotonic mappings that place the existing scores of 23 IQA datasets on a common quality scale while preserving within-dataset rankings. The resulting dataset, MOSAIQ-500K, contains over 500,000 images and is, to our knowledge, the largest IQA dataset with perceptually aligned subjective scores. We also propose MOSAIQ-Bench, an inter-dataset benchmark, and use it to evaluate 31 IQA methods, revealing substantial gaps between intra- and inter-dataset performance, particularly on authentic distortions. The aligned scores offer a key advantage: replacing specialized multi-dataset training mechanisms with standard regression losses yields comparable intra-dataset accuracy and better inter-dataset performance. MOSAIQ thus provides a practical foundation for combining heterogeneous IQA datasets to train more generalizable IQA models. Code and dataset will be made public at ivc.uwaterloo.ca/projects/unifying_iqa_datasets/ upon acceptance.
Comments
Log in to comment, reply, and vote.
No comments yet.