TrueVision: A Human-Aligned Explainable Visual Completion Dataset for Incomplete Artworks
Abstract
Large vision-language models are increasingly used in domains where the input image is incomplete, including damaged
paintings, cropped artworks, and restoration scenarios. In these settings, a plausible answer is not sucient: the model should
distinguish between visible evidence and inferred content, communicate uncertainty, and abstain when the available cues are too
weak. This paper presents the current repository-backed stage of TrueVision, a bilingual English-Romanian dataset and annotation
platform for explainable visual completion on artworks and other incomplete images. The current reportable experimental bundle
contains a shared 100-task overlap annotated by three production annotators, yielding 300 bilingual annotation instances and
deterministic 76/10/14 train-validation-test splits. The live platform already supports a richer structured workflow that separates
visible cues from inferred hidden content in both languages, but the committed gold overlap still follows the earlier single-prediction
annotation format and is therefore the basis of the quantitative results reported here. The paper includes four reproducible local
baselines: zero-shot and few-shot Qwen2-VL 2B prompting, a metadata-aware retrieval baseline, and a lightweight supervised
visual baseline trainable on Apple Silicon. On the held-out test split, retrieval remains the strongest non-learned reference in
English similarity (0.4820), while the trained baseline provides the strongest learned result, with 0.4746 Romanian similarity,
0.4507 English token F1, 0.3223 Romanian token F1, 0.5093 confidence macro-F1, and 0.4853 diculty macro-F1. These results
show that the TrueVision pipeline is operational end to end and already supports controlled comparison of explainable visual
completion methods in a low-resource bilingual setting.








