HALLUCINATION-AWARE UAV DATASET BOOTSTRAPPING FOR ISR
DOI:
https://doi.org/10.31891/csit-2026-3-22Keywords:
autonomous unmanned aerial vehicles (UAVs), annotation framework, grounded summaries, public aerial benchmark, open vocabulary detection, secure local deploymentAbstract
This paper studies hallucination-aware annotation for UAV-style imagery under publication-safe conditions and treats the problem as one of measurement, reliability, and deployment rather than as a small prompt-engineering demonstration. The proposed framework combines ontology normalization, overlay masking, grounded summary generation, controlled prompting, and structured metadata capture so that labels, counts, and textual descriptions can be audited jointly. To address the gap between convenient multimodal demonstrations and publishable evidence, the study combines an archival audit of 7 historical military-style frames with a public proxy benchmark of 350 aerial images and frames drawn from AerialMPT, DLR-MVDA, and DLR-ACD. The executed public benchmark contains 274,581 ground-truth objects, moving the empirical basis of the work from a narrow pilot into a hundreds-scale evaluation. The local YOLO-World baseline completed all 350 public samples with mean latency 0.1807 s per image, but it under-counted dense scenes severely, with mean absolute count error of 143.889 on AerialMPT, 297.300 on full DLR-MVDA scenes, and 6848.242 on DLR-ACD. For a matched cloud-versus-local comparison, the paper introduces a dense-tile MVDA benchmark with 10 derived 1024x1024 crops and 315 labeled vehicles. A three-profile Gemini Robotics-ER ablation shows that prompt design materially changes outcome quality: the generic profile completed all tiles but over-counted heavily, the constrained ISR profile reduced error but dropped one request, and the best profile, coarse-to-fine, achieved MAE 5.300, MAPE 10.161%, and 10 / 10 completion on this slice. A supplementary video study on publicly available tactical footage then isolates text control: the freeform condition produced ontology expansion in 21 of 21 non-empty summaries, whereas the matched grounded condition reduced expansion to 0 of 20 while cutting latency from 491.070 s to 272.711 s. The archival audit further reveals unsupported unit, status, or intent language in 4 of 7 historical frames (57.1%) and four raw label variants (mil targets, military target, military targets, vehicle) for the same armored-platform concept. The contribution of the paper is therefore an experimentally grounded framework for hallucination-aware UAV dataset bootstrapping that reports semantic control, counting quality, prompt sensitivity, and execution robustness within one evaluation design.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2026 Yehor TERESHCHENKO, Mika HÄMÄLÄINEN

This work is licensed under a Creative Commons Attribution 4.0 International License.
