Software architecture diagrams communicate software structure and relationships, but automatically estimating their connectivity remains difficult. We study whether Vision Large Language Models, e.g., LLaVA, Gemma, Mistral, and Qwen-VL, can estimate connectivity through a localized proxy: counting arrowheads, usually denoting dependencies, control flows, or communication links. We evaluate local and cloud-based VLLMs on annotated architecture diagrams, showing that the strongest models capture useful connectivity information: Qwen3-VL-Cloud achieves the best overall error-based performance (MAE 1.73, RMSE 3.65), while Gemma4 obtains the highest exact-match accuracy (48.1%). However, performance degrades on large and dense diagrams. To address this limitation, we introduce an image-slicing strategy counting arrowheads in sub-regions; a 2 × 2 partition improves tolerance-based accuracy from 65.8% to 73.3%. Overall, while estimating diagram connectivity and complexity remains challenging, preliminary results suggest VLLMs combined with simple diagram-aware preprocessing are promising for arrowhead localization and graph reconstruction.
Evaluating Software Architecture Diagram Connectivity with Vision Large Language Models / Scommegna, N., Bindini, L., Noor ali, K., Verdecchia, R., Marinai, S.. - ELETTRONICO. - (2026), pp. 1-4. (26th ACM Symposium on Document Engineering, DocEng 2026 Perolles Campus, che 2026) [10.1145/3820755.3832802].
Evaluating Software Architecture Diagram Connectivity with Vision Large Language Models
Bindini, Luca;Noor ali, Kimiya;Verdecchia, Roberto;Marinai, Simone
2026
Abstract
Software architecture diagrams communicate software structure and relationships, but automatically estimating their connectivity remains difficult. We study whether Vision Large Language Models, e.g., LLaVA, Gemma, Mistral, and Qwen-VL, can estimate connectivity through a localized proxy: counting arrowheads, usually denoting dependencies, control flows, or communication links. We evaluate local and cloud-based VLLMs on annotated architecture diagrams, showing that the strongest models capture useful connectivity information: Qwen3-VL-Cloud achieves the best overall error-based performance (MAE 1.73, RMSE 3.65), while Gemma4 obtains the highest exact-match accuracy (48.1%). However, performance degrades on large and dense diagrams. To address this limitation, we introduce an image-slicing strategy counting arrowheads in sub-regions; a 2 × 2 partition improves tolerance-based accuracy from 65.8% to 73.3%. Overall, while estimating diagram connectivity and complexity remains challenging, preliminary results suggest VLLMs combined with simple diagram-aware preprocessing are promising for arrowhead localization and graph reconstruction.| File | Dimensione | Formato | |
|---|---|---|---|
|
3820755.3832802.pdf
accesso aperto
Tipologia:
Pdf editoriale (Version of record)
Licenza:
Creative commons
Dimensione
582.61 kB
Formato
Adobe PDF
|
582.61 kB | Adobe PDF |
I documenti in FLORE sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



