By Oguzhan Baser
Accepted to the Fifth Learning on Graphs Conference (LoG 2026).
Authors: Hakan Emre Gedik, Andrew Martin, Mustafa Munir, Oguzhan Baser, Radu Marculescu, Sandeep P. Chinchali, and Alan C. Bovik.
TL;DR: AttentionViG teaches Vision GNNs which neighbors matter. Cross-attention dynamically weights neighboring nodes, helping sparse graphs suppress irrelevant features while reaching up to 83.9% ImageNet-1K top-1 accuracy.
| Read the paper |
Vision Graph Neural Networks (ViGs) represent image patches as nodes and exchange information through graph edges. Their performance depends on graph construction and on how each node aggregates its neighbors.
Dynamic k-nearest-neighbor graphs can capture semantic relationships, but neighbor search is expensive. Fixed sparse graphs such as SVGA are more efficient, yet may connect semantically unrelated nodes. Common aggregation operators also lack an explicit mechanism for learning how important each proposed neighbor should be.
We ask:
The key idea is to separate graph construction from neighbor importance. SVGA supplies candidate neighbors; cross-attention learns which should contribute strongly and which should be suppressed. This makes message passing adaptive to image content without expensive dynamic neighbor search. The parallelized Grapher implementation retains linear scaling with input resolution.
On ImageNet-1K, AttentionViG-S/M/B achieve 81.3% / 83.1% / 83.9% top-1 accuracy at 1.6 / 3.2 / 4.8 GFLOPs. AttentionViG-B matches GreedyViG-B at 83.9% while using 4.8 rather than 5.2 GFLOPs. These results use the paper’s training setup with knowledge distillation from a RegNetY-16GF teacher.
AttentionViG-B also reaches 46.4 box AP and 42.3 mask AP on MS-COCO, and 47.8 mIoU on ADE20K, using Mask R-CNN and Semantic FPN, respectively.
In vanilla ViG, cross-attention reaches 74.3% top-1 at 1.6 GFLOPs, matching EdgeConv’s accuracy at about two-thirds of its FLOPs. In the affinity-function ablation, exponential affinity reaches 81.3% top-1 versus 80.8% with softmax.
Similarity visualizations show the model emphasizing semantically related regions while suppressing unrelated ones.
AttentionViG shows that efficient graph construction does not require every proposed edge to be equally useful. Learning neighbor relevance allows Vision GNNs to remain sparse and efficient while becoming more selective about the information they propagate. Extending the aggregation idea to video, point clouds, and other graph-based domains is a promising direction for future work.
For the complete experiments, ablations, and implementation details, read the paper.