Finding core labels for maximizing generalization of graph neural networks.

Journal: Neural networks : the official journal of the International Neural Network Society

Published Date: Aug 14, 2024

Abstract

Graph neural networks (GNNs) have become a popular approach for semi-supervised graph representation learning. GNNs research has generally focused on improving methodological details, whereas less attention has been paid to exploring the importance of labeling the data. However, for semi-supervised learning, the quality of training data is vital. In this paper, we first introduce and elaborate on the problem of training data selection for GNNs. More specifically, focusing on node classification, we aim to select representative nodes from a graph used to train GNNs to achieve the best performance. To solve this problem, we are inspired by the popular lottery ticket hypothesis, typically used for sparse architectures, and we propose the following subset hypothesis for graph data: "There exists a core subset when selecting a fixed-size dataset from the dense training dataset, that can represent the properties of the dataset, and GNNs trained on this core subset can achieve a better graph representation". Equipped with this subset hypothesis, we present an efficient algorithm to identify the core data in the graph for GNNs. Extensive experiments demonstrate that the selected data (as a training set) can obtain performance improvements across various datasets and GNNs architectures.

Authors

Sichao Fu

The General Hospital of Western Theater Command PLA, Chengdu, China.
Xueqi Ma

School of Computing and Information Systems, The University of Melbourne, Parkville, VIC 3010, Australia. Electronic address: xueqim@student.unimelb.edu.au.
Yibing Zhan

JD Explore Academy, Beijing, China. Electronic address: zhanyibing@jd.com.
Fanyu You

University of Southern California, Los Angeles 90005, USA. Electronic address: fyou@usc.edu.
Qinmu Peng

School of Electronic Information and Communications, Huazhong University of Science and Technology, Wuhan, China; Shenzhen Huazhong University of Science and Technology Research Institute, China. Electronic address: pengqinmu@hust.edu.cn.
Tongliang Liu
James Bailey
Danilo Mandic

Department of Electrical and Electronic Engineering, Imperial College London, London, United Kingdom. Electronic address: d.mandic@imperial.ac.uk.

Keywords

Algorithms Humans Neural Networks, Computer Supervised Machine Learning

External Resources

View on PubMed Access via DOI PubMed (39173205)

Finding core labels for maximizing generalization of graph neural networks.

Abstract

Authors

Keywords

External Resources

Popular Topics

Recent Journals