Skip to content

Graph Datasets

K3Node provides multi-backend graph datasets compatible with k3_node.data.Data and k3_node.data.HeteroData. All datasets are built on top of k3_node.data.InMemoryDataset or k3_node.data.Dataset, outputting native backend-agnostic tensors that seamlessly run across PyTorch, TensorFlow, and JAX.


Benchmark Datasets

KarateClub

k3_node.datasets.karate.KarateClub

Bases: InMemoryDataset

Zachary's karate club network from the `"An Information Flow Model for Conflict and Fission in Small Groups" paper, containing 34 nodes and 156 undirected edges labeled into 4 community classes.

Planetoid

k3_node.datasets.planetoid.Planetoid

Bases: InMemoryDataset

The citation network datasets "Cora", "CiteSeer" and "PubMed" from the "Revisiting Semi-Supervised Learning with Graph Embeddings" paper. Nodes represent documents and edges represent citation links. Training, validation and test splits are given by binary masks.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
name str

The name of the dataset ("Cora", "CiteSeer", "PubMed").

required
split str

The type of dataset split ("public", "full", "geom-gcn", "random"). (default: "public")

'public'
num_train_per_class int

The number of training samples per class for "random" split. (default: 20)

20
num_val int

The number of validation samples for "random" split. (default: 500)

500
num_test int

The number of test samples for "random" split. (default: 1000)

1000
transform callable

A function/transform that takes in a Data object and returns a transformed version.

None
pre_transform callable

A function/transform that takes in a Data object and returns a transformed version.

None
force_reload bool

Whether to re-process the dataset. (default: False)

False

TUDataset

k3_node.datasets.tu_dataset.TUDataset

Bases: InMemoryDataset

A variety of graph kernel benchmark datasets, e.g., "IMDB-BINARY", "REDDIT-BINARY" or "PROTEINS", collected from the TU Dortmund University.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
name str

The name of the dataset.

required
transform callable

Transform function for Data objects.

None
pre_transform callable

Pre-transform function.

None
pre_filter callable

Pre-filter function.

None
force_reload bool

Whether to re-process the dataset.

False
use_node_attr bool

Whether to include continuous node attributes.

False
use_edge_attr bool

Whether to include continuous edge attributes.

False
cleaned bool

Whether to use cleaned dataset version.

False

Amazon

k3_node.datasets.amazon.Amazon

Bases: InMemoryDataset

The Amazon Computers and Amazon Photo networks from the "Pitfalls of Graph Neural Network Evaluation" paper. Nodes represent goods and edges represent that two goods are frequently bought together. Given product reviews as bag-of-words node features, the task is to map goods to their respective product category.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
name str

The name of the dataset ("Computers", "Photo").

required
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
force_reload bool

Whether to re-process the dataset.

False

Coauthor

k3_node.datasets.coauthor.Coauthor

Bases: InMemoryDataset

The Coauthor CS and Coauthor Physics networks from the "Pitfalls of Graph Neural Network Evaluation" paper. Nodes represent authors that are connected by an edge if they co-authored a paper. Given paper keywords for each author's papers, the task is to map authors to their respective field of study.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
name str

The name of the dataset ("CS", "Physics").

required
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
force_reload bool

Whether to re-process the dataset.

False

CitationFull & CoraFull

k3_node.datasets.citation_full.CitationFull

Bases: InMemoryDataset

The full citation network datasets from the "Deep Gaussian Embedding of Graphs: Unsupervised Inductive Learning via Ranking" paper. Nodes represent documents and edges represent citation links. Datasets include "Cora", "Cora_ML", "CiteSeer", "DBLP", "PubMed".

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
name str

The name of the dataset ("Cora", "Cora_ML", "CiteSeer", "DBLP", "PubMed").

required
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
to_undirected bool

Whether the original graph is converted to undirected.

True
force_reload bool

Whether to re-process the dataset.

False

k3_node.datasets.citation_full.CoraFull

Bases: CitationFull

Alias for CitationFull with name="Cora".

WikiCS

k3_node.datasets.wikics.WikiCS

Bases: InMemoryDataset

The semi-supervised Wikipedia-based dataset from the "Wiki-CS: A Wikipedia-Based Benchmark for Graph Neural Networks" paper.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
is_undirected bool

Whether the graph is undirected. (default: True)

None
force_reload bool

Whether to re-process the dataset.

False

WebKB

k3_node.datasets.webkb.WebKB

Bases: InMemoryDataset

The WebKB datasets used in the "Geom-GCN: Geometric Graph Convolutional Networks" paper. Nodes represent web pages and edges represent hyperlinks between them.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
name str

The name of the dataset ("Cornell", "Texas", "Wisconsin").

required
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
force_reload bool

Whether to re-process the dataset.

False

Actor

k3_node.datasets.actor.Actor

Bases: InMemoryDataset

The actor-only induced subgraph of the film-director-actor-writer network used in the "Geom-GCN: Geometric Graph Convolutional Networks" paper.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
force_reload bool

Whether to re-process the dataset.

False

PolBlogs

k3_node.datasets.polblogs.PolBlogs

Bases: InMemoryDataset

The Political Blogs dataset containing 1,490 vertices and 19,025 edges.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
force_reload bool

Whether to re-process the dataset.

False

Airports

k3_node.datasets.airports.Airports

Bases: InMemoryDataset

The Airports dataset from the "struc2vec: Learning Node Representations from Structural Identity" paper.

EmailEUCore

k3_node.datasets.email_eu_core.EmailEUCore

Bases: InMemoryDataset

An e-mail communication network of a large European research institution.

GitHub

k3_node.datasets.github.GitHub

Bases: InMemoryDataset

The GitHub Web and ML Developers dataset.

FacebookPagePage

k3_node.datasets.facebook.FacebookPagePage

Bases: InMemoryDataset

The Facebook Page-Page network dataset.

LastFMAsia

k3_node.datasets.lastfm_asia.LastFMAsia

Bases: InMemoryDataset

The LastFM Asia Network dataset.

Twitch

k3_node.datasets.twitch.Twitch

Bases: InMemoryDataset

The Twitch Gamer networks.


Synthetic & Procedural Datasets

FakeDataset

k3_node.datasets.fake.FakeDataset

Bases: InMemoryDataset

A fake dataset that returns randomly generated k3_node.data.Data objects.

FakeHeteroDataset

k3_node.datasets.fake.FakeHeteroDataset

Bases: InMemoryDataset

A fake dataset that returns randomly generated k3_node.data.HeteroData objects.

BAShapes

k3_node.datasets.ba_shapes.BAShapes

Bases: InMemoryDataset

The BA-Shapes dataset from the "GNNExplainer: Generating Explanations for Graph Neural Networks" paper, containing a Barabasi-Albert (BA) graph with 300 nodes and a set of 80 "house"-structured graphs connected to it.

Parameters:

Name Type Description Default
connection_distribution str

Specifies how the houses and the BA graph get connected ("random", "uniform"). (default: "random")

'random'
transform callable

Transform function.

None

BA2MotifDataset

k3_node.datasets.ba2motif_dataset.BA2MotifDataset

Bases: InMemoryDataset

The synthetic BA-2motifs graph classification dataset for evaluating explainability algorithms, containing 1000 random Barabasi-Albert graphs.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
force_reload bool

Whether to re-process the dataset.

False

StochasticBlockModelDataset

k3_node.datasets.sbm_dataset.StochasticBlockModelDataset

Bases: InMemoryDataset

A synthetic graph dataset generated by the stochastic block model.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
block_sizes [int] or array

The sizes of blocks.

required
edge_probs [[float]] or array

The density of edges between blocks.

required
num_graphs int

The number of graphs. (default: 1)

1
num_channels int

The number of node features. (default: None)

None
is_undirected bool

Whether the graph is undirected. (default: True)

True
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
force_reload bool

Whether to re-process the dataset.

False

RandomPartitionGraphDataset

k3_node.datasets.sbm_dataset.RandomPartitionGraphDataset

Bases: StochasticBlockModelDataset

The random partition graph dataset from the "How to Find Your Friendly Neighborhood: Graph Attention Design with Self-Supervision" paper.

ExplainerDataset

k3_node.datasets.explainer_dataset.ExplainerDataset

Bases: InMemoryDataset

Generates a synthetic dataset for evaluating explainability algorithms, as described in the "GNNExplainer: Generating Explanations for Graph Neural Networks" paper.

Parameters:

Name Type Description Default
graph_generator GraphGenerator or str

The graph generator to use.

required
motif_generator MotifGenerator or str

The motif generator to use.

required
num_motifs int

The number of motifs to attach to the graph.

required
num_graphs int

The number of graphs to generate. (default: 1)

1
graph_generator_kwargs dict

Keyword arguments for graph generator.

None
motif_generator_kwargs dict

Keyword arguments for motif generator.

None
transform callable

Transform function.

None

Relational & Knowledge Graph Datasets

Entities

k3_node.datasets.entities.Entities

Bases: InMemoryDataset

The relational entities networks "AIFB", "MUTAG", "BGS" and "AM".

WordNet18 & WordNet18RR

k3_node.datasets.word_net.WordNet18

Bases: InMemoryDataset

The WordNet18 dataset containing 40,943 entities, 18 relations and 151,442 fact triplets.

k3_node.datasets.word_net.WordNet18RR

Bases: InMemoryDataset

The WordNet18RR dataset.

FB15k_237

k3_node.datasets.freebase.FB15k_237

Bases: InMemoryDataset

The FB15K237 dataset containing 14,541 entities, 237 relations and 310,116 fact triples.

Parameters:

Name Type Description Default
root str

Root directory where the dataset should be saved.

required
split str

"train", "val", or "test". (default: "train")

'train'
transform callable

Transform function.

None
pre_transform callable

Pre-transform function.

None
force_reload bool

Whether to re-process the dataset.

False

Heterogeneous Graph Datasets

DBLP

k3_node.datasets.dblp.DBLP

Bases: InMemoryDataset

A subset of the DBLP computer science bibliography website containing four types of entities: authors, papers, terms, and conferences.

IMDB

k3_node.datasets.imdb.IMDB

Bases: InMemoryDataset

A subset of the Internet Movie Database (IMDB) containing three types of entities: movies, actors, and directors.


Molecular Datasets

QM7b

k3_node.datasets.qm7.QM7b

Bases: InMemoryDataset

The QM7b dataset consisting of 7,211 molecules with 14 regression targets.