Node embeddings from random walks (Node2Vec)
Author: K3-Node Team
Backend: Multi-Backend
Dataset: Cora (Planetoid)
Description: Learn an embedding for every paper of the Cora graph from random walks, without looking at features or labels.
Node embeddings from random walks (Node2Vec)
Learn an embedding for every paper of the Cora graph from random walks, without looking at features
or labels. Node2Vec (Grover & Leskovec, 2016) trains the embeddings
so that nodes appearing close together on walks get similar embeddings (skip-gram with negative
sampling); p and q bias the walks towards returning or exploring.
Same model as PyG's examples/node2vec.py; with Adam instead of SparseAdam.
Install K3-Node, then choose a backend: "tensorflow", "torch" or "jax"
Load the data
import keras
from keras import ops
from k3_node.datasets import Planetoid
from k3_node.models import Node2Vec
dataset = Planetoid("data/Planetoid", name="Cora")
data = dataset[0]
Define the model
Every node starts 10 walks of length 20; each window of 10 consecutive nodes on a walk is a positive example.
model = Node2Vec(data.edge_index, embedding_dim=128, walk_length=20, context_size=10, walks_per_node=10,
num_negative_samples=1, p=1.0, q=1.0, num_nodes=data.num_nodes)
Train
model.compile(optimizer=keras.optimizers.Adam(learning_rate=0.01))
history = model.fit(epochs=100, batch_size=128)
Evaluate
A logistic regression on the embeddings of the training papers predicts the topics of the test papers.
z = model() # embeddings of all nodes
accuracy = model.test(z[data.train_mask], data.y[data.train_mask], z[data.test_mask], data.y[data.test_mask],
max_iter=150)
print(f"Test accuracy: {accuracy:.4f}")
Visualize
t-SNE projects the embeddings to 2D; colors are the topics of the papers.