Skip to content

Recommendation with LightGCN (MovieLens)

Author: K3-Node Team
Backend: Multi-Backend
Dataset: MovieLens100K
Description: Recommend movies to users.

View in Colab   GitHub source


Recommendation with LightGCN (MovieLens)

Recommend movies to users. LightGCN (He et al., 2020) learns an embedding for every user and item and smooths them over the user-item interaction graph with simple, parameter-free propagation. It is trained with the BPR loss: a movie the user rated should score higher than a random movie.

Same model as PyG's examples/lightgcn.py; on MovieLens 100K instead of AmazonBook, to keep the example small.

Install K3-Node, then choose a backend: "tensorflow", "torch" or "jax"

!pip install k3-node[examples]
import os
os.environ["KERAS_BACKEND"] = "tensorflow"

Load the data

Users and movies are nodes of one graph (movies are numbered after the users). edge_index holds the training interactions in both directions; edge_label_index the held-out test interactions.

import keras
import numpy as np
from keras import ops
from k3_node.datasets import MovieLens100K
from k3_node.metrics import precision_recall_at_k
from k3_node.models import LightGCN
from k3_node.training import gradient_step

dataset = MovieLens100K("data/MovieLens100K")
data = dataset[0]
num_users, num_items = dataset.num_users, dataset.num_items
train_edges = np.asarray(ops.convert_to_numpy(data.edge_index))
train_edges = train_edges[:, train_edges[0] < train_edges[1]]  # each (user, movie) pair once
print(data)

Define the model

model = LightGCN(num_nodes=data.num_nodes, embedding_dim=64, num_layers=2)

Train

Each step takes a batch of user-movie pairs and pairs every user with a random movie as the negative.

optimizer = keras.optimizers.Adam(learning_rate=0.001)
batch_size = 8192


def train_epoch():
    losses = []
    for start in range(0, train_edges.shape[1], batch_size):
        pos = train_edges[:, start:start + batch_size]
        neg = np.stack([pos[0], np.random.randint(num_users, num_users + num_items, pos.shape[1])])
        edge_label_index = np.concatenate([pos, neg], axis=1)

        def loss():
            pos_rank, neg_rank = ops.split(model(data.edge_index, edge_label_index), 2)
            return model.recommendation_loss(pos_rank, neg_rank, node_id=np.unique(edge_label_index))

        losses.append(gradient_step(loss, model.trainable_variables, optimizer))
    return float(np.mean(losses))


model(data.edge_index, train_edges[:, :2])  # create the weights
for epoch in range(1, 101):
    np.random.shuffle(train_edges.T)
    loss = train_epoch()
    if epoch % 10 == 0:
        emb = model.get_embedding(data.edge_index)
        precision, recall = precision_recall_at_k(emb[:num_users], emb[num_users:], train_edges, data.edge_label_index, k=20)
        print(f"Epoch {epoch:03d}: loss {loss:.4f}, Precision@20 {precision:.4f}, Recall@20 {recall:.4f}")

Evaluate

emb = model.get_embedding(data.edge_index)
precision, recall = precision_recall_at_k(emb[:num_users], emb[num_users:], train_edges, data.edge_label_index, k=20)
print(f"Test Precision@20: {precision:.4f}, Recall@20: {recall:.4f}")