Attention Layers
The k3_node.layers.attention module provides standalone attention building blocks used by K3-Node's scalable graph transformers (Polynormer, SGFormer, GPSE) and can also be composed into custom architectures. Unlike k3_node.layers.conv attention layers (GATConv, TransformerConv, ...), these operate on dense [batch, nodes, channels] tensors rather than sparse edge_index message passing.
Linear / Kernelized Attention
PerformerAttention
FAVOR+ kernelized linear attention (from the Performer paper), used by Polynormer for its global attention branch. Approximates full softmax attention in linear time/memory by projecting queries/keys through random Fourier features.
k3_node.layers.attention.PerformerAttention
Bases: Layer
k3_node.layers.PerformerAttention
Initialization Arguments:
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
channels
|
The number of output channels. |
required | |
heads
|
The number of attention heads. |
required | |
head_channels
|
The number of attention heads. |
64
|
|
kernel
|
activation function. |
relu
|
|
qkv_bias
|
activation function. |
False
|
|
attn_out_bias
|
Bias in Attention Out. |
True
|
|
dropout
|
Dropout rate. |
0.0
|
Example
import numpy as np
from k3_node.layers import PerformerAttention
x = np.random.rand(1, 10, 8).astype("float32") # [batch, num_nodes, channels]
mask = np.ones((1, 10), dtype=bool) # which nodes are real (not padding)
attn = PerformerAttention(channels=8, heads=2) # linear-complexity attention
print(tuple(attn(x, mask).shape)) # (1, 10, 8)
PerformerProjection
The random-feature projection used internally by PerformerAttention to approximate the softmax kernel; useful standalone when building a custom linear-attention layer.
k3_node.layers.attention.PerformerProjection
Bases: Layer
Layer PerformerProjection.
Example
import numpy as np
from k3_node.layers import PerformerProjection
q = k = v = np.random.rand(1, 2, 10, 8).astype("float32") # [batch, heads, num_nodes, head_dim]
proj = PerformerProjection(num_cols=8) # random-feature approximation of softmax attention
print(tuple(proj(q, k, v).shape)) # (1, 2, 10, 8)
Usage:
from keras import ops
from k3_node import layers as k3_layers
attn = k3_layers.PerformerAttention(channels=64, heads=4)
mask = ops.ones((1, num_nodes)) # [batch, nodes]
out = attn(x, mask) # x: [batch, nodes, 64]
Polynomial & Scalable Graph Attention
PolynormerAttention
Local-to-global polynomial-expressive attention from Polynormer, combining a local propagation term with a global attention term whose polynomial expansion is computed via PerformerAttention-style kernelization.
k3_node.layers.attention.PolynormerAttention
Bases: Layer
Layer PolynormerAttention.
Example
SGFormerAttention
All-pair, single-layer linear attention from SGFormer, designed to replace deep attention stacks with one global mixing layer for scalable graph transformers.
k3_node.layers.attention.SGFormerAttention
Bases: Layer
The simple global attention mechanism from the
"SGFormer: Simplifying and Empowering Transformers for
Large-Graph Representations"
<https://arxiv.org/abs/2306.10759>_ paper.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
channels
|
int
|
Size of each input sample. |
required |
heads
|
int
|
Number of parallel attention heads.
(default: :obj: |
1
|
head_channels
|
int
|
Size of each attention head.
(default: :obj: |
64
|
qkv_bias
|
bool
|
If specified, add bias to query, key
and value in the self attention. (default: :obj: |
False
|
Example
Usage:
Q-Former (Query Transformer)
QFormer
A BERT-style querying transformer (as used in BLIP-2 / GPSE) that distills a variable-length input sequence into a fixed set of learned query tokens via cross-attention, useful for pooling variable-size node sets into a fixed-size graph representation.
k3_node.layers.attention.QFormer
QFormerEncoderLayer
A single self-attention + feed-forward encoder block used inside QFormer.
k3_node.layers.attention.QFormerEncoderLayer
Bases: Layer
Layer QFormerEncoderLayer.
Example
Usage: