Instructions to use marcsun13/ggml-gated-delta-net with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Kernels
How to use marcsun13/ggml-gated-delta-net with Kernels:
# !pip install kernels from kernels import get_kernel kernel = get_kernel("marcsun13/ggml-gated-delta-net") - Notebooks
- Google Colab
- Kaggle
ggml-gated-delta-net
The gated delta rule from llama.cpp as a single kernel
(gated_delta_net) โ the linear-attention recurrence behind Qwen3-Next and Qwen3.5, which eager torch
spells out in ~200 ops per layer. l2_norm is exposed alongside it.
q and k carry one head per value head rather than being pre-expanded, and the recurrent state is
indexed [value][key] โ store what the op returns rather than transposing it. Ask
supports_gated_delta_net for the head dims it covers.
Usage
import torch
from kernels import get_kernel
gdn = get_kernel("marcsun13/ggml-gated-delta-net", version=1)
n_seqs, n_tokens, n_heads, head_dim = 1, 1, 32, 128
q = torch.randn(n_seqs, n_tokens, n_heads, head_dim, device="mps")
k = torch.randn_like(q) # one head per value head, not expanded
v = torch.randn_like(q)
g = torch.randn(n_seqs, n_tokens, n_heads, device="mps") # log-domain gate
beta = torch.rand(n_seqs, n_tokens, n_heads, device="mps")
state = torch.zeros(n_seqs, n_heads, head_dim, head_dim, device="mps")
out, state = gdn.gated_delta_net(q, k, v, g, beta, state) # (1, 1, 32, 128), (1, 32, 128, 128)
- Downloads last month
- 26
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support