Unofficial PyTorch reproduction for Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.
-
Updated
Jul 2, 2026 - Python
Unofficial PyTorch reproduction for Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.
Log-linear sparse attention (coarse-to-fine hierarchical top-k with KV enrichment) plus a small pixel-space diffusion transformer, in MLX.
Tabular foundation model experiments with learned context sampling and efficient attention
Hybrid Turkish LLM research: replacing 20% of Qwen3-14B softmax attention layers with Gated DeltaNet, with reproducible benchmarks and confidence intervals.
Element-wise attention in MLX: squared-Euclidean distance scores with a t-th order Taylor expansion, giving linear-time training and a recurrent inference form.
Adaptive Vector Quantized Attention — production-grade attention infrastructure for PyTorch (drop-in AVQ-Attention backend)
Minimal implementation of Samba by Microsoft in PyTorch
Sub-quadratic O(n log n) replacement for self-attention — efficiency benchmark (results preview; mechanism held back pending write-up)
Run local AI models on your machine with a secure, Rust-based inference engine that keeps your data private and provides controlled system access.
HiCI: Hierarchical Construction-Integration for Long-Context Attention
O(N) attention with a bounded inference KV cache. D4 Daubechies wavelet field + content-gated Q·K gather at dyadic offsets.
DWARF v2 is a research prototype that expands on the original DWARF.
Nonparametric Modern Hopfield Models
Two small-scale research threads with pre-registered falsifiable bars + adversarial referee audits: Prizma-Seq (a parameter-free quadratic delta-state sequence mixer, an efficient-attention-replacement candidate) and Prizma (backprop-free, fully-local continual learning).
Official Implementation of SEA: Sparse Linear Attention with Estimated Attention Mask (ICLR 2024)
Implementation of: Hydra Attention: Efficient Attention with Many Heads (https://arxiv.org/abs/2209.07484)
Pytorch implementation of "Compact Global Descriptor for Neural Networks" (CGD).
Official repository for "SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space"
Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing (Technical Report)
Unofficial PyTorch implementation of the paper "cosFormer: Rethinking Softmax In Attention".
To associate your repository with the efficient-attention topic, visit your repo's landing page and select "manage topics."