Skip to content
#

free-gpu

Here are 17 public repositories matching this topic...

Frontier-class open models on a free Kaggle TPU v5e-8: GLM-5.3-Flash 320B MoE (~64 tok/s, our own JAX engine) and Qwen3.8-27B bf16 (~130 tok/s), 262k context, prefix caching. Works with Claude Code, Codex, opencode and pi.

  • Updated Sep 15, 2026
  • Python

Reproducible experiments for running and optimizing open-weight LLMs on free GPU compute — GGUF, llama.cpp, CUDA, quantization, speculative decoding, and inference benchmarking.

  • Updated Sep 28, 2026
  • Python

Add this topic to your repo

To associate your repository with the free-gpu topic, visit your repo's landing page and select "manage topics."

Learn more