Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
-
Updated
Oct 2, 2026 - Python
Evidence-backed AMD Strix Halo local-AI setup and benchmarks: Qwen3.8, Ollama, llama.cpp, Vulkan/ROCm, large GGUFs, and cross-OEM results.
Honest local LLM deployment planning and benchmarking for high unified-memory Macs and future Linux/NVIDIA rigs.
Operator-grade GPU monitor for NVIDIA GPUs with native GB10 / DGX Spark coherent UMA support — PSI pressure, clock detection, ConnectX-7 network layer
Local inference server for Apple Silicon — hot-swaps MLX models (LLM, vision, embeddings, TTS, STT) via OpenAI API
A CUDA implementation of the transpose-free Quasi-Minimal Residual method
Unified Memory Abstraction Layer for AI Inference on AMD APUs and Intel iGPUs
Apache Arrow compute on Apple silicon GPUs through Metal: Arrow columns in unified memory, copy-free where the producer's buffers are page aligned; null-aware GPU kernels; Arrow C Data, C Device and C Stream interop
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
gpu thrashingNVIDIA GPU Unified Memory diagnostic tool — architecture-aware, measurement-based, PCIe/coherent transport detection
Apple Silicon Unified Memory for GPU-Accelerated Analytics — TPC-H benchmarks across DuckDB, NumPy, and MLX
Talos-O (Omni): A sovereign, embodied agentic organism forged on AMD Strix Halo. Integrating the Chimera Kernel (Linux 7.0), Zero-Copy Introspection, and the Phronesis Engine. Built from First Principles.
Unlock fast, local LLM inference on AMD-powered mini PCs delivering 65-87 t/s for large models without cloud or subscription costs
Fundamentals of Accelerated Computing C/C++ is a course provided by NVIDIA.
NVML unified memory shim for NVIDIA DGX Spark Grace Blackwell GB10 - enables MAX Engine, PyTorch, and GPU monitoring
Empirical kernel scheduling characterization for NVIDIA GB10 (SM121a). Sweeps GEMM tile configurations, classifies PTX instruction paths, captures hardware telemetry
Run LLMs larger than your RAM — native GGUF inference engine with SSD streaming, no GPU required
System monitor that reports real GPU memory on unified-memory AMD parts (Strix Halo / Ryzen AI Max), where other tools show only the firmware carve-out. Split collector daemon + TUI, no dependencies.
The real-time coordination layer for teams of developers running Claude Code agents. Git coordinates code at rest; Datum coordinates agents in motion.
Performance comparison of two different forms of memory management in CUDA
Research into CUDA Unified Memory as a VRAM extension for LLM inference
To associate your repository with the unified-memory topic, visit your repo's landing page and select "manage topics."