Run AI models too large for your Mac's memory — at near-full speed. Intelligent expert caching, speculative execution, and 15+ research techniques for MoE inference on Apple Silicon.
-
Updated
Oct 6, 2026 - Python
Run AI models too large for your Mac's memory — at near-full speed. Intelligent expert caching, speculative execution, and 15+ research techniques for MoE inference on Apple Silicon.
Lightweight Python lazy imports that defer module loading to reduce startup time and initial memory use.
Pale Moon is a free full-version web browser for Windows that enhances Firefox's performance by 25%. Enjoy faster loading times, improved memory usage, and easy profile migration. Download the latest version free.
🚀 Benchmark Rust performance to test the "pitfall theory," revealing the true value of optimized techniques over naive implementations.
Compress and run LLM KV caches with KVTC to cut memory use 6-9x and support 2M+ token context on one GPU
Run large MLX models on Apple Silicon with flash weight streaming, using native precision beyond RAM limits
Counts each recorded trade close exactly once across JSONL/CSV copies and retries, streaming rows through size-capped sorted runs on disk. SYNTHETIC 100k closes / 225k rows: peak RSS 683.8 MB -> 27.5 MB (-96.0%) and 11.8% less median time than the in-memory reader, byte-identical reports. Python stdlib only.
Compress LLM KV cache by 5–7x with near-zero accuracy loss for longer context and lower GPU use
A tiered-memory system design for workloads that don't fit in RAM: measure the working set, pin the hot tier, stream the cold tier from flash. Ships the residency calculator, measurement harnesses, and the build recipes behind it. Predictions validated against public benchmarks.
YaiJS Component Library - Advanced VanillaJS web components built on YEH
Making DuckDB's capabilities available to the JVM efficiently. Includes a DuckDB-backed CQEngine persistence plugin: 7x less memory, and joins across collections.
A semantic network and reasoning engine in which rules, facts and numbers are nodes of one graph, so a statement is itself a node. Importing the statements of the 1.7 TB Wikidata dump builds 983 million nodes on a single machine, and reasoning over them has surfaced thousands of consistency violations in the Wikidata ontology.
Memory intelligence and optimization for local AI in Go — hardware detection, KV-cache estimation, and VeloxQuant runtime client
Diagnosing non-monotonic activation memory in torch.compile's memory-budget partitioner, with a use-site rematerialization fix.
Optimize memory allocation in Rust with Auto-Allocator. Enjoy smart, platform-aware performance boosts effortlessly. 🌟🚀
Production-focused portfolio tracking a systematic transition from Linux Infrastructure Administration to AI System Architecture. Specializing in low-level CPython runtime mechanics, memory lifecycle management, and high-density MLOps pipelines.
Fused, logit-free linear-cross-entropy loss, RAM-fit planner, and benchmark harness for MLX fine-tuning on Apple Silicon
Lightweight Windows browser shell on Edge WebView2 (.NET 10 / C#) with a 20-list ad-block engine, 3-tier tab memory policy, and CRX extension support.
Portable Windows app and process manager with whitelists, startup/uninstall controls, hidden-window recovery, adaptive downloads and performance overlays. WinUI x64/ARM64 + WPF compatibility UI.
Hardware-aware memory intelligence for local LLM inference.
To associate your repository with the memory-optimization topic, visit your repo's landing page and select "manage topics."