memory-hierarchy
Here are 66 public repositories matching this topic...
MIT Course 6.004 - Computation Structures
-
Updated
Sep 28, 2024 - Assembly
Live GPU emulator for High-Bandwidth Flash: applies HBF timing, capacity and thermal effects to a real LLM inference workload while it executes on a real GPU. arXiv:2609.09800
-
Updated
Sep 28, 2026 - Python
My solutions of Computer Systems: A Programmer’s Perspective, Third Edition (CS:APP3e) book, the text book for the course, CMU15-213: Introduction to Computer Systems.
-
Updated
Mar 25, 2022 - C
A Fast DNN Accelerator Design Space Exploration Framework.
-
Updated
Aug 10, 2022 - Python
A high-performance C++20 cache simulator with power/area modeling, MESI coherence, prefetching, and multi-level hierarchy support for architecture research and education.
-
Updated
Feb 10, 2026 - C++
Exerting coherency between caches with protocols in a Memory-Shared Multiprocessors system whether it has uniform memory access(UMA, symmetric) or not(non-UMA).
-
Updated
Mar 6, 2026 - Verilog
Hardware compression for memory
-
Updated
Jul 19, 2019 - Jupyter Notebook
EVict and recOver KV cache Entries. Selective KV cache eviction and recovery for long-context LLM inference.
-
Updated
Oct 1, 2026 - Python
A comprehensive C++20 cache simulator for analyzing memory hierarchy performance with configurable cache levels, replacement policies, and inclusion strategies
-
Updated
May 27, 2025 - C++
AOS Agent is an open-source, multi-agent AI system that acts as your extended cognitive partner, featuring structured memory, specialized agents, and deep integration with various softwares via MCP.
-
Updated
Apr 30, 2026 - Python
A superscalar out-of‐order architectural simulator (With Memory Hierarchy).
-
Updated
Dec 10, 2016 - Java
Implementation of Hierarchy Oblivious Algorithms
-
Updated
Sep 3, 2019 - C++
A modular and fully synthesizable 3-level (L1/L2/L3) cache memory subsystem implemented in SystemVerilog, featuring split L1 caches, a centralized L2, and a Last-Level L3 cache with coherence support.
-
Updated
May 14, 2026 - SystemVerilog
How much of GEMM performance is memory access order? Five CPU variants of the same matrix product, a shared-memory tiled CUDA kernel and a CUDA sum reduction, all measured on one shape. Loop reordering alone is worth 37.9x. C++17, CMake, no dependencies.
-
Updated
Sep 1, 2026 - C++
-
Updated
Oct 7, 2021 - Verilog
(ECE) A collection of helper scripts for the assignments for the course "Advanced Computer Architecture" [3.4.37.8]
-
Updated
Apr 14, 2020 - Shell
Silicon workbench for chip blueprints, roofline analysis, tile-memory flow, and AI accelerator co-design.
-
Updated
Sep 25, 2026 - HTML
The task is to design a "family" of three microprocessors that differ in performance and cost for the same computational task, as a project in "Computer Architecture 2" course.
-
Updated
Jan 15, 2025 - Assembly
Add this topic to your repo
To associate your repository with the memory-hierarchy topic, visit your repo's landing page and select "manage topics."