llm theoretical performance analysis tools and support params, flops, memory and latency analysis.
-
Updated
Jul 11, 2025 - Python
llm theoretical performance analysis tools and support params, flops, memory and latency analysis.
code for benchmarking GPU performance based on cublasSgemm and cublasHgemm
Hands-on Machine Learning Infrastructure on Kubernetes. Using Microk8s/Ubuntu on Paperspace Cloud.
A systematic CPU/GPU performance study of lightgbm and xgboost classifiers for different data shapes and hardware setups.
Disable GPU Thermal and Change GPU Governor to performance. (Only for snapdragon devices).
gpu thrashingNVIDIA GPU Unified Memory diagnostic tool — architecture-aware, measurement-based, PCIe/coherent transport detection
Evidence-driven PresentMon diagnostics and policy modeling paired with bounded owned-lab D3D11 runtime actuation.
This repository provides the latest benchmarks for the CHARMM/pyCHARMM program on GPUs
Comprehensive performance analysis of DeepSeek V3 quantization levels (FP16, Q8_0, Q4_0) on 16GB GPU environments.
Interactive theoretical Kimi-K3 inference roofline calculator for H200, B300, and GB300
Windows runtime for read-only process monitoring, cooperative D3D11/D3D12/Vulkan execution and bounded shared-memory research. C++20 + .NET, explicit safety gates and reproducible evidence.
📊 Mobile VR Performance Optimization Project
Cycle-accurate UMA fault latency and bandwidth measurement for NVIDIA GPUs. C and PTX. No Python. Pascal (SM 6.0) through Blackwell GB10 (SM 12.1).
Reproducible long-context inference benchmark comparing vLLM, SGLang, and TensorRT-LLM on NVIDIA GB10.
Agent Skills and an MCP server for GPU performance profiling, benchmarking, optimization, and reporting, with an inference focus.
Reproducible Qwen3-1.7B prefix cache benchmark on RTX 3070 Laptop (8GB). Hand-written reference inference loop + vLLM v1 comparison.
Python lab for exploring memory bandwidth, cache effects, and locality in accelerator workloads
WMMA FP16-->FP32 Tensor Core GEMM with shared-memory tiling and cp.async-style pipelining, benchmarked against cuBLAS.
Professional GPU Performance Testing Suite for dual GPU setup (AMD RX 6600 + NVIDIA RTX 3050) with comprehensive monitoring tools, crash-safe scripts, and thermal management optimized for ASRock X570 Taichi + Ryzen 7 5700X
Reproducible Instruction Roofline analysis of cuSPARSE and Ginkgo SpMM on RTX 4090 using Nsight Compute metrics.
To associate your repository with the gpu-performance topic, visit your repo's landing page and select "manage topics."