I like machines close to the metal and results you can measure. I run tiyuvta, a research lab that helps companies choose, adapt and deploy open models on their hardware or cloud account. I also build in-memory data systems at AWS ElastiCache, and maintain open-source tools people actually run.
Research notes and working papers live at avifenesh.ai. Issues, questions, and counterexamples are always welcome.
tiyuvta is my research lab. We help companies run open models on their hardware or cloud account. The work includes model and hardware assessment, deployment and serving-engine optimization, fine-tuning and evaluation, and cloud setup. We compare options on your workload and budget, then agree the deliverables, handover and support.
Our research in model execution, quantization, pruning, speculative decoding and Hebrew models informs that work. See the services and discuss your project.
The lab's models are on Hugging Face: NVFP4 builds of GLM, DeepSeek, MiMo, Qwen and DictaLM, MTP draft heads in GGUF, and Phonon-2 in ONNX.
- memra: tiyuvta's from-scratch Rust + CUDA research engine for RTX PRO 6000 Blackwell and RTX 5090, an instrument behind the studies below. Safetensors is the tuned path, GGUF stays supported, and a mechanism that wins on one card and loses on the other becomes a per-device default rather than a compromise. On crates.io with prebuilt binaries.
- hqmtp: MTP draft-head lab, concluded. Function cuts (pruning, low-rank, distillation) pay a 10–19-point off-distribution tax that fidelity cuts don't; the zero-training trimmed-vocabulary recipe won at 1.8–2.7× end to end. The negative results stay in the ledger.
- recipe-lab: layer-loop weight sharing + ε=λ/(N√L) residual scaling, combined for the first time and tested from zero in 11 pre-registered rounds. In the data-constrained regime the looped model beat FLOPs-matched vanilla in all three mixer families (attention, pure SSM, and hybrid; seven paired runs, zero sign flips), with 26–34% fewer parameters. Rule isolated: loop the state-mixer, never the retriever.
- mem-retrofit: grafted a product-key memory layer onto a stock dense 4B and ran it against LoRA over sequential updates. The retrofit is free at lr/10; the published forgetting advantage failed 12/12 confidence intervals.
- More studies with receipts: fixed-compute-frontier (a preregistered kill-gate ledger, ~84 theory lanes) · gemma-expert-atlas (26B MoE expert surgery, 3,840 experts traced) · block-routed-swiglu (near-free kernel, refuted capability) · moe-lab · assumption-excavator.
- Working papers: small-vocabulary MTP heads · prune, heal, quantize. Methods, failed arms, and evidence in the open.
- In review upstream:
- SGLang: SM120 (RTX PRO 6000 / 5090) fixes for sparse MLA, DeepSeek V4 decode and long FP8 prefill, plus DeepGEMM BF16 output divergence. HiCache: a hybrid host tier for Mamba state, SWA match gating, NextN layer handling. Tool calls: strict DSML parsing for DeepSeek V3.2/V4, Qwen3Coder object parameters. Also MiMo ModelOpt mappings and credential redaction in logs.
- vLLM: hybrid-KV loads through LMCache, activation caching in llm-compressor.
- LMCache: binary buffers and safe read failures in the local disk backend.
- llama.cpp: imatrix-aware NVFP4 quantization.
- Valkey and its ecosystem. I maintain Valkey GLIDE, the official multi-language client (Rust core, Java/JNI, Node/N-API), and valkey-skills, the official AI skills for the ecosystem, which I started. I contribute to Valkey itself, and a good part of my open-source time goes to the people around it: the clients team, talks, and helping the people who run it.
- agent-sh, my org: an ecosystem of tools for agent-assisted development, working across Claude Code, Codex, OpenCode, Cursor, and Kiro.
- glide-mq: Node.js queue on Valkey Streams with a Rust N-API core, plus adapters for Hono, Fastify, Hapi, NestJS and a dashboard.
- Also around: RustOwl (runtime, memory, and CI work).
- Valkey: core contributor; sync-from-replica replication in review. On the side: CRIU copy-on-write live-migration research, under 50 ms of freeze while migrating a 200 GB loaded instance.
- SGLang router on Valkey: a shared, restart-safe placement index for sgl-kv-indexer, then an event log with worker replay and indexer failover. Measured under load in sglang-valkey-demo.
- ferrings: io_uring TCP transport for Node.js with a Rust N-API core. 2.5x Node
httpthroughput and 37-57% fewer syscalls per connection in the published bench. On npm. - FlowFabric: durable-execution engine in Rust for Valkey, Postgres, and SQLite: lease-safe workers, waitpoints, budgets.
- layout-audit: DWARF memory-layout analysis: padding, layout diffs, size budgets for C/C++/Rust/Go.
- scrump: format-aware secret scrubber for binary capture artifacts: perf.data, core dumps, nsys traces, JFR.
- ocaml-valkey: OCaml 5 + Eio Valkey client, published on opam.
- agnix: linter and language server for AI agent configs: 444 rules with autofixes, a GitHub Action, an MCP server, and editor plugins.
- computer-use-linux / agent-workspace-linux: Linux desktop control over MCP, and isolated agent-owned desktops so an agent never has to touch your real machine.
- parlar: voice mode for Claude Code and Codex. You talk to a running session, idle or mid-work, and it answers out loud. Local CPU only: VAD, Phonon-2 recognition, Kokoro speech. On npm and crates.io.
- agentsys: the agent-sh plugin set (workflow, review, ship, deslop, perf) for Claude Code, OpenCode, Codex, Cursor, and Kiro.
- revuto: local PR reviewer that works with any model and learns each repo from its PR history and maintainer feedback. It reviews my own repos.
- harness tools: read, write, grep, glob, bash, webfetch, lsp and skill tools built for LLM callers, as
@agent-sh/harness-*on npm with Rust ports at parity. - linubot: Linux desktop app for a team of AI helpers: per-bot providers, memory, and isolated computer workspaces.
- Codex Desktop for Linux: contributor (120+ commits) to the unofficial Linux build of the ChatGPT/Codex desktop app. Linux Computer Use, Read Aloud and conversation mode, the in-app updater, launcher hardening.
- Talks: Inside Valkey GLIDE on the AWS Developers Podcast · Glide into resiliency on Let's Talk About Data
- Writing: avifenesh.ai/writing · answering on Stack Overflow
- Hugging Face: Avifenesh · tiyuvta
- 📫 aviarchi1994@gmail.com · LinkedIn · X
If something here saved you time, sponsoring helps me keep doing it.






