You can now run 70B LLMs on a 4GB GPU. AirLLM just made massive models usable on low-memory hardware. 𝗪𝗵𝗮𝘁 𝗷𝘂𝘀𝘁 𝗵𝗮𝗽𝗽𝗲𝗻𝗲𝗱 AirLLM released memory-optimized inference for large language models. It runs 70B models on 4GB VRAM. It can even run 405B Llama 3.1 on 8GB VRAM. 𝗛𝗼𝘄 𝗶𝘁 𝘄𝗼𝗿𝗸𝘀 AirLLM loads models one layer at a time. Instead of loading everything: → Load a layer → Run computation → Free memory → Load the next layer This keeps GPU memory usage extremely low. 𝗞𝗲𝘆 𝗱𝗲𝘁𝗮𝗶𝗹𝘀 • No quantization required by default • Optional 4-bit or 8-bit weight compression • Same API as Hugging Face Transformers • Supports CPU and GPU inference • Works on Linux and macOS Apple Silicon 𝗪𝗵𝗮𝘁 𝘆𝗼𝘂 𝗰𝗮𝗻 𝗱𝗼 • Run Llama, Qwen, Mistral, Mixtral locally • Test large models without cloud GPUs • Prototype agents on cheap hardware
Large Language Models Insights
Explore top LinkedIn content from expert professionals.
-
-
𝗧𝗵𝗶𝘀 𝗶𝘀 𝗵𝗮𝗻𝗱𝘀 𝗱𝗼𝘄𝗻 𝗼𝗻𝗲 𝗼𝗳 𝘁𝗵𝗲 𝗕𝗘𝗦𝗧 𝘃𝗶𝘀𝘂𝗮𝗹𝗶𝘇𝗮𝘁𝗶𝗼𝗻 𝗼𝗳 𝗵𝗼𝘄 𝗟𝗟𝗠𝘀 𝗮𝗰𝘁𝘂𝗮𝗹𝗹𝘆 𝘄𝗼𝗿𝗸. ⬇️ 𝘓𝘦𝘵'𝘴 𝘣𝘳𝘦𝘢𝘬 𝘪𝘵 𝘥𝘰𝘸𝘯: 𝗧𝗼𝗸𝗲𝗻𝗶𝘇𝗮𝘁𝗶𝗼𝗻 & 𝗘𝗺𝗯𝗲𝗱𝗱𝗶𝗻𝗴𝘀: - Input text is broken into tokens (smaller chunks). - Each token is mapped to a vector in high-dimensional space, where words with similar meanings cluster together. 𝗧𝗵𝗲 𝗔𝘁𝘁𝗲𝗻𝘁𝗶𝗼𝗻 𝗠𝗲𝗰𝗵𝗮𝗻𝗶𝘀𝗺 (𝗦𝗲𝗹𝗳-𝗔𝘁𝘁𝗲𝗻𝘁𝗶𝗼𝗻): - Words influence each other based on context — ensuring "bank" in riverbank isn’t confused with financial bank. - The Attention Block weighs relationships between words, refining their representations dynamically. 𝗙𝗲𝗲𝗱-𝗙𝗼𝗿𝘄𝗮𝗿𝗱 𝗟𝗮𝘆𝗲𝗿𝘀 (𝗗𝗲𝗲𝗽 𝗡𝗲𝘂𝗿𝗮𝗹 𝗡𝗲𝘁𝘄𝗼𝗿𝗸 𝗣𝗿𝗼𝗰𝗲𝘀𝘀𝗶𝗻𝗴) - After attention, tokens pass through multiple feed-forward layers that refine meaning. - Each layer learns deeper semantic relationships, improving predictions. 𝗜𝘁𝗲𝗿𝗮𝘁𝗶𝗼𝗻 & 𝗗𝗲𝗲𝗽 𝗟𝗲𝗮𝗿𝗻𝗶𝗻𝗴 - This process repeats through dozens or even hundreds of layers, adjusting token meanings iteratively. - This is where the "deep" in deep learning comes in — layers upon layers of matrix multiplications and optimizations. 𝗣𝗿𝗲𝗱𝗶𝗰𝘁𝗶𝗼𝗻 & 𝗦𝗮𝗺𝗽𝗹𝗶𝗻𝗴 - The final vector representation is used to predict the next word as a probability distribution. - The model samples from this distribution, generating text word by word. 𝗧𝗵𝗲𝘀𝗲 𝗺𝗲𝗰𝗵𝗮𝗻𝗶𝗰𝘀 𝗮𝗿𝗲 𝗮𝘁 𝘁𝗵𝗲 𝗰𝗼𝗿𝗲 𝗼𝗳 𝗮𝗹𝗹 𝗟𝗟𝗠𝘀 (𝗲.𝗴. 𝗖𝗵𝗮𝘁𝗚𝗣𝗧). 𝗜𝘁 𝗶𝘀 𝗰𝗿𝘂𝗰𝗶𝗮𝗹 𝘁𝗼 𝗵𝗮𝘃𝗲 𝗮 𝘀𝗼𝗹𝗶𝗱 𝘂𝗻𝗱𝗲𝗿𝘀𝘁𝗮𝗻𝗱𝗶𝗻𝗴 𝗵𝗼𝘄 𝘁𝗵𝗲𝘀𝗲 𝗺𝗲𝗰𝗵𝗮𝗻𝗶𝗰𝘀 𝘄𝗼𝗿𝗸 𝗶𝗳 𝘆𝗼𝘂 𝘄𝗮𝗻𝘁 𝘁𝗼 𝗯𝘂𝗶𝗹𝗱 𝘀𝗰𝗮𝗹𝗮𝗯𝗹𝗲, 𝗿𝗲𝘀𝗽𝗼𝗻𝘀𝗶𝗯𝗹𝗲 𝗔𝗜 𝘀𝗼𝗹𝘂𝘁𝗶𝗼𝗻𝘀. Here is the full video from 3Blue1Brown with exaplantion. I highly recommend to read, watch and bookmark this for a further deep dive: https://lnkd.in/dAviqK_6 𝗜 𝗲𝘅𝗽𝗹𝗼𝗿𝗲 𝘁𝗵𝗲𝘀𝗲 𝗱𝗲𝘃𝗲𝗹𝗼𝗽𝗺𝗲𝗻𝘁𝘀 — 𝗮𝗻𝗱 𝘄𝗵𝗮𝘁 𝘁𝗵𝗲𝘆 𝗺𝗲𝗮𝗻 𝗳𝗼𝗿 𝗿𝗲𝗮𝗹-𝘄𝗼𝗿𝗹𝗱 𝘂𝘀𝗲 𝗰𝗮𝘀𝗲𝘀 — 𝗶𝗻 𝗺𝘆 𝘄𝗲𝗲𝗸𝗹𝘆 𝗻𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿. 𝗬𝗼𝘂 𝗰𝗮𝗻 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲 𝗵𝗲𝗿𝗲 𝗳𝗼𝗿 𝗳𝗿𝗲𝗲: https://lnkd.in/dbf74Y9E
-
Multi-Head Attention (MHA) is the engine of LLMs. But over the years, we have added several tweaks to make it more efficient for long-context settings, especially when using KV caching during inference. I implemented the most common variants from scratch: 1) Grouped-Query Attention (GQA): Instead of having a unique key and value for each query head, multiple queries share the same key and value. As long as the sharing ratio is not too extreme, this has minimal impact on model quality. It is the most widely used variant today and found in almost every modern LLM including Llama 2-4, GPT-OSS, Gemma 3, Qwen 3, GLM 4.6, and many others. 2) Multi-Head Latent Attention (MLA): This variant introduces a compressed latent representation for the keys and values that are stored in the KV cache. During inference, these latent keys and values are up-projected back to the full dimension. The extra projection adds a small computational cost, but the memory savings make it worth it. This approach is currently used by DeepSeek V3 and Kimi K2. 3) Sliding-Window Attention (SWA): SWA restricts each token’s attention span to a fixed local window, which reduces memory needs by shrinking the KV cache in long-context regimes. It is usually applied selectively, for example every other layer, or in Gemma 3's case, five SWA layers for each full-attention layer. While less common today, it remains an important optimization, notably in Gemma 3. All three variants, GQA, MLA, and SWA, can also be combined freely within the same model. Here's a link to check them out: 1️⃣ GQA: https://lnkd.in/grDPXUUi 2️⃣ MLA: https://lnkd.in/gm4FzE32 3️⃣ SWA: https://lnkd.in/g7x-fdgn
-
LLMs process text from left to right — each token can only look back at what came before it, never forward. This means that when you write a long prompt with context at the beginning and a question at the end, the model answers the question having "seen" the context, but the context tokens were generated without any awareness of what question was coming. This asymmetry is a basic structural property of how these models work. The paper asks what happens if you just send the prompt twice in a row, so that every part of the input gets a second pass where it can attend to every other part. The answer is that accuracy goes up across seven different benchmarks and seven different models (from the Gemini, ChatGPT, Claude, and DeepSeek series of LLMs), with no increase in the length of the model's output and no meaningful increase in response time — because processing the input is done in parallel by the hardware anyway. There are no new losses to compute, no finetuning, no clever prompt engineering beyond the repetition itself. The gap between this technique and doing nothing is sometimes small, sometimes large (one model went from 21% to 97% on a task involving finding a name in a list). If you are thinking about how to get better results from these models without paying for longer outputs or slower responses, that's a fairly concrete and low-effort finding. Read with AI tutor: https://lnkd.in/ene242cx Get the PDF: https://lnkd.in/e9tbUTNv
-
The Voice Stack is improving rapidly. Systems that interact with users via speaking and listening will drive many new applications. Over the past year, I’ve been working closely with DeepLearning.AI, AI Fund, and several collaborators on voice-based applications, and I will share best practices I’ve learned in this and future posts. Foundation models that are trained to directly input, and often also directly generate, audio have contributed to this growth, but they are only part of the story. OpenAI’s RealTime API makes it easy for developers to write prompts to develop systems that deliver voice-in, voice-out experiences. This is great for building quick-and-dirty prototypes, and it also works well for low-stakes conversations where making an occasional mistake is okay. I encourage you to try it! However, compared to text-based generation, it is still hard to control the output of voice-in voice-out models. In contrast to directly generating audio, when we use an LLM to generate text, we have many tools for building guardrails, and we can double-check the output before showing it to users. We can also use sophisticated agentic reasoning workflows to compute high-quality outputs. Before a customer-service agent shows a user the message, “Sure, I’m happy to issue a refund,” we can make sure that (i) issuing the refund is consistent with our business policy and (ii) we will call the API to issue the refund (and not just promise a refund without issuing it). In contrast, the tools to prevent a voice-in, voice-out model from making such mistakes are much less mature. In my experience, the reasoning capability of voice models also seems inferior to text-based models, and they give less sophisticated answers. (Perhaps this is because voice responses have to be more brief, leaving less room for chain-of-thought reasoning to get to a more thoughtful answer.) When building applications where I need a more control over the output, I use agentic workflows to reason at length about the user’s input. In voice applications, this means I end up using a pipeline that includes speech-to-text (STT) to transcribe the user’s words, then processes the text using one or more LLM calls, and finally returns an audio response to the user via TTS (text-to-speech). This, where the reasoning is done in text, allows for more accurate responses. However, this process introduces latency, and users of voice applications are very sensitive to latency. When DeepLearning.AI worked with RealAvatar (an AI Fund portfolio company led by Jeff Daniel) to build an avatar of me, we found that getting TTS to generate a voice that sounded like me was not very hard, but getting it to respond to questions using words similar to those I would choose was. Even after much tuning, it remains a work in progress. You can play with it at https://lnkd.in/gcZ66yGM [At length limit. Full text, including latency reduction technique: https://lnkd.in/gjzjiVwx ]
-
For the last couple of years, Large Language Models (LLMs) have dominated AI, driving advancements in text generation, search, and automation. But 2025 marks a shift—one that moves beyond token-based predictions to a deeper, more structured understanding of language. Meta’s Large Concept Models (LCMs), launched in December 2024, redefine AI’s ability to reason, generate, and interact by focusing on concepts rather than individual words. Unlike LLMs, which rely on token-by-token generation, LCMs operate at a higher abstraction level, processing entire sentences and ideas as unified concepts. This shift enables AI to grasp deeper meaning, maintain coherence over longer contexts, and produce more structured outputs. Attached is a fantastic graphic created by Manthan Patel How LCMs Work: 🔹 Conceptual Processing – Instead of breaking sentences into discrete words, LCMs encode entire ideas, allowing for higher-level reasoning and contextual depth. 🔹 SONAR Embeddings – A breakthrough in representation learning, SONAR embeddings capture the essence of a sentence rather than just its words, making AI more context-aware and language-agnostic. 🔹 Diffusion Techniques – Borrowing from the success of generative diffusion models, LCMs stabilize text generation, reducing hallucinations and improving reliability. 🔹 Quantization Methods – By refining how AI processes variations in input, LCMs improve robustness and minimize errors from small perturbations in phrasing. 🔹 Multimodal Integration – Unlike traditional LLMs that primarily process text, LCMs seamlessly integrate text, speech, and other data types, enabling more intuitive, cross-lingual AI interactions. Why LCMs Are a Paradigm Shift: ✔️ Deeper Understanding: LCMs go beyond word prediction to grasp the underlying intent and meaning behind a sentence. ✔️ More Structured Outputs: Instead of just generating fluent text, LCMs organize thoughts logically, making them more useful for technical documentation, legal analysis, and complex reports. ✔️ Improved Reasoning & Coherence: LLMs often lose track of long-range dependencies in text. LCMs, by processing entire ideas, maintain context better across long conversations and documents. ✔️ Cross-Domain Applications: From research and enterprise AI to multilingual customer interactions, LCMs unlock new possibilities where traditional LLMs struggle. LCMs vs. LLMs: The Key Differences 🔹 LLMs predict text at the token level, often leading to word-by-word optimizations rather than holistic comprehension. 🔹 LCMs process entire concepts, allowing for abstract reasoning and structured thought representation. 🔹 LLMs may struggle with context loss in long texts, while LCMs excel in maintaining coherence across extended interactions. 🔹 LCMs are more resistant to adversarial input variations, making them more reliable in critical applications like legal tech, enterprise AI, and scientific research.
-
Have you ever noticed that after a point, the more logs you paste into an LLM, the worse it gets at actually helping you? The responses get vague, it drops important context, or just loses track of the original error. Here’s what happens. LLMs charge by the token. A single log line = 80 to 150 tokens. Of those, 60 to 100 are noise, timestamps, field separators, UUID hyphens, repeated key names. The actual information your LLM needs is maybe just 30 to 40 tokens per line. And the thing is, most of your logs aren't even unique. In a typical 200,000 line log file, the top 10 templates cover 80-95% of all messages, same structure, repeated thousands of times just with a different IP, a different UUID, a slightly different timestamp. So you're not sending your LLM logs. You're sending it the same sentence, repeated thousands of times that the model doesn’t even need. It hits the context limit, starts dropping earlier context, and loses track of the error you pasted at the top. That's the problem CtrlB Decompose fixes at the source. It takes your raw logs and separates what's repeated, the template, from what actually changes, the values. Instead of sending 200,000 lines, you send a few patterns and the variables extracted from them. The model gets more context, not less, because the noise is gone. The result is token reduction of up to 99.9%, so lower cost, and the LLM actually performs better because the signal is cleaner. CtrlB just open-sourced this, and if you're doing any kind of LLM-powered debugging or observability, this is definitely worth your time. You can check it out here: https://lnkd.in/gmgWibHP Ps: Recently visited CtrlB office and was genuinely impressed by what Adarsh Srivastava and his team are building!
-
NVIDIA just exposed the dirty secret about LLMs. A new research paper from NVIDIA shows what many suspected: 👉 Small Language Models (SLMs) can outperform massive LLMs in real-world applications. This flips the current AI playbook on its head. For years, every agentic task — no matter how simple — has been run through massive models like GPT-4 or Claude. NVIDIA’s findings? That approach is wasteful, unnecessary, and about to change. I have a few takeaways that will change how we build AI agents: SLMs are fast, cheap, and effective. Tasks like summarizing docs, extracting info, writing templates, or calling APIs are predictable. For these, SLMs aren’t just “good enough” — they’re better. Smaller ≠ weaker. • Toolformer (6.7B) beats GPT-3 (175B) on API use. • DeepSeek-R1-Distill (7B) outperforms Claude 3.5 and GPT-4o on reasoning. Efficiency is unmatched. • 10–30x cheaper to run • Lower energy use • Faster response times • Easy to deploy locally They’re easy to fine-tune. Techniques like LoRA and QLoRA make overnight customization possible without GPU farms. Perfect fit for structured outputs. SLMs align better with strict formats (JSON, XML, Python) — ideal for agents that need reliability instead of creativity. So why keep running everything through massive LLMs? The smarter path is modular agents: • Default to SLMs. • Call an LLM only when absolutely necessary. This architecture is cheaper, faster, and more controllable. The paper even outlines the migration path: 1. Log usage data 2. Cluster tasks 3. Fine-tune SLMs 4. Replace LLM calls 5. Iterate Why hasn’t the industry switched yet? • Heavy sunk costs in LLM infrastructure • Benchmarks biased toward general tasks • Lack of attention on SLMs But none of these are technical blockers. The future of AI agents isn’t bigger models. It’s smarter architecture. SLMs give you control, speed, and affordability. The paper is worth a read for those at the application layer of AI. Read NVIDIA’s full paper here: https://lnkd.in/gdQRYxyw
-
I have one simple ask, stop calling it "Hallucination" LLMs are working exactly as designed. After writing two books on machine learning and diving deep into the mathematical foundations of large language models, I feel I need to highlight a persistent misconception in the industry: the term "AI hallucination" is not just misleading, it's fundamentally wrong. > What's Really Happening When an LLM generates incorrect information, it's not "hallucinating." It's executing a stochastic sampling process based on learned probability distributions. Every token generated is the result of mathematical operations on attention weights, layer normalizations, and softmax functions operating over vast parameter spaces. The model is doing exactly what it was trained to do: predict the most probable next token given the context, with some degree of randomness introduced through temperature scaling and sampling strategies. There's no perception involved, no false imagery, no cognitive breakdown—just probability mathematics in action. > Why "Hallucination" Misleads Us Anthropomorphic framing creates several problems: 1. It implies malfunction when there isn't one. The model is operating within its design parameters. When it generates plausible sounding but incorrect information, that's an expected outcome of the training process, not a bug. 2. It obscures the real challenges. Instead of focusing on improving training data quality, refining sampling strategies, or developing better uncertainty quantification methods, we get distracted by the metaphor. 3. It sets unrealistic expectations. Stakeholders hear "hallucination" and think the problem is binary that either the AI is working correctly or it's having episodes. The reality is that confidence and accuracy exist on a spectrum defined by the underlying probability distributions. > A Better Framework? What we're actually dealing with is distributional sampling under uncertainty. When training data is sparse for a particular domain, the model's learned representations become less reliable. When the temperature parameter is high, we get more diverse but potentially less accurate outputs. When the context doesn't provide sufficient signal, the model falls back on broader statistical patterns from its training corpus. This isn't pathology, it's statistics. What's your take? Are we ready to retire this misleading terminology in favour of more precise technical language? Or are the blind going to continually lead the blind? #MachineLearning #LLMs #AI #DataScience #TechnicalLeadership
-
Everyone interacts with ChatGPT. But can you build it from scratch? I received a PhD in Machine Learning from MIT in 2022. Then discovered my passion in teaching machine learning from scratch. 3 months back, I started a project to teach “How to Build Large Language Models from scratch” without any libraries! The goal is to empower students and industry professionals to master the building blocks of large language models and ChatGPT. The result is a mega-project with 15 videos covering everything about large language models. I have uploaded all videos on Youtube. Lecture 1: Building LLMs from scratch: Series introduction https://lnkd.in/dFJgqxxf Lecture 2: Large Language Models (LLM) Basics https://lnkd.in/dfACgdPX Lecture 3: Pretraining LLMs vs Finetuning LLMs https://lnkd.in/dga-NhmN Lecture 4: What are transformers? https://lnkd.in/dicb7rEk Lecture 5: How does GPT-3 really work? https://lnkd.in/dxt4GSkS Lecture 6: Stages of building an LLM from Scratch https://lnkd.in/dwfU8R5d Lecture 7: Code an LLM Tokenizer from Scratch in Python https://lnkd.in/d_6XC9PE Lecture 8: The GPT Tokenizer: Byte Pair Encoding https://lnkd.in/d4yZyNsT Lecture 9: Creating Input-Target data pairs using Python DataLoader https://lnkd.in/dbKVPCgt Lecture 10: What are token embeddings? https://lnkd.in/duVKJKvz Lecture 11: The importance of Positional Embeddings https://lnkd.in/dsP7vGJ5 Lecture 12: The entire Data Preprocessing Pipeline of Large Language Models (LLMs) https://lnkd.in/dFHfKPtc Lecture 13: Introduction to the Attention Mechanism in Large Language Models (LLMs) https://lnkd.in/d_kVY-Q2 Lecture 14: Simplified Attention Mechanism - Coded from scratch in Python | No trainable weights https://lnkd.in/dk7U7C5s Lecture 15: Coding the self attention mechanism with key, query and value matrices https://lnkd.in/dYR8u_Fp I have spent a lot of time and effort in making these lectures. I show everything on a whiteboard and then show it through Python code. Nothing is assumed. Everything is spelled out. P.S: Want to learn about all of this live - with me as your course instructor? Join our live bootcamp starting from January 3rd week 2025: https://vizuara.ai/spit/ 100+ students have already registered and we are closing registrations very soon!