Of course you need to use open-source models if you’re an enterprise leader. Close model providers, that are now forcing data retention, are gaining immense leverage on your business if you don’t. As you connect models to your business context, they see it and learn from it, and have a track record of going after their most successful customers thanks to this information. But that’s not enough, you also need to store your data and records in open systems, or your software vendors might block you from building AI systems outside of the walled garden they have set up for you. If you can’t convince them to give you complete access to the data they manage for you, AI fortunately allows you to migrate quite fast. Once you’ve got hold of your data, you’ll need to manage how AI systems can access this data on behalf of human users, because you don’t always want Bob to see what Alice is doing in your company. That’s hard and merciless, since AI models are great at finding need-to-know errors. It takes systems that check hard access rules and models that check soft access rules. Now comes the most important part. You need to set up your own continuous training flywheel, so that you can improve your AI systems based on their interaction with your employees and your users. This is how you turn the edges of your business into AI systems your vendors and competitors cannot replicate. It’s also how you reduce deployment cost as well, as you can shrink models according to model input distribution. Those bills are getting substantial, we need to collectively become efficient if we want AI development to continue, so that matters. All of these efforts might seem daunting – they are. This is both a complete replatforming of your IT, and a complete change in the way you’re developing software, and operating your business. AI lifecycle management requires understanding human behavior and gradient descent, that’s a stretch. At Mistral, we facilitate that work by providing all primitives that you need in a single control plane, Studio, and a training platform, Forge. With our applied AI engineers and scientists working hand-in-hand with our customers, we ensure that we transfer knowledge, and that we can disappear once the systems are up and running. We deploy on our customers' infrastructure, or through our zero-data-retention hosted services, so that your edges remain your edges, and the switch button can be fully in your hand. Frontier AI can accelerate the growth of your business, but if it’s not in your hands, it’s not going to be your growth.
AI Safety and Risk Management
Explore top LinkedIn content from expert professionals.
-
-
𝗡𝗮𝗶𝘃𝗲 𝗥𝗔𝗚 𝘄𝗼𝗿𝗸𝘀 𝗶𝗻 𝗮 𝗱𝗲𝗺𝗼. 𝗜𝘁 𝗳𝗮𝗶𝗹𝘀 𝘁𝗵𝗲 𝗺𝗼𝗺𝗲𝗻𝘁 𝗿𝗲𝗮𝗹 𝘂𝘀𝗲𝗿𝘀 𝘀𝗵𝗼𝘄 𝘂𝗽. Embed → retrieve → generate looks clean in a notebook. Real requirements break it: → Questions whose answer is spread across many documents → Industry terms that embeddings get wrong → Bad chunks the pipeline never catches → Answers that live in how things connect, not in any single chunk → PDFs full of tables and images a text-only index cannot read These 5 architectures are how serious teams stay ahead in the agentic AI era: 𝟬𝟭 𝗛𝘆𝗯𝗿𝗶𝗱 𝗥𝗔𝗚 → Dense vectors find meaning. BM25 finds exact words. → Reciprocal Rank Fusion combines both ranked lists. → A safe baseline for almost every team. 𝟬𝟮 𝗚𝗿𝗮𝗽𝗵𝗥𝗔𝗚 → Pull entities and their relationships into a knowledge graph. → Retrieve subgraphs and community summaries, not chunks. → Best when the answer lives in how things connect. 𝟬𝟯 𝗔𝗴𝗲𝗻𝘁𝗶𝗰 𝗥𝗔𝗚 → A planner agent picks the right tool: vector, web, or SQL. → A reasoner agent keeps trying until the answer is solid. → Retrieval becomes a plan, not a single step. 𝟬𝟰 𝗖𝗼𝗿𝗿𝗲𝗰𝘁𝗶𝘃𝗲 𝗥𝗔𝗚 (𝗖𝗥𝗔𝗚) → Grade every retrieval before you trust it. → Correct → answer. Unclear → rewrite the query. Wrong → search the web. → This is what production RAG actually looks like. 𝟬𝟱 𝗠𝘂𝗹𝘁𝗶𝗺𝗼𝗱𝗮𝗹 𝗥𝗔𝗚 → One embedding model (CLIP, ColPali) for text, images, and tables. → One vector index. One multimodal LLM. → No more separate pipelines for PDFs with charts. I built a runnable example for each of the five patterns. GitHub link in the first comment. The best teams in 2026 do not pick one. They combine them — hybrid retrieval inside an agentic loop, with a corrective grader, over a multimodal index. Naive RAG is a starting point, not a finish line. That is why most enterprise GenAI projects stall at the demo. Which of these five becomes the default RAG stack in the next 18 months — and which stays a specialized tool?
-
🪂 How To Make Your Design System AI-Ready (https://lnkd.in/dtnpy7CM), a practical guide on how to reduce drifts, minimize mistakes, maintain context and improve the quality of AI-generated prototypes — with structured spec files, automated auditing and token layers. Put together by Hardik Pandya from Atlassian. --- 🔹 1. Design Decisions Are Infrastructure AI-generated prototypes often don't deliver consistently decent results because of tiny inconsistencies scattered all across a design system. Often it's decisions made but not documented, hard-coded values never cleaned up, or relying too much on AI making sense of mock-ups or design flows on its own. Unsurprisingly, better AI prototypes come from better data — but also from better human guidance. We shouldn’t assume that AI knows how to choose the right component, and how to design with accessibility in mind. It needs priorities, a clear path on how we make decisions, design principles, examples, do's and don'ts. In fact, we should treat design decisions as infrastructure. That means that every time we make a decision — not just a design decision, but even decision on how actually prioritize our work and how we make decisions around here — it must find a path into the spec file that is then consumed by AI. --- 🔶 2. Three Layers: Spec Files + Token Layer + Audit To ensure quality, we establish design principles, guidelines, rules in a form of “spec files”). It's structured Markdown files that include spacing rules, color choices, component usage guidelines, priorities etc. AI is going to read and reuse that spec file every time it's going to generate a prototype. Because the spec files are text files, it's much more cost-effective, but also much more accurate just because we don't rely on AI recognizing or decoding patterns from mock-ups, but gets specific guidelines instead. In fact, extending code is often a more effective way than generating code from mock-ups. Token layer lists and keeps updated all tokens used throughout the design system. AI always chooses from a closed set of named variables instead of inventing plausible values ad-hoc. An audit script catches what AI gets wrong. It scans the prototype and flags every hard-coded value and flags it if necessary. It can be a regular software doing that, with AI waiting for its feedback to come back. Finally, when a design system ships updates, a sync routine flags which spec files need updating. The goal is to make sure that AI always reads up-to-date, current specs, not the ones written against an outdated version. --- 🔺 3. Examples of AI-Ready Design Systems ⌾ Atlassian: https://lnkd.in/dVsGc3Cp ⌾ Carbon: https://lnkd.in/d4zq4WWb ⌾ CMS Design System: https://lnkd.in/dHHzV3en ⌾ Nordhealth: https://lnkd.in/d8C4j2ZA Yet again, AI can’t magically resolve technical debt or design debt — it needs guidance, decisions, priorities and principles.
-
AI security/securing the use of AI is going to kill me. I use Claude Code almost daily. It's a problem.... Here's what I have to change AGAIN this week. Security researcher Ari Marzuk disclosed 30+ vulnerabilities across AI coding tools. Cursor. GitHub Copilot. Windsurf. Claude Code. All of them. He called it IDEsaster. The attack chain includes prompt injection, hijacking LLM context, and auto-approved tool calls executing without permission. Then, legitimate IDE features are weaponized for data exfiltration and RCE. Your .env files. Your API keys. Your source code. Accessible through features you thought were safe. Most studies I read claim that around 85% of developers now use AI coding tools daily. Most have no idea their IDE treats its own features as inherently trusted. 𝗦𝗼... 𝗮𝗳𝘁𝗲𝗿 𝗿𝗲𝘃𝗶𝗲𝘄𝗶𝗻𝗴 𝗔𝗿𝗶'𝘀 𝗿𝗲𝘀𝗲𝗮𝗿𝗰𝗵, 𝗵𝗲𝗿𝗲'𝘀 𝗜 𝘄𝗶𝗹𝗹 𝗯𝗲 𝗱𝗼𝗶𝗻𝗴... Be warned: All this is SO much easier said than done! Audit every MCP server connection. Checked for tool poisoning vectors where legitimate tools might parse attacker-controlled input from GitHub PRs or web content. Removed servers I couldn't verify. Disabled auto-approve for file writes. The attack chains weaponize configuration files and project instructions like .claude/settings.json and CLAUDE.md. One malicious write to these files can alter agent behavior or achieve code execution without additional user interaction. Move all credentials to a secrets manager. No .gitignored .env files in agent-accessible directories. API keys live in 1Password CLI. Environment variables inject at runtime through a wrapper script the LLM never sees. Start running Claude Code in isolated containers. Mounted volumes limited to specific project directories. No access to ~/.ssh, ~/.aws, or ~/.config. If the agent gets compromised, blast radius stays contained. Enable all security warnings. Claude Code added explicit warnings for JSON schema exfiltration and settings file modifications. These exist because Anthropic knows the attack surface. Add pre-commit hooks for hidden characters. Prompt injections hide in pasted URLs, READMEs, and file names using invisible Unicode. Flag non-ASCII characters in any file the agent might ingest. The fix isn't to stop using AI coding tools. The fix is to stop trusting them implicitly. What controls do you have for AI tools with write access to your codebase? 👉 Follow for more AI and cybersecurity insights with the occasional rant #AISecurity #DevSecOps
-
Microsoft just released a 35-page report on medical AI - and it’s a reality check for healthcare. The paper, “The Illusion of Readiness”, tested six of the most popular models (OpenAI, Gemini, etc)… across six multimodal medical benchmarks. And the verdict? The models scored high on medical exams. But they’re not even close to being real-world ready. Here’s what the stress tests revealed: ▶ 1. Shortcut learning Models often answered correctly even when key information, like medical images, was removed. They weren’t reasoning - they were exploiting statistical shortcuts. That means benchmark wins may hide shallow understanding. ▶ 2. Fragile under small changes Making small tweaks caused big swings in predictions. This fragility shows how unreliable model reasoning becomes under stress. In visual substitution tests, accuracy dropped from 83% to 52% when images were swapped - exposing shallow visual–answer pairings. ▶ 3. Fabricated reasoning Models produced confident, step-by-step medical explanations - but many were medically unsound… or entirely fabricated. Convincing to the eye, dangerous in practice. And more importantly, healthcare isn’t a multiple-choice exam. It’s uncertainty, incomplete data, and high stakes. So Microsoft’s team calls for new standards: - Stress tests that expose fragility - Clinician-guided guidelines that profile benchmarks - Evaluation of robustness and trustworthiness - not just leaderboard scores The takeaway is simple: Medical AI may ace tests today. But until it proves reliable under stress, it’s not ready for the clinic. When do you think popular LLMs will be clinic-ready? #entrepreneurship #healthtech #AI
-
🚨 BREAKING: OpenAI has just launched ChatGPT Agent. Below are important privacy & security risks everybody should be aware of: Agentic AI applications differ from non-agentic ones particularly regarding the access rights/permissions they require in order to engage with external tools on the user's behalf. The more autonomous the agent, the more permissions/access rights it will require. For example, if a user wants an AI agent to search and buy a dress for them without asking further questions, besides accessing the internet, the AI agent will need to access their wallet. If the user wants the agent to schedule an event and invite friends, it will need access to at least the calendar and contact list. Having said that, any permission given to a third-party app or system has potential privacy and security risks. ChatGPT already presents privacy risks due to the way it's trained, the way it processes personal data, the user's privacy settings, and the type of personal information being input by users. The privacy risks from ChatGPT Agent will be exponentially higher as many people will be giving access rights to external tools containing personal information (calendar, email, wallet, and more). OpenAI knows that malicious actors will try to trick other people's AI agents into sharing private information, including address, email, phone, credit card information, and more. Sam Altman has just posted on X, recommending that people give agents "the minimum access required to complete a task." In many cases, the privacy and security risks of letting an AI agent perform a task will greatly outweigh any productivity benefits it can offer (but people will use AI agents anyway, because of hype, curiosity, or because their company is "AI first") Unfortunately, the pace of AI development is much faster than the pace of AI literacy. Most people haven't yet understood ChatGPT's privacy risks, but they will be thrown a new feature with exponentially MORE risks. - 👉 Never miss my analyses on AI: join my newsletter's 68,200+ subscribers (below).
-
McKinsey & Company 𝗮𝗻𝗮𝗹𝘆𝘇𝗲𝗱 𝟭𝟱𝟬+ 𝗲𝗻𝘁𝗲𝗿𝗽𝗿𝗶𝘀𝗲 𝗚𝗲𝗻𝗔𝗜 𝗱𝗲𝗽𝗹𝗼𝘆𝗺𝗲𝗻𝘁𝘀 — 𝗮𝗻𝗱 𝗳𝗼𝘂𝗻𝗱 𝗼𝗻𝗲 𝗰𝗼𝗺𝗺𝗼𝗻 𝘁𝗵𝗿𝗲𝗮𝗱: ⬇️ One-off solutions don’t scale. The most successful projects take a different path: They use open, modular architectures that enable speed, reuse, and control. → Designed for reuse → Able to plug in best-in-class capabilities → Free from vendor lock-in This is the reference architecture McKinsey now recommends — optimized to scale what works while staying compliant. It consists of five core components: ⬇️ 𝟭. 𝗦𝗲𝗹𝗳-𝘀𝗲𝗿𝘃𝗶𝗰𝗲 𝗽𝗼𝗿𝘁𝗮𝗹: → A secure, compliant “pane of glass” where teams can launch, monitor, and manage GenAI apps. → Preapproved patterns, validated capabilities, shared libraries. → Observability and cost controls built-in. 𝟮. 𝗢𝗽𝗲𝗻 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲 → Services are modular, reusable, and provider-agnostic. → Core functions like RAG, chunking, or prompt routing are shared across apps. → Infra and policy as code, built to evolve fast. 𝟯. 𝗔𝘂𝘁𝗼𝗺𝗮𝘁𝗲𝗱 𝗴𝗼𝘃𝗲𝗿𝗻𝗮𝗻𝗰𝗲 𝗴𝘂𝗮𝗿𝗱𝗿𝗮𝗶𝗹𝘀 → Every prompt and response is logged, audited, and cost-attributed. → Hallucination detection, PII filters, bias audits — enforced by default. → LLMs accessed only through a centralized AI gateway. 4. 𝗙𝘂𝗹𝗹-𝘀𝘁𝗮𝗰𝗸 𝗼𝗯𝘀𝗲𝗿𝘃𝗮𝗯𝗶𝗹𝗶𝘁𝘆 → Centralized logging, analytics, and monitoring across all solutions → Built-in lifecycle governance, FinOps, and Responsible AI enforcement → Secure onboarding of use cases and private data controls → Enables policy adherence across infrastructure, models, and apps 5. 𝗣𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝗼𝗻-𝗴𝗿𝗮𝗱𝗲 𝗨𝘀𝗲 𝗖𝗮𝘀𝗲𝘀 → Modular setup for user interface, business logic, and orchestration → Integrated agents, prompt engineering, and model APIs → Guardrails, feedback systems, and observability built into the solution → Delivered through the AI Gateway for consistent compliance and scale The message is clear: If your GenAI program is stuck, don’t look at the LLM. Look at your platform. 𝗜 𝗲𝘅𝗽𝗹𝗼𝗿𝗲 𝘁𝗵𝗲𝘀𝗲 𝗱𝗲𝘃𝗲𝗹𝗼𝗽𝗺𝗲𝗻𝘁𝘀 — 𝗮𝗻𝗱 𝘄𝗵𝗮𝘁 𝘁𝗵𝗲𝘆 𝗺𝗲𝗮𝗻 𝗳𝗼𝗿 𝗿𝗲𝗮𝗹-𝘄𝗼𝗿𝗹𝗱 𝘂𝘀𝗲 𝗰𝗮𝘀𝗲𝘀 — 𝗶𝗻 𝗺𝘆 𝘄𝗲𝗲𝗸𝗹𝘆 𝗻𝗲𝘄𝘀𝗹𝗲𝘁𝘁𝗲𝗿. 𝗬𝗼𝘂 𝗰𝗮𝗻 𝘀𝘂𝗯𝘀𝗰𝗿𝗶𝗯𝗲 𝗵𝗲𝗿𝗲 𝗳𝗼𝗿 𝗳𝗿𝗲𝗲: https://lnkd.in/dbf74Y9E
-
AI is not failing because of bad ideas; it’s "failing" at enterprise scale because of two big gaps: 👉 Workforce Preparation 👉 Data Security for AI While I speak globally on both topics in depth, today I want to educate us on what it takes to secure data for AI—because 70–82% of AI projects pause or get cancelled at POC/MVP stage (source: #Gartner, #MIT). Why? One of the biggest reasons is a lack of readiness at the data layer. So let’s make it simple - there are 7 phases to securing data for AI—and each phase has direct business risk if ignored. 🔹 Phase 1: Data Sourcing Security - Validating the origin, ownership, and licensing rights of all ingested data. Why It Matters: You can’t build scalable AI with data you don’t own or can’t trace. 🔹 Phase 2: Data Infrastructure Security - Ensuring data warehouses, lakes, and pipelines that support your AI models are hardened and access-controlled. Why It Matters: Unsecured data environments are easy targets for bad actors making you exposed to data breaches, IP theft, and model poisoning. 🔹 Phase 3: Data In-Transit Security - Protecting data as it moves across internal or external systems, especially between cloud, APIs, and vendors. Why It Matters: Intercepted training data = compromised models. Think of it as shipping cash across town in an armored truck—or on a bicycle—your choice. 🔹 Phase 4: API Security for Foundational Models - Safeguarding the APIs you use to connect with LLMs and third-party GenAI platforms (OpenAI, Anthropic, etc.). Why It Matters: Unmonitored API calls can leak sensitive data into public models or expose internal IP. This isn’t just tech debt. It’s reputational and regulatory risk. 🔹 Phase 5: Foundational Model Protection - Defending your proprietary models and fine-tunes from external inference, theft, or malicious querying. Why It Matters: Prompt injection attacks are real. And your enterprise-trained model? It’s a business asset. You lock your office at night—do the same with your models. 🔹 Phase 6: Incident Response for AI Data Breaches - Having predefined protocols for breaches, hallucinations, or AI-generated harm—who’s notified, who investigates, how damage is mitigated. Why It Matters: AI-related incidents are happening. Legal needs response plans. Cyber needs escalation tiers. 🔹 Phase 7: CI/CD for Models (with Security Hooks) - Continuous integration and delivery pipelines for models, embedded with testing, governance, and version-control protocols. Why It Matter: Shipping models like software means risk comes faster—and so must detection. Governance must be baked into every deployment sprint. Want your AI strategy to succeed past MVP? Focus and lock down the data. #AI #DataSecurity #AILeadership #Cybersecurity #FutureOfWork #ResponsibleAI #SolRashidi #Data #Leadership
-
The AI bank of the future ! To embed AI seamlessly across the enterprise, banks can implement a comprehensive capability stack that goes beyond just AI models. This AI bank stack contains four key capability layers: engagement, decision making, data and core tech, and operating model. Each layer will need to receive investment and attention to unlock the full power of AI for the enterprise. Each layer’s foundational elements are supplemented by several new elements. To create sustainable value, banks need to put AI first and revamp the entire technology stack. The rise of innovative technologies such as gen AI has prompted an update to the technology stack from a previous version published in 2020, with new elements highlighted in shades of blue. — Engagement layer Banks will need to reimagine how they engage with customers, making their experiences as intelligent, personalized, and frictionless as possible through the use of AI. Leading banks’ customers are experiencing human-like conversational interactions with AI via text and voice chats and are moving seamlessly across channels such as mobile apps, websites, branches, and contact centers, thanks to powerful AI capabilities. — AI-powered decision-making layer The brain of the bank, this layer makes and orchestrates decisions. Historically, banks have focused on deploying traditional analytics modules such as models, but as AI technologies mature, this layer has expanded to include agent and AI orchestration sublayers working in unison with the traditional analytics layer to drive superior outcomes. — Core technology and data layer This layer includes the technology and data needed for an AI transformation, including reusable tools and pipelines equipped with machine learning operations capabilities needed to run large language models (LLMs) at scale. Other portions of this layer include the data needed to train multiagent systems, as well as modern application programming interface (API) architecture and robust cybersecurity. — Operating model By integrating business and technology in platforms run by cross-functional teams, banks can break up organizational silos, boost agility and speed, and better align goals and priorities across the enterprise. An AI control tower tracks the value realized from AI initiatives, among other tasks. — All together now Elements across the four layers of the AI bank stack work together to enable transformative change and deliver value for the enterprise. Bottomline - Enabling value through an AI stack powered by multiagent systems ; To gain material value from AI, banks need to move beyond experimentation to transform critical business areas, including by reimagining complex workflows with multiagent systems. Source - https://lnkd.in/dd4Cen3B