lakeFS’ cover photo
lakeFS

lakeFS

Software Development

New York, NY 8,110 followers

The Control Plane for AI-Ready Data

About us

lakeFS is the control plane for AI-ready data, bridging the infrastructure gap that slows down enterprise AI initiatives. Built on a highly scalable data version control architecture, lakeFS ensures data quality, makes AI training and agent runs reproducible, and reduces data access friction for tools, users, and AI agents, without replacing your existing infrastructure. lakeFS sits between your data storage and everything that consumes it: pipelines, models, and agents. It handles unstructured, structured, and metadata across back-end source systems through a single interface, with zero data copying or duplication. Trusted by AI/MLOps teams, data engineers, and data scientists at thousands of organizations including Arm, Bosch, Lockheed Martin, NASA, Volvo, and the U.S. Department of Energy.

Website
https://lakefs.io/
Industry
Software Development
Company size
11-50 employees
Headquarters
New York, NY
Type
Privately Held
Founded
2020

Employees at lakeFS

View 41 employees at lakeFS

or

By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.

See all employees

Locations

Updates

  • The shift from "models as a feature" to "agents that do work" quietly rewrote the infrastructure requirements underneath. Single-turn inference became long-running, multi-step execution. Stateless calls became sessions that need memory. A model that only generated text became software that reads and writes your data and calls external systems. Each of those shifts spun up a new category of infrastructure. That is the story the table is trying to tell. Which assumption did agents break hardest for your team? #AgenticAI #AIInfrastructure

  • View organization page for lakeFS

    8,110 followers

    Red Hat published an AI quickstart for fraud detection on OpenShift AI, with lakeFS managing the data underneath. Data scientists train a new fraud model on an isolated lakeFS branch of production data, without duplicating physical storage. When the model starts flagging false positives, they compare data versions to isolate the cause. If a dataset gets corrupted, they roll back. For financial institutions, being able to recreate the exact dataset a model was trained on is what makes the audit conversation short. The guide walks the full path: deploying MinIO and lakeFS, configuring OpenShift AI, training and registering the model, and serving it with KServe. It was tested with lakeFS 1.73.0. Thanks to the Red Hat team for writing it up! Link in comments. 👇🏼

    • No alternative text description for this image
  • You can now run lakeFS on Impossible Cloud! Because Impossible Cloud speaks S3, the setup is the standard lakeFS S3 blockstore config with a regional endpoint and your keys. Import your existing bucket with zero-copy, and you get branches, commits, merges and reverts on data that never moves. For ML teams, that means a training run can point at an exact commit instead of a path that might have changed since Tuesday. The guide covers the config and the import. Turn off bucket versioning and Object Lock first, since lakeFS manages versions itself. Thanks to the team at Impossible Cloud for their help with this integration!

    View organization page for Impossible Cloud

    9,323 followers

    🔁 AI data needs version control too Impossible Cloud can be configured as a datastore for lakeFS, bringing Git-like version control to data stored in our S3-compatible object storage. With lakeFS, ML teams can branch, commit, merge and revert data without creating full copies of their datasets. That makes it easier to govern changes, reproduce the data behind training runs and give AI workloads access to a specific, traceable version of the data. 🤖 For teams building AI on object storage, this adds an important control layer between the data and the workloads using it. Using lakeFS for AI or data workflows? Here's how to connect it to Impossible Cloud: https://lnkd.in/dkG68pZJ #ImpossibleCloud #lakeFS #ObjectStorage #AIInfrastructure

    • No alternative text description for this image
  • #FabCon and SQLCon Barcelona brought a lot of OneLake news, and we're glad lakeFS got a mention in Microsoft's recap. OneLake shortcuts now give Fabric users a complementary way to access versioned data managed in lakeFS. A shortcut can point at a branch for the latest data, or at a specific commit for a fixed version. We're happy to be working with the Microsoft Fabric team on this, and we're looking forward to what comes next. Microsoft's full recap on what’s new at the event at link in comments. 👇🏼

    • No alternative text description for this image
  • Here are the 15 categories that make up modern agent infrastructure: Compute and infrastructure. Foundation models. Coding agents. Agent harnesses. Agent frameworks. Workflow and orchestration. Agent sandboxes. Memory management. Data connectors and tool integrations. Data storage. Data query and analytics engines. Metadata management and data catalogs. AI gateways and cost control. Observability and evaluation. Governance and compliance. Each one is a foundational capability, not a nice-to-have, once an agent runs real work in production. We put the companies operating in each into a single reference. Find the link to the full interactive table, with a short write-up of every category at the link in comments. #AgenticAI #AIInfrastructure #AIGovernance

    • No alternative text description for this image
    • No alternative text description for this image
  • Most data governance was designed around a person doing the work: someone who moves at human speed and leaves a trail you can read afterward. Autonomous agents change that. They touch data far more often than any approval process expects, and when an agent gets something wrong, one audit log isn't enough. You need two histories: what the agent did, and what happened to your data. On October 22, lakeFS CTO and co-founder Oz Katz and Ravind Kumar from MinIO will walk through what it takes to run agents in production, and how enterprise teams in regulated environments are building that foundation now: → Why human-speed governance can't keep up once agents become a new type of data consumer → Why agent memory should live on infrastructure you control → Where to put governance so you get isolation, reproducibility, and guardrails without slowing teams down → A reference architecture for agentic AI on enterprise data → A live demo of governed agent workflows on MinIO AIStor and lakeFS 📅 October 22 | 12:00 PM ET | Virtual Save your spot at the link in comments 👇🏼

    • No alternative text description for this image
  • Join Ravind Kumar from MinIO and Oz Katz from lakeFS on October 22 at 12:00 PM Eastern Time to learn what it takes to move autonomous agents from pilot to production, what that requires from your data infrastructure, and how enterprise teams in complex and regulated environments are building that foundation now. Register today for the LinkedIn Live Event or Zoom event (link in comments)

    Governing Autonomous AI Agents with MinIO and lakeFS

    Governing Autonomous AI Agents with MinIO and lakeFS

    www.linkedin.com

  • View organization page for lakeFS

    8,110 followers

    Which data trained this model? And was one specific patient's scan in it? Our latest DVC webinar is now up on the DVC YouTube. AWS solutions architects Sandeep Raveesh and Paolo Di Francesco walked through the lineage patterns they built with Amazon SageMaker AI, SageMaker AI MLflow Apps, and DVC by lakeFS. How it fits together: - A SageMaker Processing job reads the raw data and a consent registry, filters out records that opted out, and versions the result with DVC - The training job clones the repo at that Git tag, pulls that exact data, and logs the run to MLflow, which syncs to SageMaker Model Registry - Every link in the chain, from the endpoint back to the consent registry, can be queried Paolo's demo used a public chest X-ray dataset. His audit queries showed which dataset versions a patient appeared in, and confirmed that a patient who opted out was absent from every version after. Regulated teams will recognize one detail: a record in the test split still counts, because it shaped the model's metrics. He also had an AI coding assistant run the same audit against a live endpoint, with no special instructions. Then Joe Pringle, our VP of Customer Success, took the same problem to team scale with lakeFS. He showed branching in place of copying S3 buckets, access control modeled on IAM, pull requests for data changes, and training on a lakeFS tag instead of an S3 path that can change. Thanks to Sandeep and Paolo, and to their blog co-authors Manuwai Korber and Nick McCarthy. Watch the recording at the link in comments.

    • No alternative text description for this image
  • Migrating from on-prem NetApp to AWS shouldn't force you to choose between the NFS and SMB your data scientists rely on and the S3 your AI stack expects. A Fortune 500 pharma team ran into exactly this. The storage answer was Amazon FSx for NetApp ONTAP. But storage protocols were only half the problem. The harder half: when multiple people edit the same clinical trial dataset, how do you keep partially tested work out of production without spawning ten copies named trial-data-final-v2-reviewed? Their approach was to put lakeFS above the storage. Every change happens on a zero-copy branch, gets validated in isolation, and merges only after review, whether the change comes from a person, a pipeline, or an AI agent. The data stays on NetApp. Full write-up here: https://hubs.la/Q04ybWQ10

  • Embeddings are first-class data assets. Most pipelines still treat them like scratch files. When you convert images, audio, and documents into vectors for similarity search and RAG, those vectors are only meaningful next to the exact model, preprocessing code, and source asset that produced them. Break that link and you get silent failures: retrieval quality drops, and there's no error to trace it back to. A few practices keep the semantic layer honest. Store each embedding with a reference to its source asset and the model version that generated it, so any vector traces back to its origin. Re-embed as a versioned operation, not an in-place overwrite. A new model or chunking strategy should produce a new version, not silently replace the old vectors. Validate that every source asset has a corresponding embedding before you promote a dataset, so a partial re-embedding run gets caught before production. Because vector stores like LanceDB read from object storage, lakeFS can version the embedding files alongside the images, transcripts, and metadata they came from, keeping every vector aligned with the source that produced it. Link to the full article in the comments. 👇🏽 #RAG #VectorSearch #DataVersion#Control

    • No alternative text description for this image

Similar pages

Browse jobs