lakefs

module

v1.16.0 Latest Latest Go to latest Published: Apr 3, 2024 License: Apache-2.0

Details

Valid go.mod file

The Go module system was introduced in Go 1.11 and is the official dependency management solution for Go.
Redistributable license

Redistributable licenses place minimal restrictions on how software can be used, modified, and redistributed.
Tagged version

Modules with tagged versions give importers more predictable builds.
Stable version

When a project reaches major version v1 it is considered stable.
Learn more about best practices

Repository

github.com/treeverse/lakefs

Links

Open Source Insights

README ¶

lakeFS is Data Version Control (Git for Data)

lakeFS is an open-source tool that transforms your object storage into a Git-like repository. It enables you to manage your data lake the way you manage your code.

With lakeFS you can build repeatable, atomic, and versioned data lake operations - from complex ETL jobs to data science and analytics.

lakeFS supports AWS S3, Azure Blob Storage, and Google Cloud Storage as its underlying storage service. It is API compatible with S3 and works seamlessly with all modern data frameworks such as Spark, Hive, AWS Athena, DuckDB, and Presto.

For more information, see the documentation.

Getting Started

You can spin up a standalone sandbox instance of lakeFS using Docker:

docker run --pull always \
		   --name lakefs \
		   -p 8000:8000 \
		   treeverse/lakefs:latest \
		   run --quickstart

Once you've got lakeFS running, open http://127.0.0.1:8000/ in your web browser.

Quickstart

👉🏻 For a hands-on walk through of the core functionality in lakeFS head over to the quickstart to jump right in!

Make sure to also have a look at the lakeFS samples. These are a rich resource of examples of end-to-end applications that you can build with lakeFS.

Why Do I Need lakeFS?

ETL Testing with Isolated Dev/Test Environment

When working with a data lake, it’s useful to have replicas of your production environment. These replicas allow you to test these ETLs and understand changes to your data without impacting downstream data consumers.

Running ETL and transformation jobs directly in production without proper ETL Testing is a guaranteed way to have data issues flow into dashboards, ML models, and other consumers sooner or later. The most common approach to avoid making changes directly in production is to create and maintain multiple data environments and perform ETL testing on them. Dev environment to develop the data pipelines and test environment where pipeline changes are tested before pushing it to production. With lakeFS you can create branches, and get a copy of the full production data, without copying anything. This enables a faster and easier process of ETL testing.

Reproducibility

Data changes frequently. This makes the task of keeping track of its exact state over time difficult. Oftentimes, people maintain only one state of their data––its current state.

This has a negative impact on the work, as it becomes hard to:

Debug a data issue.
Validate machine learning training accuracy (re-running a model over different data gives different results). Comply with data audits.

In comparison, lakeFS exposes a Git-like interface to data that allows keeping track of more than just the current state of data. This makes reproducing its state at any point in time straightforward.

CI/CD for Data

Data pipelines feed processed data from data lakes to downstream consumers like business dashboards and machine learning models. As more and more organizations rely on data to enable business critical decisions, data reliability and trust are of paramount concern. Thus, it’s important to ensure that production data adheres to the data governance policies of businesses. These data governance requirements can be as simple as a file format validation, schema check, or an exhaustive PII(Personally Identifiable Information) data removal from all of organization’s data.

Thus, to ensure the quality and reliability at each stage of the data lifecycle, data quality gates need to be implemented. That is, we need to run Continuous Integration(CI) tests on the data, and only if data governance requirements are met can the data can be promoted to production for business use.

Everytime there is an update to production data, the best practice would be to run CI tests and then promote(deploy) the data to production. With lakeFS you can create hooks that make sure that only data that passed these tests will become part of production.

Rollback

A rollback operation is used to to fix critical data errors immediately.

What is a critical data error? Think of a situation where erroneous or misformatted data causes a signficant issue with an important service or function. In such situations, the first thing to do is stop the bleeding.

Rolling back returns data to a state in the past, before the error was present. You might not be showing all the latest data after a rollback, but at least you aren’t showing incorrect data or raising errors. Since lakeFS provides versions of the data without making copies of the data, you can time travel between versions and roll back to the version of the data before the error was presented.

Community

Stay up to date and get lakeFS support via:

Share your lakeFS experience and get support on our Slack.
Follow us and join the conversation on Twitter and Mastodon.
Learn from video tutorials on our YouTube channel.
Read more on data versioning and other data lake best practices in our blog.
Feel free to contact us about anything else.

More information

Read the documentation.
See the contributing guide.
Take a look at our roadmap to peek into the future of lakeFS.

Licensing

lakeFS is completely free and open-source and licensed under the Apache 2.0 License.

Who Uses lakeFS?

lakeFS is used by numerous companies, including those below. If you use lakeFS and would like to be included here please open a PR.

AirAsia
APEX Global
AppsFlyer
Auburn University
BAE Systems
Bureau of Labor Statistics
Cambridge Consultants
Connor, Clark & Lunn Financial Group
Context Labs Bv
Daimler Truck
Enigma
EPCOR
Ford Motor Company
Generali
Giesecke+Devrient
greehill
Karius
Lockheed Martin
Luxonis
Mixpeek
Netflix
Paige
PETRONAS
Pollinate
Proton Technologies AG
ProtonMail
Renaissance Computing Institute
RHEA Group
RMS
Sensum
Similarweb
State Street Global Advisors
Terramera
Tredence
Volvo Cars
Webiks
Windward
Woven by Toyota

Directories ¶

Path	Synopsis
cmd
lakectl
lakectl/cmd
lakefs
lakefs-loadtest
lakefs-loadtest/cmd
lakefs/cmd
esti
pkg
actions
actions/lua
actions/lua/crypto/aes
actions/lua/crypto/hmac
actions/lua/crypto/md5
actions/lua/crypto/sha256
actions/lua/databricks
actions/lua/encoding/base64
actions/lua/encoding/hex
actions/lua/encoding/json
actions/lua/encoding/parquet
actions/lua/encoding/yaml
actions/lua/formats
actions/lua/lakefs
actions/lua/net/http
actions/lua/net/url
actions/lua/path
actions/lua/regexp
actions/lua/storage/aws
actions/lua/storage/azure
actions/lua/storage/gcloud
actions/lua/strings
actions/lua/time
actions/lua/util
actions/lua/uuid
actions/mock Package mock is a generated GoMock package.	Package mock is a generated GoMock package.
api
api/apigen Package apigen provides generated code for our OpenAPI	Package apigen provides generated code for our OpenAPI
api/apiutil
api/helpers Package helpers provide useful wrappers for clients using the lakeFS OpenAPI.	Package helpers provide useful wrappers for clients using the lakeFS OpenAPI.
api/params
auth
auth/acl
auth/crypt
auth/keys
auth/mock Package mock is a generated GoMock package.	Package mock is a generated GoMock package.
auth/model
auth/oidc/encoding Package encoding defines Claims for interoperable external services to use in JWTs.	Package encoding defines Claims for interoperable external services to use in JWTs.
auth/params
auth/remoteauthenticator
auth/setup
auth/testutil
auth/wildcard
authentication
authentication/mock Package mock is a generated GoMock package.	Package mock is a generated GoMock package.
batch
block
block/azure
block/blocktest
block/factory
block/gs
block/local
block/mem
block/params
block/s3
block/transient
cache
catalog
catalog/testutils
cloud
cloud/aws
cloud/azure
cloud/gcp
cmdutils
config
diff
dockertest Package dockertest provides a wrapper for running a lakeFS container in integration tests.	Package dockertest provides a wrapper for running a lakeFS container in integration tests.
fileutil
gateway
gateway/errors
gateway/multipart
gateway/operations
gateway/path
gateway/serde
gateway/sig Package sig This file implements helper functions to validate Streaming AWS Signature Version '4' authorization header.	Package sig This file implements helper functions to validate Streaming AWS Signature Version '4' authorization header.
gateway/testutil
git
graveler
graveler/branch
graveler/committed
graveler/committed/mock Package mock is a generated GoMock package.	Package mock is a generated GoMock package.
graveler/mock Package mock is a generated GoMock package.	Package mock is a generated GoMock package.
graveler/ref
graveler/retention
graveler/settings
graveler/sstable
graveler/staging
graveler/testutil
httputil
ident
ingest/store
kv
kv/cosmosdb
kv/dynamodb
kv/kvparams
kv/kvtest
kv/local
kv/mem
kv/migrations
kv/mock Package mock is a generated GoMock package.	Package mock is a generated GoMock package.
kv/postgres
loadtest
local
logging
metastore
metastore/errors
metastore/glue
metastore/hive
metastore/hive/gen-go/hive_metastore
metastore/hive/gen-go/hive_metastore/thrift_hive_metastore-remote
metastore/mock
mock Package mock is a generated GoMock package.	Package mock is a generated GoMock package.
permissions Code generated by extract_actions.	Code generated by extract_actions.
pyramid
pyramid/mock Package mock is a generated GoMock package.	Package mock is a generated GoMock package.
pyramid/params
samplerepo
samplerepo/assets
stats
testutil
testutil/stress
upload
uri
validator
version
webui

?	: This menu
/	: Search site
f or F	: Jump to
y or Y	: Canonical URL