Skip to content

Concepts ​

LakeSail is built from a small number of objects, arranged in layers. Knowing which layer you are looking at is most of what you need to navigate the console.

Infrastructure
Cloud account
Network · a VPC
ClusterKubernetes compute
Data & compute
Catalog
A pointer to your tables. The data stays in your own storage.
Compute profile
Which cluster, what size, which libraries and catalogs.
Workloads
Jobs, sessions & notebooks
What you define and run, each through a compute profile.
each execution
Runs
Status, timing, logs, and output.

Workloads run on a cluster, against a catalog. Infrastructure nests inward: a cluster lives in a network, which lives in your cloud account.

Organization ​

Everything lives in an organization: its members, teams, roles, and every resource below. You will not think about it as a solo user. It matters the moment a second person joins, because roles are where "who can create a cluster in production?" gets answered.

Infrastructure ​

Compute runs in your AWS account, not LakeSail's, and it is split into three nested layers because each has a different lifecycle.

  • A cloud account is a trust relationship: an IAM role in your AWS account that LakeSail can assume. One account can host many networks.
  • A network is a VPC that LakeSail provisions in that account, in one region. It is its own layer so you can run several clusters, such as dev and prod, without rebuilding networking, and delete a cluster without tearing down the VPC. Each network has a workspace bucket for job files, notebook contents, and results.
  • A cluster is Kubernetes compute where jobs, sessions, and notebooks run. The console calls the one it provisions in your network an external cluster: external to LakeSail, not to you. Compute is added per workload, so the cluster is not sized for the largest job you will ever run.

Two separate things control access to a cluster, and they answer different questions. Its access policy decides what can reach the cluster endpoint on the network: private-only (reachable only from inside the VPC) or public with a CIDR allowlist. Its teams decide who may run work on it. A job or notebook can only be assigned to a team the cluster includes. A compute profile has no team of its own, so it needs only a cluster that has teams, reached either through an organization role or through membership of one of them. Creating a cluster grants access to the creator's teams, so a solo user never meets this; it starts to matter once a second team exists. You can change both later without rebuilding the cluster. See Control which teams can use a cluster.

Where your data lives ​

A catalog tells the engine where your tables are. Connecting one copies nothing; the data stays in your storage. Catalogs are separate from clusters so that one cluster can read from several catalogs and several clusters can read the same data. The pairing is chosen per workload.

Compute profiles ​

A compute profile answers the question every workload asks: which cluster, what size, which libraries, which catalogs. You define it once and point jobs, sessions, and notebooks at it. It is the link between the workload layer and everything above it.

Workloads ​

Three kinds of workload, all on the same clusters, reading the same catalogs, governed by the same roles:

  • Jobs are SQL or Python you define once and run many times, on a schedule or on demand. Each execution is a run with its own status, timing, and output. SQL jobs can reference a saved query so several jobs share one statement.
  • Sessions are live Spark Connect endpoints for your own clients.
  • Notebooks run in the console, backed by a session.

Can't find the answer here? Email us: support@lakesail.com