Skip to content

Frequently Asked Questions ​

About LakeSail ​

What is the LakeSail Platform?

LakeSail is a fully managed data and AI platform that runs in your own AWS account, on a Rust-native engine with no JVM. You keep your data in S3 in open formats and your compute in your VPC; LakeSail runs the control plane: the console, API, job scheduler, catalog registry, access control, and monitoring. You write PySpark and Spark SQL the way you do today.

What's the difference between Sail and the LakeSail Platform?

Sail is the open-source, Rust-native engine that runs every job, query, and session. LakeSail is the managed platform around it: it provisions and scales the infrastructure in your own AWS account and adds the console, API, scheduler, catalogs, access control, and monitoring. You can run Sail yourself for free; LakeSail is what turns it into a production platform you don't have to operate.

Which clouds does LakeSail run in?

AWS today. LakeSail provisions compute and networking in your own AWS account and reads and writes your data in S3. Other clouds are not supported at this time.

How it compares ​

How is it different from Databricks, EMR, or running Spark myself?

Same PySpark and Spark SQL, but the engine is Rust with no JVM, so it starts in seconds and needs less memory. It runs in your own AWS account rather than a vendor's, over open table formats, and you pay for the compute your jobs use with no proprietary compute-unit markup. Against self-managed Spark, LakeSail automates provisioning, scaling, tear-down, scheduling, and access control.

Is LakeSail's engine built on Apache Spark, or a separate engine?

It is a separate engine, written in Rust, that implements Spark's API and speaks the Spark Connect protocol. You keep your PySpark, DataFrame, and Spark SQL; underneath, there is no JVM.

Why not contribute the changes to Spark instead? Because the change that matters most, removing the JVM, means replacing Spark's shuffle, executor lifecycle, memory model, and Python bridge, the internals every connector and workload depends on; that is a different engine, not a patch.

Accelerators like Photon and Comet take the other route and run some operators in native code, but they still live inside Spark's JVM, so its costs (garbage-collection pauses, allocation overhead, Py4J serialization for Python, heavy container footprints) stay for the executor, planning, shuffle, and Python bridge, and every native-to-JVM handoff pays a conversion cost.

Removing the JVM entirely is what makes native columnar execution and in-process Python possible end to end. Keeping Spark's API, the language millions of engineers already use, means you get that engine without rewriting your applications.

How is this different from a Spark accelerator like Photon or Comet?

Accelerators like Databricks' Photon, DataFusion Comet, and Gluten (Velox) speed up parts of Spark's execution with native code, but they keep the JVM: Spark still plans and orchestrates the query, and anything the accelerator does not cover falls back to JVM execution across a conversion boundary. That is why, per Databricks' own Photon documentation, Photon does not support Python UDFs, RDDs, or Dataset APIs, supports only stateless streaming, and does not speed up queries that already run in under two seconds. LakeSail is not an accelerator; it removes the JVM entirely, a Rust engine on Arrow and DataFusion with Arrow Flight shuffle and in-process Python UDFs, so your PySpark runs with nothing to configure and no operators silently falling back to the JVM. For the full comparison, see how Sail compares to Photon and other accelerators.

Is it really faster than Spark, and how was that measured?

On Sail's derived TPC-H benchmark, Sail runs the suite 10x faster than Spark on the same hardware with far less memory; the methodology and per-query numbers are on that page. It also leads the public ClickBench benchmark against Spark and the Spark accelerators. These are analytical benchmarks; the real test is your own workload, and the engine is free to run locally.

If DuckDB or Polars can handle my data, why do I need this?

If your data fits comfortably on one machine, they may be all you need. Sail runs single-node too. The difference is continuity: Sail gives you one Spark-compatible API from a laptop to a distributed, fault-tolerant cluster, plus native Delta and Iceberg and the Spark ecosystem, so you do not rewrite when a workload outgrows a single node.

Compatibility ​

Will my existing PySpark run on LakeSail?

LakeSail runs PySpark and Spark SQL as they are: a job is the file or SQL you already have, and a session is an endpoint your existing Spark Connect client attaches to. The Sail docs list what carries over and what does not.

Can I run Python and AI workloads?

Yes. Beyond Spark SQL and DataFrames, Python is a first-class workload: Python pipelines, UDFs, and AI and ML work like embedding generation and feature engineering all run on the same engine, in your own cluster against your own data. You don't stand up a separate service for the Python and AI side.

Do Python UDFs still pay PySpark's serialization cost?

On PySpark, Python UDFs cross the JVM-to-Python boundary and pay a serialization cost. Sail has no JVM, so Python runs in-process with the engine and shares Arrow data directly. See the Sail docs on UDFs.

Does it read my existing Delta and Iceberg tables in place?

Yes, once the table data is in the network's workspace bucket. Managed compute reaches S3 table data there, so you first make your existing S3 tables available in it. LakeSail then reads Delta, Iceberg, and Parquet in place, with no format conversion, and your tables stay in open formats, readable by any other engine too.

I'm on Databricks. What changes when I move?

Your PySpark, Spark SQL, and Delta tables carry over. The work is the Databricks-specific pieces: Workflows become LakeSail jobs with schedules and retries, and notebook magics and dbutils calls move to standard Spark or your own tooling.

Your account, data & access ​

What can LakeSail access in my account, and does my data leave?

LakeSail holds no long-lived credentials. It assumes one scoped IAM role on demand with a per-connection external ID, behind a permissions boundary that denies everything outside a fixed allow-list. Your data stays in your account; the control plane holds metadata such as job definitions, cluster identifiers, and run status.

Where does my data live?

Your data stays in your AWS account, in the region where you deploy the connection. LakeSail's control plane and its metadata run in us-east-1. See Security & IAM for the trust model and data handling.

Which catalogs and BI tools work?

LakeSail connects to AWS Glue, Hive Metastore, Unity Catalog, OneLake, and Iceberg REST catalogs. A workload can attach several at once, with one as the default that resolves unqualified table names. For BI, you run the transform in LakeSail and land the result as a table in your lake, then point Tableau, Power BI, or Looker at it through their usual connector; BI tools connect to SQL or JDBC endpoints, not to a Spark Connect session. Programmatic clients like PySpark and notebooks connect directly over Spark Connect.

Am I locked in?

No. The engine is open source, your data stays in open formats (Delta, Iceberg, Parquet) in your own S3, and the connection protocol is Spark Connect, which is open. If you ever leave, there is nothing to export or convert; other engines read the same files directly, and the workspace bucket is retained if you disconnect.

Running it in production ​

How do I schedule and monitor production jobs?

Jobs are first-class: you define a job once, version it, give it a cron schedule with retry and timeout policy, and each run keeps a full history. The run view shows the run's lifecycle, a status message, execution time, the exact job version that ran, and a link to browse the output in your S3. For alerting, route run events like failed or timeout to email, Slack, a webhook, PagerDuty, or Rootly.

What happens when a job fails? Do I get retries?

Every run keeps a full record: its status, timing, the exact job version that ran, and the engine's error output, so a Failed run usually points straight at the cause, whether a SQL error, a missing column, or a Python exception. Retries are built in: set a max retry count on the job, or leave it to inherit the compute profile's default (0 disables retries, -1 retries indefinitely). Give a job a per-run timeout and a stuck run is stopped rather than running indefinitely. From the run view you can retry a failed run by hand; the retry re-runs it in place, keeping its original ID and history. Route failed and timeout events to email, Slack, PagerDuty, Rootly, or a webhook so you hear about them right away. See Runs & Debugging and Notifications.


For a question that isn't covered here, email support@lakesail.com.

Can't find the answer here? Email us: support@lakesail.com