Why LakeSail?
Self-managed Spark is operationally expensive, and Spark itself was built for an earlier era of data workloads. The JVM runtime and the assumptions baked into the engine predate today's mix of interactive queries, Python-heavy pipelines, and AI workloads moving Arrow buffers and tensors. LakeSail's engine is written from the ground up in Rust for today's data and AI workloads, not retrofitted onto a JVM runtime.
Powered by a Rust-native engine
LakeSail's engine, Sail, is a Rust-native runtime that runs batch, SQL, Python, and AI on one foundation, with no JVM, no garbage collector, and compatibility with the Spark API your team already writes.
| JVM-based Spark | Sail | |
|---|---|---|
| Runtime | JVM, stop-the-world GC | Rust, no garbage collector |
| Idle memory | Gigabytes of heap | Tens of MB |
| Cold start | Slow to warm up | Seconds |
| Python | Py4J, serialized on every hop | Zero-copy Arrow, in-process |
| Engine | JVM executors, even with accelerators | Native engine, JVM removed |
On Sail's derived TPC-H benchmark, Sail runs the 22-query suite 10x faster than JVM-based Spark on the same hardware, peaking at 26 GB of memory where Spark peaks at 72 GB, with zero disk spill. It also leads the public ClickBench benchmark against Spark and every Spark accelerator.
The choice to build in Rust rather than extend the JVM has practical advantages throughout the stack:
- No JVM to tune. Rust's ownership model gives memory safety without a garbage collector, so there are no stop-the-world GC pauses and no heap or GC tuning. Latency stays predictable at scale, which matters most for interactive queries and agent loops.
- A native engine. Photon and Comet add native operators to Spark, but Spark still plans every query on the JVM, and anything they do not support falls back to the Spark engine. Sail removes the JVM entirely. See Sail and Spark accelerators.
- Less memory, more compute. Rust carries no JVM object headers, no boxed-type overhead, and no reserved GC heap, and it frees memory deterministically, so the engine holds RAM only while a query runs and releases it right after, instead of a JVM sitting on a large heap the whole session. The same workload fits in far less RAM, leaving more of each node for computation.
Because Sail is open source, the engine at the core of your platform is not a black box: you can read the source, run it locally, and audit exactly how your data is processed.
Spark-compatible
Sail is designed as a drop-in replacement for the Spark compute engine. It speaks the Spark API (SQL, DataFrames, and Spark Connect) from one binary, so the interface your analysts and engineers already target continues to work. What changes is where the code runs, not what you wrote.
- Your PySpark and Spark SQL run on Sail. Sail maintains the Spark API contract, so evaluating a different engine does not mean rewriting jobs.
- Jobs take your files as they are. A job is the
.pyfile, wheel, or SQL you already have, with a schedule and a compute profile around it. - Spark Connect for your own clients. A session is a Spark Connect endpoint. PySpark, Scala, Go, and Rust clients, local notebooks, and IDEs connect to it without changes.
- Standard Spark SQL everywhere. SQL jobs, saved queries, notebooks, and
spark.sql()from a client all run the same Spark SQL semantics on the same engine.
Python-native, built for AI and agentic workloads
PySpark on the JVM has to serialize and deserialize every batch crossing the Python-to-JVM boundary; Sail passes Arrow array pointers directly, so a Python UDF reads the same buffer the engine just wrote.
- Native Python, in-process. Pandas UDFs and embedding generation run inline with SQL and DataFrame work, sharing the engine's Arrow buffers directly with no JVM serialization tax.
- A better substrate for AI-era workloads. No GC pauses, predictable latency, and zero-copy interop matter more for tool-using agents and tight LLM-in-the-loop pipelines than they do for nightly batch. The unpredictability of stop-the-world GC compounds where every step is interactive.
- Agents are first-class clients. A session hands out a scoped, time-boxed Spark Connect token, so an LLM agent can query tables, run Python UDFs, and iterate in a loop on the same engine your analysts use, no human in the loop. The predictable, GC-free latency that makes interactive work fast is exactly what tight agent loops need.
Native lakehouse formats
Sail implements Iceberg and Delta Lake directly in Rust, inside its own query planner and runtime.
- Delta Lake and Iceberg, read and write. Native implementations of both formats, inside the engine's own planner and runtime.
- The catalog you already use. Sail connects to the common catalog providers. On the platform, see Connect a Catalog.
- Direct object storage access. Sail reads from S3 directly. There's no intermediate staging layer; the control plane handles schema and run metadata, not row data.
- Standard formats in, standard formats out. Because Sail writes Iceberg and Delta, any other tool that reads those formats (a BI tool, a different engine, a notebook) can read the same tables without a migration step.
Your data, your account, open standards
Compute runs in your own AWS account, and data stays in your object storage in open formats. The rest of the architecture follows from that.
- No data egress. Query execution happens on clusters in your VPC. The control plane handles metadata (table schemas, job definitions, run statuses), not row data.
- Open table formats throughout. Iceberg, Delta, and Parquet on S3. LakeSail registers a pointer to your existing tables; the data doesn't move.
- Open wire protocol. The connection protocol is Spark Connect. Any compliant client can use it, without a vendor-specific SDK or driver.
- Reversible by default. If you stop using LakeSail, the tables are already in the format your next tool expects. There's nothing to extract or convert.
This matters for every team, and especially those with data residency, isolation, or audit obligations.
Operationally lighter, and cheaper
Running Sail yourself is possible. Running it well, with autoscaling, job scheduling, retries, observability, secrets management, and access control wired together, is heavy infrastructure work that needs constant attention. That takes a dedicated platform layer around it.
LakeSail is that layer:
- Compute follows the work. Capacity comes up when a job, session, or notebook starts and is released when it stops. Sail starts in seconds, so there is no reason to keep a cluster warm between workloads.
- Jobs are first-class. Versioned definitions, configurable retries, scheduling, run history, and notifications are built in.
- Access control is built in. Role-based access, team management, SSO, and MFA are part of LakeSail.
What it costs
You pay your cloud provider for the compute, and LakeSail a flat hourly rate per vCPU and per GB of memory that your workloads use. There is no proprietary compute unit, no per-query markup, and no license fee for the engine, because Sail is open source. Usage and rates are visible in the console at any time.
The same workload runs on smaller, fewer, or shorter-lived nodes than the JVM-based equivalent, and with no JVM heap to reserve it fits in less RAM to begin with. The pricing page compares LakeSail with EMR and Databricks on common instance types.
Next steps
- Concepts: how the control plane, engine, and your data fit together.
- Quickstart: try it.