Skip to content

Check Your Spark Code ​

Whether your code runs on LakeSail is a question about your code, not a general claim. Three steps answer it: scan the code, port one real job, and look up anything the scan flags in the Sail docs.

1. Scan your code ​

Point the Spark-to-Sail scanner at your code before you change anything. Paste a snippet, upload a file, or run it over a whole repository. It reports which of the Spark features your code uses are supported by Sail.

A scan tells you which calls are supported. It does not tell you that the output matches on your data. That is step 2.

2. Port one job and diff ​

  1. Pick a job that uses the patterns you care about: your common joins, your UDFs, your write path.
  2. Run it on LakeSail as a job with the file or SQL as it is, or from your own client through a session.
  3. Run the same job where it runs today.
  4. Compare the output: row counts, schema, and a content hash.

Do this on at least one real workload before moving anything else. A clean diff on a representative job covers the functions, data, and edge cases your code hits.

3. Look up what the scan flags ​

Most flagged items have a short path: a Scala or Java UDF becomes a Python UDF, a Hive-specific function gets swapped. Code built on the RDD API does not run on Sail and needs a rewrite on DataFrames or SQL.

Next ​

  • Quickstart: set up the platform and run your first job.
  • Introduction: what LakeSail is and what you can run on it.

Can't find the answer here? Email us: support@lakesail.com