local e2e · ~10 minutes · laptop only

Quickstart: run the whole demo on your laptop

Boot a local Spark Connect server, seed demo data (1M customers × 10M orders), start the gateway, and run the join/group-by query — ending with a real query profile of 21 nodes, the same scenario that powers the blog-post screenshots. All AWS calls are mocked; nothing leaves your machine.

21nodes in the query profile
10Morders + 1M customers seeded
2–3sjoin query on the demo cluster
0 AWSeverything mocked locally
step 0

Prerequisites

# check what you have
git --version
python3 --version
uv --version
java -version
Stuck?

JDK problems are the #1 failure here — Homebrew installs Java 26 by default, and Spark dies on it with a jdk.internal.ref.Cleaner error.

Open DeepSeek chat
macOS: I installed JDK via `brew install openjdk` and it's version 26. I need JDK 17 for Apache Spark 4.2. What's the exact Homebrew command to install openjdk@17, and how do I point Spark at it (JAVA_HOME + PATH)?
step 1

Clone the repo

git clone https://github.com/prabodh1194/flashpoint.git
cd flashpoint
Stuck?

If you can't clone, check SSH keys or use the web UI's Download ZIP instead.

Open DeepSeek chat
`git clone https://github.com/prabodh1194/flashpoint.git` fails with a permission/network error on macOS. The repo is public. What should I check (SSH keys, proxy, git config)?
step 2

Recommended: one-shot script

A single script does everything — boots Spark Connect, seeds data, starts the gateway, creates a warehouse and runs the query. Run it from the repo root:

python3 scripts/e2e_demo.py

Expected end of output (durations will vary):

── step 5/5 ── running the join/group-by query
query id:   faae3bb101fa8e1e
columns:    ['region', 'cnt']
rows:       5 — [['central', '2000000'], ['west', '2000000'], ['north', '2000000']]
duration:   2211 ms (api round-trip 2253 ms)
profile:    21 nodes, 17 with column treatments
plan tree root: AdaptiveSparkPlan

done.

Flags: --keep leaves the servers running (for the UI step below), --reseed regenerates the demo data, --skip-seed assumes data exists. Press Ctrl-C to tear everything down.

Stuck?

Everything goes to /tmp/flashpoint-demo/*.log — check those first. Common failures: a stale server already on port 15002/8080 (kill it with lsof -nP -iTCP:15002 -sTCP:LISTEN), or the JDK issue from step 0.

Open DeepSeek chat
I ran `python3 scripts/e2e_demo.py` from the flashpoint repo root. It fails at booting the Spark Connect server: [paste the last 15 lines of /tmp/flashpoint-demo/spark-connect.log here]. I'm on macOS with Homebrew JDK [your java -version]. What's wrong?
step 3

Manual path (same steps, by hand)

Useful when you want each component in its own terminal. First, install gateway deps:

uv sync --directory gateway

Terminal A — boot the Spark Connect server (JDK 17/21 required):

export JAVA_HOME=$(/usr/libexec/java_home -v 21)   # or 17
cd gateway/.venv/lib/python3.*/site-packages/pyspark
bin/spark-class org.apache.spark.deploy.SparkSubmit \
  --master local[*] \
  --conf spark.ui.port=4040 \
  --class org.apache.spark.sql.connect.service.SparkConnectServer \
  spark-internal

Wait for Spark session available at sc://localhost:15002 in the log. Terminal B — seed the demo data (this uses the server, then exits):

cd flashpoint
gateway/.venv/bin/python -c "
from pyspark.sql import SparkSession
s = SparkSession.builder.remote('sc://localhost:15002').getOrCreate()
s.range(1000000).selectExpr(
    'CAST(id AS INT) AS customer_id', \"CONCAT('user_', id) AS name\",
    'CASE WHEN id % 5 = 0 THEN \'north\' WHEN id % 5 = 1 THEN \'south\' WHEN id % 5 = 2 THEN \'east\' '
    'WHEN id % 5 = 3 THEN \'west\' ELSE \'central\' END AS region', 'CAST(id % 3 AS INT) AS tier',
).write.mode('overwrite').parquet('/tmp/spark-data/customers')
s.range(10000000).selectExpr(
    'id', 'CAST(id % 1000000 AS INT) AS customer_id', 'CAST(id % 5000 AS INT) AS product_id',
    'CAST(id AS DECIMAL(10,2)) AS amount',
    \"CAST(date_add(CAST('2024-01-01' AS DATE), CAST(id % 365 AS INT)) AS STRING) AS order_date\",
).write.mode('overwrite').parquet('/tmp/spark-data/orders')
print('seeded')"

Terminal C — boot the gateway (AWS mocked, in-memory DynamoDB):

cd gateway
.venv/bin/python local_dev.py
Why register views through the gateway? Spark Connect sessions are isolated: a view registered from your seeding client is invisible to the gateway's session. The gateway must register the views itself — that's the next step.

Terminal D — create a warehouse, register the parquet data as views through the gateway, and run the query:

curl -s -X POST localhost:8080/warehouses -H 'Content-Type: application/json' \
  -d '{"name":"demo","size":"S"}'

curl -s -X POST localhost:8080/warehouses/demo/query -H 'Content-Type: application/json' \
  -d '{"sql":"CREATE OR REPLACE TEMPORARY VIEW customers USING parquet OPTIONS (path '\''/tmp/spark-data/customers'\'')"}'
curl -s -X POST localhost:8080/warehouses/demo/query -H 'Content-Type: application/json' \
  -d '{"sql":"CREATE OR REPLACE TEMPORARY VIEW orders USING parquet OPTIONS (path '\''/tmp/spark-data/orders'\'')"}'

curl -s -X POST localhost:8080/warehouses/demo/query -H 'Content-Type: application/json' \
  -d '{"sql":"SELECT c.region, count(*) AS cnt FROM orders o JOIN customers c ON o.customer_id = c.customer_id GROUP BY c.region ORDER BY cnt DESC"}' \
  | python3 -m json.tool
Stuck?

Include the full error from the gateway log (gateway.log when run via the script) or the JSON detail field from the curl response.

Open DeepSeek chat
Flashpoint manual quickstart: `curl -X POST localhost:8080/warehouses/demo/query` returns 400 with: [paste the JSON `detail` error]. I followed the quickstart steps, Spark Connect is on :15002, gateway on :8080. What does this error mean and how do I fix it?
step 4

Open the UI and view the profile

With the gateway running (from the script's --keep or terminal C), start the web UI:

cd web
npm install
npm run dev

Open http://localhost:5173. Run a query in the worksheet (or re-run it from History), then click the query id to open its profile — the result-at-top operator tree with codegen chips. Every profile has a URL like #/history/<query_id> that survives reload.

Stuck?

If the UI can't reach the gateway, check the API base URL in web/src/ — it should point at http://localhost:8080.

Open DeepSeek chat
Flashpoint UI: `npm run dev` works and the page loads at localhost:5173, but the worksheet shows a connection error / offline banner while the gateway is definitely on localhost:8080 (curl /healthz works). What should I check in web/src for the API base URL?
step 5

Teardown

# if you used the script: Ctrl-C in its terminal is enough
# by hand: stop each process, then free the ports
lsof -nP -iTCP:15002 -sTCP:LISTEN | awk 'NR==2{print $2}' | xargs kill
lsof -nP -iTCP:8080  -sTCP:LISTEN | awk 'NR==2{print $2}' | xargs kill
Stuck?

Leftover processes make the next run "reuse" a stale server — when in doubt, kill ports 15002 and 8080, then re-run the script fresh.

Open DeepSeek chat
On macOS, my next run of the flashpoint demo script says "port 15002 already in use — reusing" but queries fail with TABLE_OR_VIEW_NOT_FOUND. How do I reliably kill the previous Spark Connect server and all its children on macOS?
next

What now?