Boot a local Spark Connect server, seed demo data (1M customers × 10M orders), start the gateway, and run the join/group-by query — ending with a real query profile of 21 nodes, the same scenario that powers the blog-post screenshots. All AWS calls are mocked; nothing leaves your machine.
gituv — install uv# check what you have
git --version
python3 --version
uv --version
java -version
JDK problems are the #1 failure here — Homebrew installs Java 26 by default, and Spark
dies on it with a jdk.internal.ref.Cleaner error.
macOS: I installed JDK via `brew install openjdk` and it's version 26. I need JDK 17 for Apache Spark 4.2. What's the exact Homebrew command to install openjdk@17, and how do I point Spark at it (JAVA_HOME + PATH)?
git clone https://github.com/prabodh1194/flashpoint.git cd flashpoint
If you can't clone, check SSH keys or use the web UI's Download ZIP instead.
`git clone https://github.com/prabodh1194/flashpoint.git` fails with a permission/network error on macOS. The repo is public. What should I check (SSH keys, proxy, git config)?
A single script does everything — boots Spark Connect, seeds data, starts the gateway, creates a warehouse and runs the query. Run it from the repo root:
python3 scripts/e2e_demo.py
Expected end of output (durations will vary):
── step 5/5 ── running the join/group-by query query id: faae3bb101fa8e1e columns: ['region', 'cnt'] rows: 5 — [['central', '2000000'], ['west', '2000000'], ['north', '2000000']] duration: 2211 ms (api round-trip 2253 ms) profile: 21 nodes, 17 with column treatments plan tree root: AdaptiveSparkPlan done.
Flags: --keep leaves the servers running (for the UI step below),
--reseed regenerates the demo data, --skip-seed assumes data exists.
Press Ctrl-C to tear everything down.
Everything goes to /tmp/flashpoint-demo/*.log — check those first. Common
failures: a stale server already on port 15002/8080 (kill it with
lsof -nP -iTCP:15002 -sTCP:LISTEN), or the JDK issue from step 0.
I ran `python3 scripts/e2e_demo.py` from the flashpoint repo root. It fails at booting the Spark Connect server: [paste the last 15 lines of /tmp/flashpoint-demo/spark-connect.log here]. I'm on macOS with Homebrew JDK [your java -version]. What's wrong?
Useful when you want each component in its own terminal. First, install gateway deps:
uv sync --directory gateway
Terminal A — boot the Spark Connect server (JDK 17/21 required):
export JAVA_HOME=$(/usr/libexec/java_home -v 21) # or 17
cd gateway/.venv/lib/python3.*/site-packages/pyspark
bin/spark-class org.apache.spark.deploy.SparkSubmit \
--master local[*] \
--conf spark.ui.port=4040 \
--class org.apache.spark.sql.connect.service.SparkConnectServer \
spark-internal
Wait for Spark session available at sc://localhost:15002 in the log. Terminal B —
seed the demo data (this uses the server, then exits):
cd flashpoint
gateway/.venv/bin/python -c "
from pyspark.sql import SparkSession
s = SparkSession.builder.remote('sc://localhost:15002').getOrCreate()
s.range(1000000).selectExpr(
'CAST(id AS INT) AS customer_id', \"CONCAT('user_', id) AS name\",
'CASE WHEN id % 5 = 0 THEN \'north\' WHEN id % 5 = 1 THEN \'south\' WHEN id % 5 = 2 THEN \'east\' '
'WHEN id % 5 = 3 THEN \'west\' ELSE \'central\' END AS region', 'CAST(id % 3 AS INT) AS tier',
).write.mode('overwrite').parquet('/tmp/spark-data/customers')
s.range(10000000).selectExpr(
'id', 'CAST(id % 1000000 AS INT) AS customer_id', 'CAST(id % 5000 AS INT) AS product_id',
'CAST(id AS DECIMAL(10,2)) AS amount',
\"CAST(date_add(CAST('2024-01-01' AS DATE), CAST(id % 365 AS INT)) AS STRING) AS order_date\",
).write.mode('overwrite').parquet('/tmp/spark-data/orders')
print('seeded')"
Terminal C — boot the gateway (AWS mocked, in-memory DynamoDB):
cd gateway .venv/bin/python local_dev.py
Terminal D — create a warehouse, register the parquet data as views through the gateway, and run the query:
curl -s -X POST localhost:8080/warehouses -H 'Content-Type: application/json' \
-d '{"name":"demo","size":"S"}'
curl -s -X POST localhost:8080/warehouses/demo/query -H 'Content-Type: application/json' \
-d '{"sql":"CREATE OR REPLACE TEMPORARY VIEW customers USING parquet OPTIONS (path '\''/tmp/spark-data/customers'\'')"}'
curl -s -X POST localhost:8080/warehouses/demo/query -H 'Content-Type: application/json' \
-d '{"sql":"CREATE OR REPLACE TEMPORARY VIEW orders USING parquet OPTIONS (path '\''/tmp/spark-data/orders'\'')"}'
curl -s -X POST localhost:8080/warehouses/demo/query -H 'Content-Type: application/json' \
-d '{"sql":"SELECT c.region, count(*) AS cnt FROM orders o JOIN customers c ON o.customer_id = c.customer_id GROUP BY c.region ORDER BY cnt DESC"}' \
| python3 -m json.tool
Include the full error from the gateway log (gateway.log when run via the
script) or the JSON detail field from the curl response.
Flashpoint manual quickstart: `curl -X POST localhost:8080/warehouses/demo/query` returns 400 with: [paste the JSON `detail` error]. I followed the quickstart steps, Spark Connect is on :15002, gateway on :8080. What does this error mean and how do I fix it?
With the gateway running (from the script's --keep or terminal C), start the web UI:
cd web npm install npm run dev
Open http://localhost:5173. Run a query in the worksheet (or re-run it from
History), then click the query id to open its profile — the result-at-top operator tree with
codegen chips. Every profile has a URL like #/history/<query_id> that
survives reload.
If the UI can't reach the gateway, check the API base URL in
web/src/ — it should point at http://localhost:8080.
Flashpoint UI: `npm run dev` works and the page loads at localhost:5173, but the worksheet shows a connection error / offline banner while the gateway is definitely on localhost:8080 (curl /healthz works). What should I check in web/src for the API base URL?
# if you used the script: Ctrl-C in its terminal is enough
# by hand: stop each process, then free the ports
lsof -nP -iTCP:15002 -sTCP:LISTEN | awk 'NR==2{print $2}' | xargs kill
lsof -nP -iTCP:8080 -sTCP:LISTEN | awk 'NR==2{print $2}' | xargs kill
Leftover processes make the next run "reuse" a stale server — when in doubt, kill ports 15002 and 8080, then re-run the script fresh.
Open DeepSeek chatOn macOS, my next run of the flashpoint demo script says "port 15002 already in use — reusing" but queries fail with TABLE_OR_VIEW_NOT_FOUND. How do I reliably kill the previous Spark Connect server and all its children on macOS?
http://localhost:8080/docs — warehouse CRUD, sync + async queries.http://localhost:4040 while a query runs.gateway/dag.py (Spark UI REST → {nodes, edges}).docs/ — architecture ADRs, warehouse sizing, fargate notes.