aws deploy · ~15 minutes · open tofu

Boot the real Flashpoint stack on AWS

One OpenTofu apply creates the full production shape: a t4g.small EC2 gateway running the FastAPI control plane, Spark Connect drivers on Fargate with spot executors, DynamoDB state, and a Cost Center fed by real Cost Explorer data. The booted gateway serves main, so this deploys the whole Cost Center feature.

~15 minapply → healthy gateway
1 × t4g.smallalways-on gateway EC2
~$14/mobaseline while idle
$0.08/hXS warehouse while running
step 0

Prerequisites

# check what you have
tofu --version
aws --version
aws sts get-caller-identity # confirms the profile resolves
Stuck?

aws sts get-caller-identity fails → you need credentials. Flashpoint was built with an IAM user profile named personal-aws-iam; use yours and export it for every command below, or set it as the default.

Open DeepSeek chat
I'm deploying the flashpoint repo's infra/ with OpenTofu. `aws sts get-caller-identity` fails with "Unable to locate credentials" on macOS. My AWS CLI profile is named [profile-name]. What exactly should I set in ~/.aws/config so `tofu apply -var-file=dev.tfvars` picks it up, and how do I verify?
step 1

One apply

State is local (infra/terraform.tfstate, gitignored) — no remote backend needed. On a fresh clone the first apply builds the whole stack; with existing state it only reconciles drift (e.g. recreating the gateway EC2 if it was terminated).

cd infra
tofu init
tofu apply -var-file=dev.tfvars -auto-approve

Expected: one aws_instance.gateway created (plus instance profile bits). If apply shows a long plan touching VPC/ECS/DynamoDB, stop — something is out of sync; the apply should be surgical.

What pins the gateway code? dev.tfvars has a gateway_branch setting. On first boot the EC2 instance clones that branch from GitHub and runs it — it is not deployed from your laptop. Point it at main unless you're deliberately testing a branch.
Stuck?

Plan shows diffs on resources you didn't touch (VPC, ECS, DynamoDB). The state file is stale vs. what's live — don't apply blindly, or you'll rebuild the world.

Open DeepSeek chat
I ran `tofu apply -var-file=dev.tfvars` in the flashpoint repo's infra/ and the plan wants to replace/modify aws_vpc, aws_ecs_cluster and DynamoDB tables I didn't change. The repo has committed terraform.tfstate. The apply previously only created aws_instance.gateway. What should I check (tofu state list, remote vs local state, tfvars drift)?
step 2

Wait for first boot

On first launch the instance runs gateway-init.sh from user-data: installs Python 3.12 + uv, clones the repo, runs uv sync (pulls pyspark — this is the slow part), writes /etc/flashpoint-gateway.env, and starts a systemd service flashpoint-gateway running the FastAPI app on :8080. Allow 3–6 minutes.

tofu output gateway_api     # http://<public-ip>:8080

# poll until it answers:
curl -s http://$(tofu output -raw gateway_public_ip):8080/healthz

Healthy response looks like:

{"status":"ok"}
Probing from inside — the instance carries the SSM agent, so you can inspect the boot directly without SSH:
aws ssm send-command --instance-ids "$(aws ssm describe-instance-information --query 'InstanceInformationList[0].InstanceId' --output text)" \
  --document-name AWS-RunShellScript \
  --parameters 'commands=["journalctl -u flashpoint-gateway -n 40 --no-pager"]' \
  --query 'Command.CommandId' --output text
Stuck?

/healthz doesn't answer after 6 minutes. Most likely the user-data script failed (bad branch name is the classic — see step 1) or uv sync is still downloading. Check journalctl -u flashpoint-gateway and the cloud-init log /var/log/cloud-init-output.log via SSM.

Open DeepSeek chat
Flashpoint deploy on AWS: I applied infra/ with OpenTofu, the t4g.small gateway instance is running, but curl http://[public-ip]:8080/healthz times out after 6+ minutes. I ran `aws ssm send-command` with `journalctl -u flashpoint-gateway -n 40` and it shows: [paste output]. What failed in gateway-init.sh / systemd?
step 3

Open the UI

The web app runs on your laptop; point it at the live gateway with one env var. The gateway serves CORS-open, so any origin works.

cd web
npm install
VITE_GATEWAY_URL=http://$(cd ../infra && tofu output -raw gateway_public_ip):8080 npm run dev

Open http://localhost:5173. The top bar should show the gateway as connected. The worksheet can run SQL against real warehouses now; Cost Center will show the real gateway EC2 row and (after warehouses run) live Fargate drivers/executors.

Stuck?

The UI shows an offline/connection banner while curl to IP:8080/healthz works. The VITE_GATEWAY_URL env var must be set at npm run dev start time — restarting the dev server with the var is the fix.

Open DeepSeek chat
Flashpoint UI on localhost:5173 shows the gateway offline banner, but `curl http://[gateway-ip]:8080/healthz` returns {"status":"ok"}. I ran `VITE_GATEWAY_URL=http://[gateway-ip]:8080 npm run dev`. What else could make the web app unable to reach the gateway? (CORS? fetch errors in devtools? the api.js BASE constant?)
step 4

First warehouse + query

Create an XS warehouse — this launches a Fargate driver task (on-demand) plus one executor (spot). The driver image is already in ECR; pulling + Spark JVM boot takes 2–3 minutes.

GATEWAY=http://$(cd ../infra && tofu output -raw gateway_public_ip):8080

curl -s -X POST $GATEWAY/warehouses -H 'Content-Type: application/json' \
  -d '{"name":"demo","size":"XS"}'

# watch the Fargate tasks come up:
aws ecs list-tasks --cluster flashpoint-dev-cluster --desired-status RUNNING
aws ecs describe-tasks --cluster flashpoint-dev-cluster --tasks $(aws ecs list-tasks --cluster flashpoint-dev-cluster --query 'taskArns[0]' --output text) \
  --query 'tasks[].{role:containers[0].name,lastStatus:lastStatus,capacity:capacityProviderName}' --output table

Once the warehouse state is RUNNING, run a query that actually exercises the cluster — range() needs no tables, and the count is computed across the driver and the executor:

curl -s -X POST $GATEWAY/warehouses/demo/query -H 'Content-Type: application/json' \
  -d '{"sql":"SELECT count(*) AS rows FROM range(100000000)"}' | python3 -m json.tool

Expect a single row, 100,000,000, with duration_ms in the response. Open the query id in History for the full operator tree profile.

Stuck?

The warehouse hangs in PROVISIONING. Check the driver task's containers — the classic failure is the task role missing a permission (read the gateway log for the exact error) or the ECR image tag referenced by ecs.tf no longer existing in the repo.

Open DeepSeek chat
Flashpoint on AWS: I POSTed a warehouse to the gateway, it stays PROVISIONING, and the Fargate driver task in `aws ecs describe-tasks` shows containers with `lastStatus: STOPPED` / exit code [X]. The gateway log shows: [paste the error]. What is wrong — task role IAM, ECR image tag in infra/ecs.tf, or something in the driver container entrypoint?
step 5

Cost Center on real data

With the gateway up and a warehouse run behind you, Cost Center is live:

Stuck?

Cost Center shows resources but the chart is all zeros. That's the cost allocation tag: check its status in Billing → Cost allocation tags, or with aws ce list-cost-allocation-tags. The tag change propagates within ~24h and does not backfill.

Open DeepSeek chat
Flashpoint Cost Center on AWS shows the EC2/Fargate resources but the daily spend chart is all $0.00. `aws ce get-cost-and-usage` with a Filter on Tags Project=flashpoint returns zero amounts, but the same query without the filter shows real numbers. I checked `aws ce list-cost-allocation-tags` — what does an Inactive status mean and how do I activate the Project tag?
step 6

What it costs & full teardown

resourcedetailcost
gateway EC2t4g.small, always on~$12.26/mo
gateway EBS20 GB gp3 root volume~$1.60/mo
warehousesXS = driver + 1 spot executor~$0.08/h while running
infra penniesECR, S3, DynamoDB, log groups~$0.10/mo

Warehouses auto-suspend after their TTL (2h default) — usage meters keep accruing per session so Cost Center reconciles what actually ran. Pausing a demo without losing the stack is one command: aws ec2 stop-instances (keeps the EBS volume and boots the already-installed gateway next time; the ~$1.60/mo volume charge stays).

Full teardown → Flashpoint cost goes to $0

tofu destroy removes the entire Flashpoint stack: gateway EC2 + its root EBS, ECS cluster and task definitions, DynamoDB tables, S3 bucket, ECR repo, log groups, and the VPC. Run it from infra/:

cd infra
tofu destroy -var-file=dev.tfvars -auto-approve

Then verify nothing is left running or billed:

# these should all come back empty:
aws ec2 describe-instances --query 'Reservations[].Instances[].InstanceId'
aws ec2 describe-volumes      --query 'Volumes[].VolumeId'
aws ecs list-tasks --cluster flashpoint-dev-cluster
aws dynamodb list-tables      --query 'TableNames[?starts_with(@,`flashpoint`)]'
aws s3 ls
Stuck?

tofu destroy fails halfway (a resource already deleted manually, a VPC dependency, an S3 bucket that isn't empty). The state may also drift if anything was changed by hand since apply.

Open DeepSeek chat
Flashpoint teardown: `tofu destroy -var-file=dev.tfvars` in the repo's infra/ fails with [paste the error]. I previously applied the stack and then manually deleted/stopped some resources in the AWS console. How do I reconcile the state file (tofu state rm?) so destroy completes and Flashpoint stops billing?
next

What now?