One OpenTofu apply creates the full production shape: a t4g.small EC2 gateway running the
FastAPI control plane, Spark Connect drivers on Fargate with spot executors, DynamoDB state,
and a Cost Center fed by real Cost Explorer data. The booted gateway serves
main, so this deploys the whole Cost Center feature.
tofu (OpenTofu 1.x) — installaws CLI with credentials for your AWS account (any profile works — the guide uses personal-aws-iam)us-east-1 (the infra is region-parameterized but dev is deployed there)# check what you have tofu --version aws --version aws sts get-caller-identity # confirms the profile resolves
aws sts get-caller-identity fails → you need credentials. Flashpoint was built
with an IAM user profile named personal-aws-iam; use yours and export it for
every command below, or set it as the default.
I'm deploying the flashpoint repo's infra/ with OpenTofu. `aws sts get-caller-identity` fails with "Unable to locate credentials" on macOS. My AWS CLI profile is named [profile-name]. What exactly should I set in ~/.aws/config so `tofu apply -var-file=dev.tfvars` picks it up, and how do I verify?
State is local (infra/terraform.tfstate, gitignored) — no remote backend
needed. On a fresh clone the first apply builds the whole stack; with existing state it
only reconciles drift (e.g. recreating the gateway EC2 if it was terminated).
cd infra tofu init tofu apply -var-file=dev.tfvars -auto-approve
Expected: one aws_instance.gateway created (plus instance profile bits). If
apply shows a long plan touching VPC/ECS/DynamoDB, stop — something is out of sync; the
apply should be surgical.
dev.tfvars has a
gateway_branch setting. On first boot the EC2 instance clones that branch from
GitHub and runs it — it is not deployed from your laptop. Point it at main
unless you're deliberately testing a branch.Plan shows diffs on resources you didn't touch (VPC, ECS, DynamoDB). The state file is stale vs. what's live — don't apply blindly, or you'll rebuild the world.
Open DeepSeek chatI ran `tofu apply -var-file=dev.tfvars` in the flashpoint repo's infra/ and the plan wants to replace/modify aws_vpc, aws_ecs_cluster and DynamoDB tables I didn't change. The repo has committed terraform.tfstate. The apply previously only created aws_instance.gateway. What should I check (tofu state list, remote vs local state, tfvars drift)?
On first launch the instance runs gateway-init.sh from user-data: installs
Python 3.12 + uv, clones the repo, runs uv sync (pulls pyspark — this is the
slow part), writes /etc/flashpoint-gateway.env, and starts a systemd service
flashpoint-gateway running the FastAPI app on :8080. Allow 3–6
minutes.
tofu output gateway_api # http://<public-ip>:8080 # poll until it answers: curl -s http://$(tofu output -raw gateway_public_ip):8080/healthz
Healthy response looks like:
{"status":"ok"}
aws ssm send-command --instance-ids "$(aws ssm describe-instance-information --query 'InstanceInformationList[0].InstanceId' --output text)" \ --document-name AWS-RunShellScript \ --parameters 'commands=["journalctl -u flashpoint-gateway -n 40 --no-pager"]' \ --query 'Command.CommandId' --output text
/healthz doesn't answer after 6 minutes. Most likely the user-data script
failed (bad branch name is the classic — see step 1) or uv sync is still
downloading. Check journalctl -u flashpoint-gateway and the cloud-init log
/var/log/cloud-init-output.log via SSM.
Flashpoint deploy on AWS: I applied infra/ with OpenTofu, the t4g.small gateway instance is running, but curl http://[public-ip]:8080/healthz times out after 6+ minutes. I ran `aws ssm send-command` with `journalctl -u flashpoint-gateway -n 40` and it shows: [paste output]. What failed in gateway-init.sh / systemd?
The web app runs on your laptop; point it at the live gateway with one env var. The gateway serves CORS-open, so any origin works.
cd web npm install VITE_GATEWAY_URL=http://$(cd ../infra && tofu output -raw gateway_public_ip):8080 npm run dev
Open http://localhost:5173. The top bar should show the gateway as connected.
The worksheet can run SQL against real warehouses now; Cost Center will show the
real gateway EC2 row and (after warehouses run) live Fargate drivers/executors.
The UI shows an offline/connection banner while curl to
IP:8080/healthz works. The VITE_GATEWAY_URL env var must be set
at npm run dev start time — restarting the dev server with the var is the fix.
Flashpoint UI on localhost:5173 shows the gateway offline banner, but `curl http://[gateway-ip]:8080/healthz` returns {"status":"ok"}. I ran `VITE_GATEWAY_URL=http://[gateway-ip]:8080 npm run dev`. What else could make the web app unable to reach the gateway? (CORS? fetch errors in devtools? the api.js BASE constant?)
Create an XS warehouse — this launches a Fargate driver task (on-demand) plus one executor (spot). The driver image is already in ECR; pulling + Spark JVM boot takes 2–3 minutes.
GATEWAY=http://$(cd ../infra && tofu output -raw gateway_public_ip):8080
curl -s -X POST $GATEWAY/warehouses -H 'Content-Type: application/json' \
-d '{"name":"demo","size":"XS"}'
# watch the Fargate tasks come up:
aws ecs list-tasks --cluster flashpoint-dev-cluster --desired-status RUNNING
aws ecs describe-tasks --cluster flashpoint-dev-cluster --tasks $(aws ecs list-tasks --cluster flashpoint-dev-cluster --query 'taskArns[0]' --output text) \
--query 'tasks[].{role:containers[0].name,lastStatus:lastStatus,capacity:capacityProviderName}' --output table
Once the warehouse state is RUNNING, run a query that actually exercises the
cluster — range() needs no tables, and the count is computed across the driver
and the executor:
curl -s -X POST $GATEWAY/warehouses/demo/query -H 'Content-Type: application/json' \
-d '{"sql":"SELECT count(*) AS rows FROM range(100000000)"}' | python3 -m json.tool
Expect a single row, 100,000,000, with duration_ms in the response. Open the
query id in History for the full operator tree profile.
The warehouse hangs in PROVISIONING. Check the driver task's containers —
the classic failure is the task role missing a permission (read the gateway log for the
exact error) or the ECR image tag referenced by ecs.tf no longer existing in
the repo.
Flashpoint on AWS: I POSTed a warehouse to the gateway, it stays PROVISIONING, and the Fargate driver task in `aws ecs describe-tasks` shows containers with `lastStatus: STOPPED` / exit code [X]. The gateway log shows: [paste the error]. What is wrong — task role IAM, ECR image tag in infra/ecs.tf, or something in the driver container entrypoint?
With the gateway up and a warehouse run behind you, Cost Center is live:
Project=flashpoint cost
allocation tag. The tag must be active in Billing for the filter to return
anything; once active, CE data has ~24h lag, so today's spend is fused from live
per-warehouse meters.FLASHPOINT_MONTHLY_BUDGET (default $20). Spike days (over 1.5× the
window average) render in red on the chart.Cost Center shows resources but the chart is all zeros. That's the cost allocation tag:
check its status in Billing → Cost allocation tags, or with
aws ce list-cost-allocation-tags. The tag change propagates within ~24h and
does not backfill.
Flashpoint Cost Center on AWS shows the EC2/Fargate resources but the daily spend chart is all $0.00. `aws ce get-cost-and-usage` with a Filter on Tags Project=flashpoint returns zero amounts, but the same query without the filter shows real numbers. I checked `aws ce list-cost-allocation-tags` — what does an Inactive status mean and how do I activate the Project tag?
| resource | detail | cost |
|---|---|---|
| gateway EC2 | t4g.small, always on | ~$12.26/mo |
| gateway EBS | 20 GB gp3 root volume | ~$1.60/mo |
| warehouses | XS = driver + 1 spot executor | ~$0.08/h while running |
| infra pennies | ECR, S3, DynamoDB, log groups | ~$0.10/mo |
Warehouses auto-suspend after their TTL (2h default) — usage meters keep accruing per
session so Cost Center reconciles what actually ran. Pausing a demo without losing the
stack is one command: aws ec2 stop-instances (keeps the EBS volume and boots
the already-installed gateway next time; the ~$1.60/mo volume charge stays).
tofu destroy removes the entire Flashpoint stack: gateway EC2 + its root EBS,
ECS cluster and task definitions, DynamoDB tables, S3 bucket, ECR repo, log groups, and the
VPC. Run it from infra/:
cd infra tofu destroy -var-file=dev.tfvars -auto-approve
Then verify nothing is left running or billed:
# these should all come back empty:
aws ec2 describe-instances --query 'Reservations[].Instances[].InstanceId'
aws ec2 describe-volumes --query 'Volumes[].VolumeId'
aws ecs list-tasks --cluster flashpoint-dev-cluster
aws dynamodb list-tables --query 'TableNames[?starts_with(@,`flashpoint`)]'
aws s3 ls
tofu destroy fails halfway (a resource already deleted manually, a VPC
dependency, an S3 bucket that isn't empty). The state may also drift if anything was
changed by hand since apply.
Flashpoint teardown: `tofu destroy -var-file=dev.tfvars` in the repo's infra/ fails with [paste the error]. I previously applied the stack and then manually deleted/stopped some resources in the AWS console. How do I reconcile the state file (tofu state rm?) so destroy completes and Flashpoint stops billing?
quickstart.html) vs. AWS — profile and cost both.aws ce get-cost-and-usage with the Project filter).docs/ — ADRs, warehouse sizing, Fargate notes.