Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
52 changes: 52 additions & 0 deletions docs/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,52 @@
# docs/

Operational documentation for the `perishdev/infra` repo. The [top-level `CLAUDE.md`](../CLAUDE.md) holds the locked design decisions; the docs below explain how those decisions play out in day-to-day use.

## Reading order for first contact

1. [`../CLAUDE.md`](../CLAUDE.md) — design decisions table. Authoritative.
2. [`../README.md`](../README.md) — what this repo is.
3. [`../CONTRIBUTING.md`](../CONTRIBUTING.md) — five-minute orientation for contributors and AI agents.
4. [`state.md`](./state.md) — how HCP state is structured + the API access pattern.
5. [`ci.md`](./ci.md) — the required-checks contract for any merge.

The rest is reference, dipped into as needed.

## Reference

### Setup and onboarding

- [`setup.md`](./setup.md) — one-time bootstrap of HCP, Cloudflare token, GitHub App, local dev. Done once per maintainer's laptop.
- [`import.md`](./import.md) — `cf-terraforming` runbook for adopting existing Cloudflare state into Terraform.
- [`worktree-workflow.md`](./worktree-workflow.md) — optional `git worktree` convention for maintainers juggling multiple branches.

### Design contracts

- [`secrets.md`](./secrets.md) — where every secret lives, how rotation works, what isn't a secret.
- [`state.md`](./state.md) — HCP backend, workspace layout, API access from CLI.
- [`ci.md`](./ci.md) — workflow contract, fork-PR policy, branch protection requirements.

### Operations

- [`recipes.md`](./recipes.md) — common-task recipes: add a DNS record, add a repo, add a label, bump a provider, cross-workspace changes.
- [`rollback.md`](./rollback.md) — six options when an apply made things worse, ranked from cheapest to last-resort.
- [`hcp-api.md`](./hcp-api.md) — HCP REST API toolkit. Read plan summaries, confirm applies, find runs, all from `curl`.
- [`limits.md`](./limits.md) — vendor free-tier limits and where the cliffs are.

## When to add a new doc

Add a new file when:

- A future maintainer or agent will need to look something up by topic. The lookup should be a single grep / open.
- The information isn't easily derived from reading code or running a command.
- The information will be stable for at least a few months — short-lived state goes in commit messages, PR descriptions, or issue threads.

Don't add a new file for:

- One-off tasks (PR description suffices).
- Things that duplicate provider documentation (link to the vendor instead).
- Things that contradict [`../CLAUDE.md`](../CLAUDE.md) — fix the design decisions table first, then write the doc.

## When to delete a doc

When it's wrong and not worth fixing. A wrong doc is worse than no doc.
179 changes: 179 additions & 0 deletions docs/hcp-api.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,179 @@
# HCP Terraform API toolkit

The HCP REST API is the same surface the UI uses. Everything in this repo's day-to-day operations can be scripted against it, often more precisely than clicking through the web app.

## Authentication

`terraform login` deposits a user API token at `~/.terraform.d/credentials.tfrc.json`. That token is what authenticates every snippet below.

```sh
HCP_TOKEN=$(jq -r '.credentials["app.terraform.io"].token' ~/.terraform.d/credentials.tfrc.json)
H_AUTH="Authorization: Bearer $HCP_TOKEN"
H_TYPE="Content-Type: application/vnd.api+json"
```

For org-level operations (managing teams, oauth clients), the user must be an org admin in HCP. For most read operations a workspace member token is enough.

## Identifiers you'll need often

| Name | Value | How to find |
|---|---|---|
| Organization | `perishdev` | the URL `app.terraform.io/app/perishdev/...` |
| `cloudflare` workspace | `ws-WWKeFPiCAjV4STNX` | `GET /organizations/perishdev/workspaces/cloudflare` |
| `github-org` workspace | `ws-iZFBsNEsUfJPRJ1Q` | `GET /organizations/perishdev/workspaces/github-org` |
| `infra` project | (find via API) | `GET /organizations/perishdev/projects` |
| GitHub OAuth client | (find via API) | `GET /organizations/perishdev/oauth-clients` |

## Find a workspace

```sh
WS_ID=$(curl -s "https://app.terraform.io/api/v2/organizations/perishdev/workspaces/cloudflare" \
-H "$H_AUTH" | jq -r '.data.id')
echo "$WS_ID"
```

## List recent runs on a workspace

```sh
curl -s "https://app.terraform.io/api/v2/workspaces/$WS_ID/runs?page%5Bsize%5D=10" \
-H "$H_AUTH" \
| jq '[.data[] | {
id,
status: .attributes.status,
plan_only: .attributes."plan-only",
message: (.attributes.message[:80])
}]'
```

Note: **speculative runs are filtered out of the default list.** To see them, add `&filter[status]=planned_and_finished,planning,planned`. That's how you find a PR's speculative plan run.

## Read a plan summary (creates / updates / destroys / imports)

```sh
RUN_ID=run-xxxxxxxxxxxxxxxx
PLAN_ID=$(curl -s "https://app.terraform.io/api/v2/runs/$RUN_ID" \
-H "$H_AUTH" | jq -r '.data.relationships.plan.data.id')

curl -sL "https://app.terraform.io/api/v2/plans/$PLAN_ID/json-output-redacted" \
-H "$H_AUTH" \
| jq '{
summary: {
creates: ([.resource_changes[]? | select(.change.actions[]? == "create" and (.change.importing // null) == null)] | length),
imports: ([.resource_changes[]? | select(.change.importing != null)] | length),
updates: ([.resource_changes[]? | select(.change.actions[]? == "update")] | length),
destroys: ([.resource_changes[]? | select(.change.actions[]? == "delete")] | length)
},
by_action: [.resource_changes[] | "\(.change.actions | join(",")) \(.address)"]
}'
```

This is faster and more precise than reading the HCP UI's plan output. Use it before merging any PR.

## Confirm an apply

After merge, HCP plans the run and sits at "needs confirmation." Confirm:

```sh
curl -s -X POST "https://app.terraform.io/api/v2/runs/$RUN_ID/actions/apply" \
-H "$H_AUTH" -H "$H_TYPE" \
-d '{"comment":"plan reviewed via API"}'
```

The HTTP 202 is the confirmation. Status transitions to `applying` then `applied` (or `errored`).

## Discard / cancel a run

```sh
# Plan reviewed, not what we wanted — discard before confirming apply
curl -s -X POST "https://app.terraform.io/api/v2/runs/$RUN_ID/actions/discard" \
-H "$H_AUTH" -H "$H_TYPE" -d '{"comment":"discarding — see revert PR"}'

# Apply currently running and we want to stop
curl -s -X POST "https://app.terraform.io/api/v2/runs/$RUN_ID/actions/cancel" \
-H "$H_AUTH"
```

Cancel is the harder of the two — partial state is possible mid-apply.

## Wait for a run to reach a terminal state

```sh
poll_run() {
local rid=$1
while :; do
local st
st=$(curl -s "https://app.terraform.io/api/v2/runs/$rid" \
-H "$H_AUTH" | jq -r '.data.attributes.status')
echo "$st"
case "$st" in
applied|errored|canceled|discarded|force_canceled|planned_and_finished)
return 0
;;
esac
sleep 15
done
}
```

## List variables on a workspace

```sh
curl -s "https://app.terraform.io/api/v2/workspaces/$WS_ID/vars" \
-H "$H_AUTH" \
| jq '[.data[] | {
id,
key: .attributes.key,
category: .attributes.category,
sensitive: .attributes.sensitive
}]'
```

Sensitive variable values are never returned by the API — only `"sensitive": true` and no `value` field. This is how it should be.

## Find the speculative plan run for a specific commit

```sh
COMMIT_SHA=...
curl -s "https://app.terraform.io/api/v2/workspaces/$WS_ID/runs?page%5Bsize%5D=20&filter%5Bstatus%5D=planned_and_finished,planning,planned" \
-H "$H_AUTH" \
| jq --arg sha "$COMMIT_SHA" '
[.data[]
| select(.attributes."plan-only" == true)
| select(.attributes.message | startswith("Merge") or contains($sha[:8]))
| {id, message: (.attributes.message[:80])}]'
```

(The PR's merge commit message contains the PR title, so the easiest match is by message substring.)

## Trigger a manual plan from the API

Useful for "the workspace's auto-trigger didn't fire" debugging.

```sh
curl -s -X POST "https://app.terraform.io/api/v2/runs" \
-H "$H_AUTH" -H "$H_TYPE" \
-d "{
\"data\": {
\"attributes\": { \"message\": \"manual diagnostic\" },
\"type\": \"runs\",
\"relationships\": {
\"workspace\": { \"data\": { \"type\": \"workspaces\", \"id\": \"$WS_ID\" } }
}
}
}"
```

## Things the API can do that the UI can't

- **Read the full structured plan JSON** (`plans/<id>/json-output-redacted`). The UI shows a rendered diff; the API gives you the underlying data to script against.
- **Atomic operations across multiple workspaces** in a script — the UI is one workspace at a time.
- **Audit / activity scripting** — list every run by every user, filter by date, dump to CSV.

## What still needs the UI

- The "Update VCS settings" button that re-registers the webhook with GitHub. The API equivalent exists but is fiddly; UI click is fastest.
- Inspecting workspace settings interactively when you don't know what you're looking for.

## Full API reference

<https://developer.hashicorp.com/terraform/cloud-docs/api-docs>
86 changes: 86 additions & 0 deletions docs/limits.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,86 @@
# Vendor limits

What free tiers cover, where the cliffs are, and how close we are. Updated by hand; numbers age — verify against vendor docs before making a decision based on these.

## HCP Terraform — Free tier

| Resource | Free tier | We're at |
|---|---|---|
| Managed resources | **500** | ~15 |
| Active users | **3** | 1 |
| Self-service workspaces | unlimited | 2 |
| Concurrent runs | 1 | enough |
| Run minutes | not metered on Free | n/a |
| API rate limit | 30 req/sec authenticated | well under |

**Cliff to watch:** the 500-resource limit. Once over, **Standard tier** charges $0.00014 per resource-hour above 500 — roughly $0.10/resource/month. At our current scale (15 resources), we have ~33× headroom. At 100 resources, still no charge. At 600, ~$10/month.

**When we'd cross 500:** unlikely without a deliberate change in scope (e.g. managing many small Cloudflare zone settings as individual resources, or onboarding a large set of GitHub repos). Worth re-checking before bulk-importing anything.

**Doc:** <https://www.hashicorp.com/products/terraform/pricing>

## Cloudflare — Free plan

| Resource | Free plan | Notes |
|---|---|---|
| Zones | 1 per account | we're at 1 (`perish.dev`) |
| DNS records | unmetered | we're at ~11 |
| Email Routing rules | unmetered (but limits on destinations) | a few |
| Workers requests | 100k/day | not using |
| R2 storage | 10 GB | not using |
| Pages projects | unlimited | not using (we use GitHub Pages) |
| Page Rules (legacy) | 3 | 0 active |
| Redirect Rules (Rulesets) | 10 | 0 |
| API rate limit | **1200 req / 5 min** | matters for big imports |

**Cliff to watch:** the API rate limit during bulk operations (cf-terraforming on a huge account, or a Terraform plan that touches many resources). For our current scale, irrelevant.

**Doc:** <https://developers.cloudflare.com/fundamentals/api/reference/limits/>

## GitHub — Free plan (for personal accounts and orgs)

| Resource | Free plan | Notes |
|---|---|---|
| Public repos | unlimited | we're at 2 |
| Private repos | unlimited | we're at 0 |
| GitHub Actions minutes | 2000/month (private repos only; public repos are free) | not metered for us |
| Storage for packages/Actions | 500 MB | unused |
| GitHub Pages bandwidth | **100 GB/month soft limit** | landing page — way under |
| GitHub Pages builds | 10/hour | a few during bootstrap |
| API rate limit (token) | **5000/hour** authenticated | well under |
| API rate limit (anon) | 60/hour | unused |
| Branch protection | available on public repos free | enabled |
| Required reviewers | available | not used (solo) |

**Cliff to watch:** GitHub Pages 100 GB/month bandwidth if `perish.dev` ever gets viral. Currently a static landing page; well under.

**Doc:** <https://docs.github.com/en/get-started/learning-about-github/githubs-plans>

## Let's Encrypt (via GitHub Pages)

| Resource | Limit | Notes |
|---|---|---|
| Certificates per registered domain per week | 50 | n/a — we get 1 |
| Failed validations per hour | 5 | bit us during the cert-wedge debug |
| Renewals | automatic, GitHub handles | every ~60 days for a 90-day cert |

**Cliff to watch:** during cert-provisioning debugging (the toggle dance), do NOT loop on failed setups — Let's Encrypt rate-limits failures and locks you out for an hour. Wait between attempts.

**Doc:** <https://letsencrypt.org/docs/rate-limits/>

## What we'd pay if every meter ran red

Pessimistic upper bound at this repo's *current* shape:

| Vendor | Free → first-paid threshold | Estimated cost at 2× current usage |
|---|---|---|
| HCP Terraform | 500 resources | $0 |
| Cloudflare | sustained traffic on Pro features | $0 (we're not on Pro) |
| GitHub | 100 GB Pages bandwidth | $0 |
| Let's Encrypt | always free | $0 |

Total: **$0/month** at current usage; no realistic path to a bill until the scope of what's managed grows substantially.

## Monitoring (deliberately absent)

We don't have alerting on any of these thresholds because none is close. If/when we approach a cliff, add a `docs/monitoring.md` for whatever check triggers it. For now: re-read this doc once a year, or whenever scope changes meaningfully.
Loading
Loading