Infrastructure and deployment¶
HOT's standard deployment process for other apps is documented in the DevOps deployment guide. fAIr differs slightly, because it has:
- Versioning of both software, as well as AI models.
- A dedicated dev instance EC2 for easier development with all components.
Currently model development happens in the fAIr-models repo, but this
might eventually move to the fAIr monorepo.
The model flow works like this:
- Each model dir has a
stac-item.json. During development it may point at moving image tags; a release copy is pinned to digests before registration. - A CI matrix workflow builds an image for each dir under
./modelswhen its contents change. Merges publishv<version>andlatesttraining images, plus the matching-inferencetags. - An admin runs
fair basemodel pinonce, then submits the same pinned item to staging and production throughPOST /api/v1/base-models/. - Registration mirrors weights into a versioned artifact path, publishes an
integer STAC version, and updates the model's single Knative service. Staging
gets a tagged route; production moves the live route only after the revision
is Ready, including its
/healthreadiness probe. - A
BaseModeltable holds the model name and its status. The version details live entirely in each environment's STAC catalog. - A scheduled reconciler restores Knative services from active STAC items. Destructive pruning remains disabled until STAC listing is paginated.
flowchart LR
A[fAIr-models PR] -->|CI: validate and test| B[Merge]
B -->|publish training + inference images| C[GHCR]
B --> D[Pin both image refs<br/>to digests once]
C --> D
D -->|same pinned item| E[Staging API]
E --> F[Staging STAC +<br/>staging Knative route]
F -->|train, publish, predict| G{Approved?}
G -->|same pinned item| H[Production API]
H --> I[Production STAC +<br/>live Knative route]
The backend image bundles models/ from the same fAIr-models release as its
fair-py-ops dependency. Registration rejects a pipeline module that is not in
that bundle, so pipeline-code changes require a backend release; metadata,
weights, and inference-image-only changes do not.
Step 1: Development¶
The environment
- Single EC2, lightweight k3s cluster.
- Manually updated and synced with dev.
- Model registration in STAC is all manual.
- Users work on models in development using the
devanddev-inferenceimages built for the pull request. - Development model images (training and inference) are pushed to GHCR.
- On the dev EC2 they run a script to update the dev STAC and knative records.
- Any changes to the frontend / API are manually synced to the dev EC2 instance.
- The dev model can be tested on the dev instance, using the dev STAC, ZenML, knative services.
Step 2: Staging¶
The environment
- Runs all the same components as production, but starts up via a PR from
stagingtomain. - The components run inside the
fair-stagingnamespace of the Kubernetes cluster, under domainhttps://stage.ai.hotosm.org. - Does not run its own
knativecontroller, instead using the cluster-wide instance.
- When we want to stabilise and push out a new model, or updates to the API / website, we use the staging setup.
- First a PR must be raised on the fAIr repo from
staging→main. This will set uphttps://stage.ai.hotosm.orgwith ZenML / STAC / Knative registration. - The staging STAC is separate from production and persists between PRs.
- After merge, pin both image references in a copy of the STAC item with
fair basemodel pin. Register that file through the staging API. It is served athttps://staging-<model>.predict.ai.hotosm.orgwithout changing production traffic. Then test training, publishing, and prediction. - Once it looks good, register the model in production (Step 3) before merging
the PR to
main. Merging shuts the staging env down.
Step 3: Production¶
The environment
- Runs through tagged releases on GitHub, where ArgoCD picks up the latest Helm chart tag and deploys.
- A new tagged version is made from the latest
maincode. - This triggers a redeploy of the fAIr website / API.
- Register the exact pinned STAC item tested in staging through the production API. This publishes the next integer STAC version and moves live traffic to the Ready revision. The images are already in GHCR.