CI/CD without keys
Build, test and deploy Sniplink from GitHub with Cloud Build and a least-privilege service account, without a single long-lived key.
What breaks
Section titled “What breaks”Nothing in that log is exotic. Someone needed CI to deploy, so they created a JSON key for a service account, gave the account Owner “to stop the permission errors”, and pasted the key into a CI secret. Two and a half years later a debug flag printed the environment, the job log was readable by anyone with access to the repo, and someone used the key to mine crypto for three days. The key was valid the whole time because a service-account key does not expire unless you make it.
Until now you deployed Sniplink from your laptop, as yourself. That does not scale past one person and it is not reviewable: nobody can see which commit is running. This module moves the deploy into CI, and the question it answers is: how does a pipeline prove to Google Cloud who it is, without a secret that can leak?
Why long-lived keys are the problem
Section titled “Why long-lived keys are the problem”A service-account key is a private key in a JSON file. Whoever holds the file is the service account, from any machine on the internet, until the key is deleted. Compare that with everything else you have used in this course:
- Your laptop uses your user credentials via
gcloud auth application-default login. They are tied to your Google account, MFA and session policies. - Cloud Run gets short-lived tokens for
sniplink-runtimefrom the metadata server. No file exists. A token lives about an hour and only works for that identity.
A JSON key breaks that model in four ways. It has no expiry by default. It is portable, so it works from an attacker’s machine as well as yours. It gets copied (CI secrets, .env files, a colleague’s Downloads folder), and you cannot know all the copies. And rotation is manual, so in practice it does not happen.
Google Cloud gives you two keyless ways to run CI:
- Run the pipeline inside Google Cloud. Cloud Build executes each build as a service account you choose. The build gets short-lived tokens from the metadata server, exactly like Cloud Run. There is no key to store.
- Federate an external CI. GitHub Actions, GitLab and others issue a signed OIDC token for each job. Workload Identity Federation lets Google Cloud trust that token, check its claims (which repository, which branch) and exchange it for a short-lived Google token. Again, no key.
The main path in this course is Cloud Build, because it is Google’s first-party tool and the identity story is the simplest. GitHub Actions with Workload Identity Federation comes as a tab later, and it is a perfectly good choice.
Make keys impossible, not just unused
Section titled “Make keys impossible, not just unused”The best defence against the incident above is a policy that refuses to create keys at all. The organisation policy constraint iam.disableServiceAccountKeyCreation does exactly that. Google enforces it by default, in its newer managed form iam.managed.disableServiceAccountKeyCreation, on every organisation created on or after 3 May 2024, as part of its security baseline (some organisations created from February to April 2024 have it too). If your project belongs to an organisation, check what applies to it:
export PROJECT_ID=your-project-idexport REGION=europe-west1
# Shows the effective policy, including inherited rulesgcloud org-policies describe iam.disableServiceAccountKeyCreation \ --project="$PROJECT_ID" --effective# The managed form that new organisations enforce by defaultgcloud org-policies describe iam.managed.disableServiceAccountKeyCreation \ --project="$PROJECT_ID" --effectiveIf your course project has no organisation (a personal Gmail account), there is nowhere to set org policies, and the commands report that. That is fine for this course: the lab ends with a self-check that no user-managed keys exist. In a company, I treat this constraint as non-negotiable and handle the rare exception with a time-boxed policy override on one project.
A build identity with least privilege
Section titled “A build identity with least privilege”When you create a Cloud Build trigger you choose the service account its builds run as. Older projects had a “legacy Cloud Build service account” with broad project-level roles; do not use it, and do not use the default compute SA either. Create a dedicated one: sniplink-build.
The rule is the same as for sniplink-runtime in module 3: grant each role on the narrowest resource that supports it. Here is what the pipeline needs, and why.
| Role | Scope | Why |
|---|---|---|
roles/artifactregistry.writer |
the sniplink repository |
Push images. Includes read, so the build can also resolve digests. Nothing outside this repo. |
roles/run.developer |
the sniplink-api service |
Update the service to a new revision. Not run.admin: developer cannot change the service’s IAM policy, so CI cannot make a private service public. |
roles/run.viewer |
project | Wait for the update to finish. Pulumi polls the long-running operation, and Cloud Run operations (run.operations.get) live at project level, outside the service’s IAM policy. Read-only: CI can list services and revisions but change none of them. |
roles/iam.serviceAccountUser |
the sniplink-runtime SA |
Deploying a revision that runs as sniplink-runtime requires permission to “act as” it. Granted on that one SA, so CI cannot deploy something that runs as any other identity. |
roles/storage.objectAdmin |
the PROJECT_ID-pulumi-state bucket |
Read and write Pulumi state and its lock files. Object-level only; CI cannot delete the bucket or change its settings. |
roles/cloudkms.cryptoKeyEncrypterDecrypter |
the pulumi/state key |
Pulumi decrypts config secrets (the lab key) and encrypts secret outputs in state. |
roles/secretmanager.admin |
the sniplink-lab-key secret |
The stack manages this secret: adding versions on rotation, destroying old ones, and the runtime SA’s binding on it. Admin on one secret, not on Secret Manager. |
roles/logging.logWriter |
project | Builds with a user-specified SA write their logs to Cloud Logging. This role has no narrower scope. |
Note what is not there: no Editor, no Owner, no project-level run.admin, no iam.serviceAccountAdmin, no access to the GitHub token secret you create later.
Add the build identity in a new file. The state bucket and the KMS key come from the bootstrap script in module 0, so the stack references them by name instead of creating them.
using Pulumi;using Gcp = Pulumi.Gcp;
namespace Sniplink.Infra;
public static class BuildIdentity{ public static Gcp.ServiceAccount.Account Create( string project, string region, Gcp.ArtifactRegistry.Repository repo, Gcp.CloudRunV2.Service service, Gcp.ServiceAccount.Account runtimeSa, Gcp.SecretManager.Secret labSecret) { var buildSa = new Gcp.ServiceAccount.Account("sniplink-build", new() { AccountId = "sniplink-build", DisplayName = "Sniplink CI/CD (Cloud Build)", }); var member = buildSa.Email.Apply(e => $"serviceAccount:{e}");
new Gcp.ArtifactRegistry.RepositoryIamMember("build-ar-writer", new() { Location = repo.Location, Repository = repo.Name, Role = "roles/artifactregistry.writer", Member = member, });
new Gcp.CloudRunV2.ServiceIamMember("build-run-developer", new() { Location = service.Location, Name = service.Name, Role = "roles/run.developer", Member = member, });
new Gcp.Projects.IAMMember("build-run-viewer", new() { Project = project, Role = "roles/run.viewer", Member = member, });
new Gcp.ServiceAccount.IAMMember("build-actas-runtime", new() { ServiceAccountId = runtimeSa.Name, Role = "roles/iam.serviceAccountUser", Member = member, });
new Gcp.Storage.BucketIAMMember("build-state-bucket", new() { Bucket = $"{project}-pulumi-state", Role = "roles/storage.objectAdmin", Member = member, });
new Gcp.Kms.CryptoKeyIAMMember("build-state-key", new() { CryptoKeyId = $"projects/{project}/locations/{region}/keyRings/pulumi/cryptoKeys/state", Role = "roles/cloudkms.cryptoKeyEncrypterDecrypter", Member = member, });
new Gcp.SecretManager.SecretIamMember("build-labkey-admin", new() { SecretId = labSecret.SecretId, Role = "roles/secretmanager.admin", Member = member, });
new Gcp.Projects.IAMMember("build-log-writer", new() { Project = project, Role = "roles/logging.logWriter", Member = member, });
return buildSa; }}using Sniplink.Infra;// …var gcpConfig = new Config("gcp");var project = gcpConfig.Require("project");var region = gcpConfig.Require("region");// … after repo, service, runtimeSa and labSecret are createdvar buildSa = BuildIdentity.Create(project, region, repo, service, runtimeSa, labSecret);BuildIdentity lives in the Sniplink.Infra namespace and the top-level statements in Program.cs do not, hence the using. The module 5 program only read gcp:region; the build identity also needs the project ID, so the program now reads gcp:project too. Both keys have been in Pulumi.dev.yaml since module 1.
resource "google_service_account" "build" { account_id = "sniplink-build" display_name = "Sniplink CI/CD (Cloud Build)"}
locals { build_member = "serviceAccount:${google_service_account.build.email}"}
resource "google_artifact_registry_repository_iam_member" "build_writer" { location = google_artifact_registry_repository.sniplink.location repository = google_artifact_registry_repository.sniplink.name role = "roles/artifactregistry.writer" member = local.build_member}
resource "google_cloud_run_v2_service_iam_member" "build_developer" { location = google_cloud_run_v2_service.sniplink_api.location name = google_cloud_run_v2_service.sniplink_api.name role = "roles/run.developer" member = local.build_member}
resource "google_project_iam_member" "build_run_viewer" { project = var.project role = "roles/run.viewer" member = local.build_member}
resource "google_service_account_iam_member" "build_actas_runtime" { service_account_id = google_service_account.runtime.name role = "roles/iam.serviceAccountUser" member = local.build_member}
resource "google_storage_bucket_iam_member" "build_state" { bucket = "${var.project}-pulumi-state" role = "roles/storage.objectAdmin" member = local.build_member}
resource "google_kms_crypto_key_iam_member" "build_state_key" { crypto_key_id = "projects/${var.project}/locations/${var.region}/keyRings/pulumi/cryptoKeys/state" role = "roles/cloudkms.cryptoKeyEncrypterDecrypter" member = local.build_member}
resource "google_secret_manager_secret_iam_member" "build_labkey" { secret_id = google_secret_manager_secret.lab_key.secret_id role = "roles/secretmanager.admin" member = local.build_member}
resource "google_project_iam_member" "build_logs" { project = var.project role = "roles/logging.logWriter" member = local.build_member}In a Terraform version of this course the state would live in a GCS backend and the KMS binding would be unnecessary; it is kept here so the tabs stay comparable.
The honest part: CI that manages IAM
Section titled “The honest part: CI that manages IAM”Read the table again with an attacker’s eyes. pulumi up in CI applies whatever the code says. If the code in main adds a binding that grants someone Owner, CI will try to apply it. Least privilege on the build SA is what stops that: sniplink-build has no permission to change project IAM, its own bindings, or the service’s IAM policy, so such a change fails with 403 instead of succeeding.
This works because Pulumi only calls the APIs for resources that changed. A normal commit changes the image, the revision name and a few env vars, so CI touches the Cloud Run service and nothing else. When a commit changes BuildIdentity.cs or the trigger, CI fails and a human with more rights has to apply it from their machine. I consider that failure a feature: changes to who-can-do-what get a second pair of eyes by construction.
The price is that one stack now has two kinds of resources: app resources that CI applies, and platform resources (the build SA, its bindings, the GitHub connection, the trigger) that only a human applies. There is a second, quieter cost. CI holds cryptoKeyEncrypterDecrypter on the key that encrypts every config secret in this stack, so anything you store in Pulumi.dev.yaml is readable by a build.
The pattern that removes both problems is two stacks: a bootstrap or platform stack, owned by an admin group and applied from a protected environment, that holds identities, triggers, the connection and its own KMS key; and an app stack that CI applies, which only receives references (the build SA email, the runtime SA) as stack references. For a course with one student and one service I keep a single stack and accept the trade-off. On a team I would split on day one. Be aware which side of the line you are on.
Connecting GitHub: the second deliberate exception
Section titled “Connecting GitHub: the second deliberate exception”Cloud Build’s 2nd-gen repositories connect a Cloud Build connection (one per GitHub account or organisation, per region) to repositories inside it. Triggers then point at a repository resource. All of that is Pulumi code, except one step that cannot be: GitHub has to agree to let Google read your repository. That consent happens by installing the Google Cloud Build GitHub App on your account, in GitHub’s UI. This is the second and last exception to the zero-ClickOps rule, and it is acceptable for the same reason as the bootstrap script: it happens once, it is in GitHub rather than in your cloud project, and everything after it is code.
-
Install the Google Cloud Build app from the GitHub Marketplace on your GitHub account. When GitHub asks which repositories, choose Only select repositories and pick your Sniplink repository. The app should see one repo, not all of them.
-
Find the installation ID. Open GitHub Settings → Applications → Installed GitHub Apps → Google Cloud Build → Configure. The URL ends with
/installations/12345678; that number is the ID. -
Create a GitHub personal access token that Cloud Build uses to read the repository. Use a classic token with the
repoandread:userscopes (addread:orgif the app is installed in an organisation), and give it an expiry date. Cloud Build does not accept fine-grained tokens here. -
Store the token and the installation ID in stack config. The token is encrypted with your KMS key, like the lab key in module 3.
cd infrapulumi config set --secret sniplink:githubToken ghp_xxx # paste your tokenpulumi config set sniplink:githubAppInstallationId 12345678pulumi config set sniplink:githubRepo YOUR_GITHUB_USER/sniplinkpulumi config set sniplink:projectNumber "$(gcloud projects describe "$PROJECT_ID" --format='value(projectNumber)')"Yes, the token now has an expiry, which means the connection breaks when it expires. That is the right failure: you rotate it with pulumi config set --secret and pulumi up, and Secret Manager gets a new version.
Connection, repository and triggers in code
Section titled “Connection, repository and triggers in code”The token goes into Secret Manager, because the connection references a secret version, not a raw string. The Cloud Build service agent (service-PROJECT_NUMBER@gcp-sa-cloudbuild.iam.gserviceaccount.com, a Google-managed identity that Cloud Build uses for its own plumbing) reads it. That is a different identity from sniplink-build, which never gets access to the token.
Two triggers point at the same repository:
- Push to
mainrunscloudbuild.yaml: test, build, push, deploy. - Pull request against
mainrunscloudbuild.pr.yaml: test andpulumi preview, no deploy.
Both run as sniplink-build, set through ServiceAccount.
using Pulumi;using Gcp = Pulumi.Gcp;
namespace Sniplink.Infra;
public static class GitHubPipeline{ public static void Create(string project, string region, Config config, Gcp.ServiceAccount.Account buildSa) { // From config, not a GetProject invoke: invokes run on every `pulumi up`, // and sniplink-build has no permission to read project metadata. var projectNumber = config.Require("projectNumber"); var cloudBuildAgent = $"serviceAccount:service-{projectNumber}@gcp-sa-cloudbuild.iam.gserviceaccount.com";
// 1. GitHub token in Secret Manager var tokenSecret = new Gcp.SecretManager.Secret("github-token", new() { SecretId = "sniplink-github-token", Replication = new Gcp.SecretManager.Inputs.SecretReplicationArgs { Auto = new Gcp.SecretManager.Inputs.SecretReplicationAutoArgs(), }, }); var tokenVersion = new Gcp.SecretManager.SecretVersion("github-token-v", new() { Secret = tokenSecret.Id, SecretData = config.RequireSecret("githubToken"), }); var agentAccess = new Gcp.SecretManager.SecretIamMember("cb-agent-token", new() { SecretId = tokenSecret.SecretId, Role = "roles/secretmanager.secretAccessor", Member = cloudBuildAgent, });
// 2. Connection (one per GitHub account, per region) var connection = new Gcp.CloudBuildV2.Connection("github", new() { Location = region, Name = "github", GithubConfig = new Gcp.CloudBuildV2.Inputs.ConnectionGithubConfigArgs { AppInstallationId = config.RequireInt32("githubAppInstallationId"), AuthorizerCredential = new Gcp.CloudBuildV2.Inputs.ConnectionGithubConfigAuthorizerCredentialArgs { OauthTokenSecretVersion = tokenVersion.Id, }, }, }, new CustomResourceOptions { DependsOn = { agentAccess } });
// 3. Repository inside the connection var githubRepo = config.Require("githubRepo"); // "owner/sniplink" var repository = new Gcp.CloudBuildV2.Repository("sniplink", new() { Location = region, Name = "sniplink", ParentConnection = connection.Name, RemoteUri = $"https://github.com/{githubRepo}.git", });
var substitutions = new InputMap<string> { ["_REGION"] = region, };
// 4a. Push to main → deploy new Gcp.CloudBuild.Trigger("deploy-main", new() { Location = region, Name = "sniplink-deploy-main", RepositoryEventConfig = new Gcp.CloudBuild.Inputs.TriggerRepositoryEventConfigArgs { Repository = repository.Id, Push = new Gcp.CloudBuild.Inputs.TriggerRepositoryEventConfigPushArgs { Branch = "^main$", }, }, Filename = "cloudbuild.yaml", ServiceAccount = buildSa.Id, // projects/PROJECT/serviceAccounts/EMAIL Substitutions = substitutions, });
// 4b. Pull request → test + preview new Gcp.CloudBuild.Trigger("preview-pr", new() { Location = region, Name = "sniplink-preview-pr", RepositoryEventConfig = new Gcp.CloudBuild.Inputs.TriggerRepositoryEventConfigArgs { Repository = repository.Id, PullRequest = new Gcp.CloudBuild.Inputs.TriggerRepositoryEventConfigPullRequestArgs { Branch = "^main$", // Forks need a maintainer's /gcbrun comment before a build runs CommentControl = "COMMENTS_ENABLED_FOR_EXTERNAL_CONTRIBUTORS_ONLY", }, }, Filename = "cloudbuild.pr.yaml", ServiceAccount = buildSa.Id, Substitutions = substitutions, }); }}// …var buildSa = BuildIdentity.Create(project, region, repo, service, runtimeSa, labSecret);GitHubPipeline.Create(project, region, config, buildSa);data "google_project" "this" {}
resource "google_secret_manager_secret" "github_token" { secret_id = "sniplink-github-token" replication { auto {} }}
resource "google_secret_manager_secret_version" "github_token" { secret = google_secret_manager_secret.github_token.id secret_data = var.github_token # sensitive variable}
resource "google_secret_manager_secret_iam_member" "cb_agent_token" { secret_id = google_secret_manager_secret.github_token.secret_id role = "roles/secretmanager.secretAccessor" member = "serviceAccount:service-${data.google_project.this.number}@gcp-sa-cloudbuild.iam.gserviceaccount.com"}
resource "google_cloudbuildv2_connection" "github" { location = var.region name = "github"
github_config { app_installation_id = var.github_app_installation_id authorizer_credential { oauth_token_secret_version = google_secret_manager_secret_version.github_token.id } }
depends_on = [google_secret_manager_secret_iam_member.cb_agent_token]}
resource "google_cloudbuildv2_repository" "sniplink" { location = var.region name = "sniplink" parent_connection = google_cloudbuildv2_connection.github.name remote_uri = "https://github.com/${var.github_repo}.git"}
resource "google_cloudbuild_trigger" "deploy_main" { location = var.region name = "sniplink-deploy-main"
repository_event_config { repository = google_cloudbuildv2_repository.sniplink.id push { branch = "^main$" } }
filename = "cloudbuild.yaml" service_account = google_service_account.build.id substitutions = { _REGION = var.region }}
resource "google_cloudbuild_trigger" "preview_pr" { location = var.region name = "sniplink-preview-pr"
repository_event_config { repository = google_cloudbuildv2_repository.sniplink.id pull_request { branch = "^main$" comment_control = "COMMENTS_ENABLED_FOR_EXTERNAL_CONTRIBUTORS_ONLY" } }
filename = "cloudbuild.pr.yaml" service_account = google_service_account.build.id substitutions = { _REGION = var.region }}A few values deserve a sentence each:
Location = regionon the connection and triggers: 2nd-gen repositories are regional, and a trigger must live in the same region as its connection.Branch = "^main$"is a regular expression. Without the anchors,mainalso matchesnot-main-yet.CommentControlmatters because your repository is public. A pull request from a fork brings its owncloudbuild.pr.yaml, and that file runs assniplink-build, which can decrypt your stack secrets. With this setting, a fork’s PR builds only after an owner or collaborator comments/gcbrun. Read the diff before you type it.ServiceAccounttakes the full resource nameprojects/PROJECT_ID/serviceAccounts/EMAIL, which is whatAccount.Idreturns.
Apply this once from your laptop. You are an Owner of the project, so you have the iam.serviceAccounts.actAs permission that creating a trigger with a custom service account requires.
cd infrapulumi upAt this point the service code is still the module 5 version, and the stack config still holds your module 5 values, so the preview must show only creates (build SA, bindings, token secret, connection, repository, two triggers) and no change to sniplink-api. If it wants to update the service, stop and find out why before you continue.
The pipeline: cloudbuild.yaml
Section titled “The pipeline: cloudbuild.yaml”Cloud Build reads a YAML file with a list of steps. Each step is a container image plus arguments. Steps run in order by default, share the /workspace directory (your checked-out repo), and share the Docker daemon, so an image built in one step is visible in the next.
Before the file, the Pulumi program has to change in three ways: revision names come from the build, the service learns which commit and build produced it, and the stack exports what it deployed.
From hand-named releases to one revision per build
Section titled “From hand-named releases to one revision per build”In module 5 you named every revision yourself (sniplink-api-v1, sniplink-api-v2) through sniplink:release, and drove traffic with three config keys. CI keeps the second half and replaces the first:
- Revision name. Every build creates a revision named
sniplink-api-<shortSha>, for examplesniplink-api-3f9c2ab. The build passes$SHORT_SHAassniplink:revisionSuffix, so you never type a name, two builds never collide, and every revision name points at a commit. Short SHAs are hex, so they can never clash with thev1/v2names from module 5, which stay in the revision list and remain valid rollback targets. - Traffic. Still stack config, committed in git, changed by pull request. One new key,
sniplink:trafficMode, chooses between two shapes:promote: the revision this build creates gets 100 %. This is the everyday mode.canary:sniplink:stableRevision(a name you commit) keeps100 - canaryPercent%, and the revision this build creates getscanaryPercent% and thecanarytag.
| Key | Set by | Example | Meaning |
|---|---|---|---|
sniplink:image |
CI, per build | …/api@sha256:… |
Image digest for this build’s revision |
sniplink:revisionSuffix |
CI, per build | 3f9c2ab |
Revision name becomes sniplink-api-3f9c2ab |
sniplink:commitSha, sniplink:buildId |
CI, per build | full SHA, UUID | COMMIT_SHA and BUILD_ID env vars |
sniplink:version |
you, by PR | 2.1.0 |
APP_VERSION env var |
sniplink:trafficMode |
you, by PR | promote |
promote or canary |
sniplink:stableRevision |
you, by PR | sniplink-api-3f9c2ab |
Keeps the rest of the traffic in canary mode |
sniplink:canaryPercent |
you, by PR | 0 |
Share for this build’s revision in canary mode |
sniplink:studentId |
you, once (module 1) | cd_… |
LABKIT_STUDENT_ID env var |
A release now looks like this. Merge a change with trafficMode: canary, stableRevision set to the revision that serves today, and canaryPercent: 0; the build deploys the new code behind the canary tag. Raise canaryPercent by PR. Promote by PR with trafficMode: promote. Roll back by PR with trafficMode: canary, stableRevision set to the old revision and canaryPercent: 0. Each of those merges runs a build, and the build gets a new revision name even when only config changed: the revision behind the tag is renamed, its code is the same. Routing changes the moment pulumi up runs; the minutes before that are the build. For an emergency, the gcloud run services update-traffic escape hatch from module 5 still works, followed by the PR that makes git agree.
Switch the committed config over once, before you push the new program:
cd infrapulumi config rm sniplink:release # replaced by revisionSuffix from CIpulumi config rm sniplink:canaryRevision # in canary mode the canary is "this build"pulumi config rm sniplink:image # CI passes the digest on every runpulumi config set sniplink:trafficMode promotepulumi config set sniplink:canaryPercent 0# sniplink:version and sniplink:stableRevision stay; from now on you change them by PRRemoving sniplink:image is deliberate: a pulumi up from your laptop now fails with a missing-config error instead of quietly redeploying an old image.
sniplink:studentId from module 1 stays as well. It is committed stack config and not a secret, so the build reads it from Pulumi.dev.yaml like any other key, and cloudbuild.yaml needs nothing extra for LABKIT_STUDENT_ID.
using Pulumi.Gcp.CloudRunV2; // same usings as in module 5using Pulumi.Gcp.CloudRunV2.Inputs;using Sniplink.Infra;// …// config is Config("sniplink"); repo, runtimeSa, labSecret, labKeyAccess,// labKeyVersion from modules 1 and 3const string serviceName = "sniplink-api";var gcpConfig = new Config("gcp");var project = gcpConfig.Require("project");var region = gcpConfig.Require("region");
var image = config.Require("image"); // repo@sha256:… from CIvar revisionSuffix = config.Require("revisionSuffix"); // $SHORT_SHA from CIvar version = config.Require("version"); // APP_VERSION, committedvar commitSha = config.Get("commitSha") ?? "local";var buildId = config.Get("buildId") ?? "local";var trafficMode = config.Get("trafficMode") ?? "promote";var stableRevision = config.Get("stableRevision");var canaryPercent = config.GetInt32("canaryPercent") ?? 0;var studentId = config.Require("studentId");
var revisionName = $"{serviceName}-{revisionSuffix}";
// Fail at preview time, not halfway through an update.if (trafficMode is not ("promote" or "canary")) throw new ArgumentException("trafficMode must be 'promote' or 'canary'");if (canaryPercent is < 0 or > 100) throw new ArgumentException("canaryPercent must be 0..100");if (trafficMode == "canary" && (stableRevision is null || stableRevision == revisionName)) throw new ArgumentException("canary mode needs a stableRevision other than this build's revision");
var traffics = new InputList<ServiceTrafficArgs>();if (trafficMode == "promote"){ traffics.Add(new ServiceTrafficArgs { Type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION", Revision = revisionName, Percent = 100, });}else{ traffics.Add(new ServiceTrafficArgs { Type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION", Revision = stableRevision!, Percent = 100 - canaryPercent, }); traffics.Add(new ServiceTrafficArgs { Type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION", Revision = revisionName, Percent = canaryPercent, Tag = "canary", });}
var service = new Service(serviceName, new ServiceArgs{ Name = serviceName, Location = region, DeletionProtection = false, Ingress = "INGRESS_TRAFFIC_ALL", Template = new ServiceTemplateArgs { Revision = revisionName, ServiceAccount = runtimeSa.Email, // … MaxInstanceRequestConcurrency, Scaling, Volumes from modules 3 and 4 Containers = { new ServiceTemplateContainerArgs { Image = image, Envs = { new ServiceTemplateContainerEnvArgs { Name = "LABKIT_STUDENT_ID", Value = studentId }, new ServiceTemplateContainerEnvArgs { Name = "APP_VERSION", Value = version }, new ServiceTemplateContainerEnvArgs { Name = "COMMIT_SHA", Value = commitSha }, new ServiceTemplateContainerEnvArgs { Name = "BUILD_ID", Value = buildId }, }, // … ports, resources, probes and volume mounts from modules 1, 3 and 4 }, }, }, Traffics = traffics,}, new CustomResourceOptions { DependsOn = { repo, labKeyAccess, labKeyVersion } });
// … public invoker binding from module 1
var buildSa = BuildIdentity.Create(project, region, repo, service, runtimeSa, labSecret);GitHubPipeline.Create(project, region, config, buildSa);
return new Dictionary<string, object?>{ ["url"] = service.Uri, ["canaryUrl"] = service.TrafficStatuses.Apply(statuses => statuses.FirstOrDefault(s => s.Tag == "canary")?.Uri ?? ""), ["deployedRevision"] = revisionName, ["deployedImage"] = image, ["deployedRevisionSuffix"] = revisionSuffix, ["deployedCommitSha"] = commitSha, ["deployedBuildId"] = buildId,};variable "image" { type = string }variable "revision_suffix" { type = string }variable "app_version" { type = string }variable "commit_sha" { type = string default = "local"}variable "build_id" { type = string default = "local"}variable "traffic_mode" { type = string default = "promote" validation { condition = contains(["promote", "canary"], var.traffic_mode) error_message = "traffic_mode must be promote or canary." }}variable "stable_revision" { type = string default = null}variable "canary_percent" { type = number default = 0}
locals { revision_name = "sniplink-api-${var.revision_suffix}"}
resource "google_cloud_run_v2_service" "sniplink_api" { name = "sniplink-api" location = var.region deletion_protection = false ingress = "INGRESS_TRAFFIC_ALL"
template { revision = local.revision_name service_account = google_service_account.runtime.email # … concurrency, scaling, volumes from modules 3 and 4 containers { image = var.image env { name = "LABKIT_STUDENT_ID" value = var.student_id } env { name = "APP_VERSION" value = var.app_version } env { name = "COMMIT_SHA" value = var.commit_sha } env { name = "BUILD_ID" value = var.build_id } # … ports, resources, probes, volume mounts } }
# promote: this build's revision gets 100 % dynamic "traffic" { for_each = var.traffic_mode == "promote" ? [1] : [] content { type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION" revision = local.revision_name percent = 100 } }
# canary: the committed stable revision keeps the rest dynamic "traffic" { for_each = var.traffic_mode == "canary" ? [1] : [] content { type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION" revision = var.stable_revision percent = 100 - var.canary_percent } }
dynamic "traffic" { for_each = var.traffic_mode == "canary" ? [1] : [] content { type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION" revision = local.revision_name percent = var.canary_percent tag = "canary" } }}
output "deployed_image" { value = var.image }What changed compared with module 5, line by line:
Revision = $"{serviceName}-{revisionSuffix}". Same prefix rule as module 5, but the suffix now comes from the build. Thereleasekey is gone.trafficModereplacescanaryRevision. In module 5 the canary was a revision you named. In CI the canary is always “the revision this build creates”, so there is nothing to name. Only the stable side is named in config, because it is the one revision you deliberately keep.- Both targets stay pinned by name. Nothing points at “latest”, so a build only gets traffic because the config committed with it says so.
COMMIT_SHAandBUILD_ID. LabKit’s/_lab/inforeads them from the environment. Locally they fall back tolocal, so a laptop deploy is visibly not a CI deploy.- The
deployed*outputs let the PR pipeline (below) and your laptop preview against what CI actually deployed.
For the lab, keep trafficMode: promote, so the main URL serves the revision of the commit you just pushed. The final project uses canary mode with exactly this program.
The deploy pipeline
Section titled “The deploy pipeline”Commit this file at the repository root.
# Runs on push to main as sniplink-build. No keys anywhere.substitutions: _REGION: europe-west1 _PULUMI_IMAGE: pulumi/pulumi-dotnet-10.0:3.267.0 # pinned, see the note below
steps: # 1. Unit tests. A red test stops the build here. - id: test name: mcr.microsoft.com/dotnet/sdk:10.0 entrypoint: dotnet args: ['test', 'Sniplink.sln', '--configuration', 'Release']
# 2. Build the image, tagged with the short commit SHA (readable in the registry) - id: build name: gcr.io/cloud-builders/docker env: ['DOCKER_BUILDKIT=1'] args: - build - --tag=$_REGION-docker.pkg.dev/$PROJECT_ID/sniplink/api:$SHORT_SHA - --file=Dockerfile - .
# 3. Push to Artifact Registry as sniplink-build - id: push name: gcr.io/cloud-builders/docker args: ['push', '$_REGION-docker.pkg.dev/$PROJECT_ID/sniplink/api:$SHORT_SHA']
# 4. Resolve the immutable digest; tags can be moved, digests cannot - id: digest name: gcr.io/cloud-builders/docker entrypoint: bash args: - -c - | set -euo pipefail docker inspect --format='{{index .RepoDigests 0}}' \ $_REGION-docker.pkg.dev/$PROJECT_ID/sniplink/api:$SHORT_SHA > /workspace/image-ref.txt echo "Deploying $$(cat /workspace/image-ref.txt)"
# 5. Deploy: the same Pulumi program you ran locally, now with CI's identity - id: deploy name: $_PULUMI_IMAGE dir: infra entrypoint: bash args: - -c - | set -euo pipefail pulumi login gs://$PROJECT_ID-pulumi-state pulumi up --yes --non-interactive --stack dev \ --config sniplink:image="$$(cat /workspace/image-ref.txt)" \ --config sniplink:revisionSuffix=$SHORT_SHA \ --config sniplink:commitSha=$COMMIT_SHA \ --config sniplink:buildId=$BUILD_ID
options: # Required when a build runs as a user-specified service account # (alternatively: a logs bucket you own). logging: CLOUD_LOGGING_ONLY
timeout: 1200sWalk through it once:
- Substitutions. Cloud Build replaces
$NAMEand${NAME}in the file before running anything.PROJECT_ID,BUILD_ID,COMMIT_SHAandSHORT_SHAare built in;COMMIT_SHAandSHORT_SHAare filled only for builds started by a trigger, not for a manualgcloud builds submit. Names starting with_are yours. Thesubstitutions:block gives defaults, and the trigger’sSubstitutionsoverride them, which is how the region in your stack config reaches the build. $$. Because Cloud Build owns$, a dollar sign meant for bash must be written$$.$$(cat …)reaches bash as$(cat …). Forget this and Cloud Build rejects the file for an unknown substitution, which is at least a loud failure.- Step 1 runs tests in the same SDK image your Dockerfile uses. The
bin/andobj/folders it leaves in/workspaceare kept out of step 2 by the.dockerignorefrom module 2. - Steps 2-3 build and push with the tag
$SHORT_SHA, so a human looking at the registry can match images to commits.DOCKER_BUILDKIT=1matters: thecloud-builders/dockerimage still defaults to the legacy builder, which does not set$BUILDPLATFORMand$TARGETARCH, so the module 2 Dockerfile fails at its first line withfailed to parse platform. - Step 4 resolves the digest.
RepoDigestsis filled in only after a push. Deployingapi@sha256:…instead ofapi:abc1234means the revision runs exactly the bytes that were tested, even if someone later pushes a different image under the same tag. - Step 5 runs Pulumi. Both the GCS state backend and the Google provider use Application Default Credentials, and in Cloud Build those come from the metadata server as
sniplink-build. Nothing to configure.--configwrites the four per-build values into the build’s copy ofPulumi.dev.yamlfor this run only; your committed file is not changed. Traffic comes from the committed keys, so what a build does to routing is whatever the merged PR said.
Steps 1 and 2 are independent, so you could start both at once with waitFor: ['-'] on each and waitFor: ['test', 'build'] on the push. I keep the file sequential in the course because it is easier to read in logs, and a failed test then saves the image build time.
The pull request pipeline
Section titled “The pull request pipeline”A PR build answers two questions: do the tests pass, and what would this change do to production? It never deploys.
# Runs on pull requests to main as sniplink-build. Never deploys.substitutions: _PULUMI_IMAGE: pulumi/pulumi-dotnet-10.0:3.267.0 # same pinned tag as cloudbuild.yaml
steps: - id: test name: mcr.microsoft.com/dotnet/sdk:10.0 entrypoint: dotnet args: ['test', 'Sniplink.sln', '--configuration', 'Release']
- id: preview name: $_PULUMI_IMAGE dir: infra entrypoint: bash args: - -c - | set -euo pipefail pulumi login gs://$PROJECT_ID-pulumi-state # Compare against what CI deployed, not against the committed config, # otherwise every preview shows the image and env vars as "changed". pulumi preview --non-interactive --diff --stack dev \ --config sniplink:image="$$(pulumi stack output deployedImage --stack dev)" \ --config sniplink:revisionSuffix="$$(pulumi stack output deployedRevisionSuffix --stack dev)" \ --config sniplink:commitSha="$$(pulumi stack output deployedCommitSha --stack dev)" \ --config sniplink:buildId="$$(pulumi stack output deployedBuildId --stack dev)"
options: logging: CLOUD_LOGGING_ONLYThe preview output in the build log is what a reviewer reads next to the code diff. If a PR that “only renames a variable” shows a Cloud Run replacement, you find out before merging, not after. For a traffic PR (say canaryPercent 0 → 10) the preview shows the traffic change against the revision that is live now; after the merge, the build applies the same split with its own, newly named revision in the canary slot.
Your laptop uses the same trick. To preview locally after module 6, pass the deployed values exactly as cloudbuild.pr.yaml does.
I use the same service account for both triggers to keep the module short. A stricter setup gives PR builds their own SA with read-only roles (state bucket objectViewer, KMS decrypter, and viewer roles on the resources), so a malicious PR cannot write anything even after a careless /gcbrun. Read access to the state bucket is enough because pulumi preview takes no lock on a self-managed backend such as a GCS bucket; pulumi up writes a lock file and needs write access.
First run
Section titled “First run”Commit both YAML files and the infra changes, push to main, and watch the build:
git add cloudbuild.yaml cloudbuild.pr.yaml infra/git commit -m "Deploy from Cloud Build"git push origin main
# Latest builds in your region, then stream one build's loggcloud builds list --region="$REGION" --limit=5gcloud builds log BUILD_ID --region="$REGION" --streamThe logs are in Cloud Logging as well, because of CLOUD_LOGGING_ONLY, so the read-only console view is Cloud Build → History for your region. When the build is green:
curl -s "$(cd infra && pulumi stack output url)/_lab/info" | jq '{commitSha, buildId, revision, version}'commitSha should be the full SHA of your commit, buildId a UUID, and revision sniplink-api- followed by the short SHA.
The same pipeline on GitHub Actions
Section titled “The same pipeline on GitHub Actions”GitHub Actions is the CI most .NET teams already use. The pipeline is the same; what changes is how the job becomes sniplink-build. Here GitHub signs an OIDC token for each job, and Google Cloud trusts it through Workload Identity Federation.
Nothing more to do. The trigger runs cloudbuild.yaml as sniplink-build, and the identity comes from the metadata server inside your project. The build definition is in your repo, the trigger and its identity are in your Pulumi stack, and the logs are in your project’s Cloud Logging.
name: deployon: push: branches: [main]
permissions: contents: read id-token: write # lets the job request a GitHub OIDC token
env: REGION: europe-west1 PROJECT_ID: your-project-id IMAGE: europe-west1-docker.pkg.dev/your-project-id/sniplink/api
jobs: deploy: runs-on: ubuntu-latest steps: - uses: actions/checkout@v7 - uses: actions/setup-dotnet@v6 with: dotnet-version: '10.0.x'
- run: dotnet test Sniplink.sln --configuration Release
# Exchange the GitHub OIDC token for short-lived Google credentials - uses: google-github-actions/auth@v3 with: workload_identity_provider: projects/PROJECT_NUMBER/locations/global/workloadIdentityPools/github/providers/github-oidc service_account: sniplink-build@your-project-id.iam.gserviceaccount.com
- uses: google-github-actions/setup-gcloud@v3 - run: gcloud auth configure-docker ${REGION}-docker.pkg.dev --quiet
- name: Build and push run: | docker build --tag "$IMAGE:${GITHUB_SHA::7}" . docker push "$IMAGE:${GITHUB_SHA::7}" echo "IMAGE_REF=$(docker inspect --format='{{index .RepoDigests 0}}' "$IMAGE:${GITHUB_SHA::7}")" >> "$GITHUB_ENV"
- uses: pulumi/actions@v7 # installs the Pulumi CLI - name: Deploy working-directory: infra run: | pulumi login "gs://${PROJECT_ID}-pulumi-state" pulumi up --yes --non-interactive --stack dev \ --config sniplink:image="$IMAGE_REF" \ --config sniplink:revisionSuffix="${GITHUB_SHA::7}" \ --config sniplink:commitSha="$GITHUB_SHA" \ --config sniplink:buildId="$GITHUB_RUN_ID"GITHUB_RUN_ID is a number, not a UUID, so this variant does not pass the lab’s buildId check. The lab asks for Cloud Build on purpose; use this tab as the pattern for your day job.
The identity side: pool, provider, binding
Section titled “The identity side: pool, provider, binding”Three resources make the GitHub token acceptable to Google Cloud:
- A workload identity pool: a container for external identities in your project.
- A provider in that pool: “trust tokens issued by
https://token.actions.githubusercontent.com”, plus an attribute mapping that copies claims from GitHub’s token into Google attributes, and an attribute condition that rejects every token not from your repository. Without the condition, any GitHub repository in the world could present a valid token to your pool; Google now requires a condition for GitHub providers for that reason. - A binding of
roles/iam.workloadIdentityUseronsniplink-buildto the principal set “all identities in this pool whoserepositoryattribute isowner/sniplink”. That is what lets a job impersonate the SA.
The Security Token Service API (sts.googleapis.com) does the token exchange. It was not in the bootstrap list, so the stack enables it.
using Pulumi;using Gcp = Pulumi.Gcp;
namespace Sniplink.Infra;
public static class GitHubActionsIdentity{ public static void Create(string githubRepo, Gcp.ServiceAccount.Account buildSa) { var sts = new Gcp.Projects.Service("sts", new() { ServiceName = "sts.googleapis.com", DisableOnDestroy = false, });
var pool = new Gcp.Iam.WorkloadIdentityPool("github", new() { WorkloadIdentityPoolId = "github", DisplayName = "GitHub Actions", }, new CustomResourceOptions { DependsOn = { sts } });
new Gcp.Iam.WorkloadIdentityPoolProvider("github-oidc", new() { WorkloadIdentityPoolId = pool.WorkloadIdentityPoolId, WorkloadIdentityPoolProviderId = "github-oidc", DisplayName = "GitHub OIDC", Oidc = new Gcp.Iam.Inputs.WorkloadIdentityPoolProviderOidcArgs { IssuerUri = "https://token.actions.githubusercontent.com", }, AttributeMapping = { ["google.subject"] = "assertion.sub", ["attribute.repository"] = "assertion.repository", ["attribute.ref"] = "assertion.ref", }, // Only tokens from this repository are accepted at all AttributeCondition = $"assertion.repository == '{githubRepo}'", });
// pool.Name = projects/NUMBER/locations/global/workloadIdentityPools/github new Gcp.ServiceAccount.IAMMember("github-impersonates-build", new() { ServiceAccountId = buildSa.Name, Role = "roles/iam.workloadIdentityUser", Member = pool.Name.Apply(n => $"principalSet://iam.googleapis.com/{n}/attribute.repository/{githubRepo}"), }); }}resource "google_project_service" "sts" { service = "sts.googleapis.com" disable_on_destroy = false}
resource "google_iam_workload_identity_pool" "github" { workload_identity_pool_id = "github" display_name = "GitHub Actions" depends_on = [google_project_service.sts]}
resource "google_iam_workload_identity_pool_provider" "github_oidc" { workload_identity_pool_id = google_iam_workload_identity_pool.github.workload_identity_pool_id workload_identity_pool_provider_id = "github-oidc" display_name = "GitHub OIDC"
oidc { issuer_uri = "https://token.actions.githubusercontent.com" }
attribute_mapping = { "google.subject" = "assertion.sub" "attribute.repository" = "assertion.repository" "attribute.ref" = "assertion.ref" }
attribute_condition = "assertion.repository == '${var.github_repo}'"}
resource "google_service_account_iam_member" "github_impersonates_build" { service_account_id = google_service_account.build.name role = "roles/iam.workloadIdentityUser" member = "principalSet://iam.googleapis.com/${google_iam_workload_identity_pool.github.name}/attribute.repository/${var.github_repo}"}Two refinements I use at work. Condition on assertion.repository_id rather than the name, because a deleted and re-created repository (or a renamed one) can reuse the name, but never the numeric ID. And for deploys, add assertion.ref == 'refs/heads/main' to the condition (or a separate provider), so a job on a feature branch cannot impersonate the deploy SA at all.
Cloud Build or GitHub Actions?
Section titled “Cloud Build or GitHub Actions?”| Cloud Build | GitHub Actions + WIF | |
|---|---|---|
| Identity | SA attached to the build; nothing to federate | OIDC federation: pool, provider, condition to get right |
| Where things live | Trigger and identity in your project and your IaC; logs in Cloud Logging | Workflow in the repo; logs and run history in GitHub |
| Ecosystem | Any container as a step; no marketplace | Huge marketplace of actions, matrix builds, environments with approvals |
| Developer experience | Logs and status reach the PR via the GitHub app, but the UI is the Cloud console | PR checks, annotations and re-runs where developers already are |
| Private networking | Private pools can reach resources in your VPC | Needs self-hosted runners for that |
| Free tier | See below | Free on standard runners for public repos; 2,000 minutes a month for private repos on GitHub Free |
My honest take: if your team lives in GitHub and does not need VPC access from CI, Actions with WIF is a fine default and the extra identity setup is a one-time cost. Cloud Build wins when you want the build identity, its logs and its audit trail inside the project, private pools, or zero third-party trust. The skill this module teaches, a keyless identity with a short list of scoped roles, is the same in both.
Cost and build logs
Section titled “Cost and build logs”Cloud Build’s free tier is 2,500 build-minutes per month per billing account, on the default e2-standard-2 machine type in the default pool. Beyond that, the same machine costs $0.006 per build-minute. Google states that the free tier is promotional and may change, so check the pricing page before you rely on it.
A Sniplink build takes a few minutes: tests, an image build, and a dotnet build of the infra project before pulumi up. At around six minutes per push you get roughly 400 builds a month inside the free tier, which a course project will not come near. The other costs from this module are cents: one more secret (the GitHub token) and an image in Artifact Registry per commit. That last one grows quietly; an Artifact Registry cleanup policy that keeps the last N versions is worth adding once you are past the course.
Logs go only to Cloud Logging (CLOUD_LOGGING_ONLY), which has its own free ingestion allowance. Treat build logs as you would treat the job log in the incident: anyone who can read them sees everything a step prints. The pipeline above never prints a secret, and pulumi up masks secret config values in its output.
Goal: a push to main on your public GitHub repository is tested, built and deployed by Cloud Build as sniplink-build, and the running service can prove which commit and build it came from.
Requirements:
- Your Sniplink repository is public on GitHub and contains
cloudbuild.yamlandcloudbuild.pr.yamlat the root. - The build SA, connection, repository and both triggers are created by Pulumi, with the roles from this module and nothing broader.
- The service’s current revision was deployed by the push trigger with
trafficMode: promote, so the main URL servessniplink-api-<shortSha>,COMMIT_SHAis the full 40-character SHA andBUILD_IDis the Cloud Build UUID. - No user-managed service-account keys exist in the project.
Check your lab
What each check verifies:
- Token and health: the baseline from every lab. The service runs on Cloud Run in your project.
- Commit SHA:
/_lab/inforeports acommitShaof 40 lowercase hex characters. Self-reported: the service returns what its env says. - Build ID:
buildIdis a UUID, the format Cloud Build uses. Self-reported as well. - GitHub commit: the platform asks GitHub’s public API whether that commit exists in the repository you entered, and whether
cloudbuild.yamlexists at that commit. This ties the running revision to real code in a real repo.
Be clear about the limit: a determined student could set those env vars by hand from a laptop and pass. A stricter check would read Cloud Build history and the service’s IAM directly, and that is what the optional Verified tier will do later. The self-check below cannot be automated through the URL at all, so it is on you.
Self-check: no keys. List user-managed keys for every service account in the project. Google-managed keys are rotated by Google and are not shown with this filter. Empty output is the pass.
for sa in $(gcloud iam service-accounts list --project="$PROJECT_ID" --format='value(email)'); do gcloud iam service-accounts keys list \ --iam-account="$sa" --managed-by=user --format='value(name)'doneIf anything prints, delete it with gcloud iam service-accounts keys delete KEY_ID --iam-account=SA_EMAIL, then find out why it existed.
Clean up
Section titled “Clean up”The final project in module 7 is a new service, but it reuses the same patterns: a build SA, a GitHub connection, a trigger. You have two options.
Keep everything (recommended if you continue soon). The connection and triggers cost nothing while idle, and the module 7 pipeline can reuse the connection. Cloud Run scales to zero. The only steady costs are cents for the KMS key version, the secrets and stored images.
Destroy the stack (if you are pausing). Run it from your laptop, because only you have the rights to delete IAM bindings and the trigger:
cd infrapulumi destroyThis removes the service, the registry with its images, the secrets, the build SA, the WIF pool (if you created it), the connection and both triggers. pulumi destroy needs the same per-build config as a preview, so pass the deployed values with --config as in cloudbuild.pr.yaml if it complains about sniplink:image. Two things stay and are yours to remove by hand: the Google Cloud Build app installation on GitHub (Settings → Applications), and the GitHub token, which you should revoke in Settings → Developer settings. Deleted workload identity pools stay in a soft-deleted state for 30 days and block the same pool ID until then.
What you learned
Section titled “What you learned”- A service-account key is a non-expiring, copyable credential; keyless CI uses an attached identity (Cloud Build) or federation (Workload Identity Federation), and the
iam.disableServiceAccountKeyCreationpolicy makes keys impossible by default. - A dedicated build SA gets each role on the narrowest resource: repo, service, runtime SA, bucket, key, secret. When CI manages its own IAM, least privilege is what turns a dangerous commit into a failed build, and a separate platform stack is the clean long-term split.
- A 2nd-gen GitHub connection needs one GitHub app install; the token secret, connection, repository and triggers are code, and PR builds from forks need a comment gate.
cloudbuild.yamltests, builds, pushes and deploys by digest, naming each revision after its commit and passingCOMMIT_SHAandBUILD_IDthrough Pulumi config into the service, while PR builds onlypulumi preview. Traffic stays in committed config, so promotion and rollback are pull requests.- GitHub Actions does the same job with an OIDC pool, a provider with an attribute condition, and a
workloadIdentityUserbinding scoped to one repository.