Skip to content

Signed in as

Sign in to Curious Dev Learn

Sign in to verify labs, save your progress and get a certificate. Lessons stay open without an account.

Module 062 h 30 minLab

CI/CD without keys

Build, test and deploy Sniplink from GitHub with Cloud Build and a least-privilege service account, without a single long-lived key.

Nothing in that log is exotic. Someone needed CI to deploy, so they created a JSON key for a service account, gave the account Owner “to stop the permission errors”, and pasted the key into a CI secret. Two and a half years later a debug flag printed the environment, the job log was readable by anyone with access to the repo, and someone used the key to mine crypto for three days. The key was valid the whole time because a service-account key does not expire unless you make it.

Until now you deployed Sniplink from your laptop, as yourself. That does not scale past one person and it is not reviewable: nobody can see which commit is running. This module moves the deploy into CI, and the question it answers is: how does a pipeline prove to Google Cloud who it is, without a secret that can leak?

A service-account key is a private key in a JSON file. Whoever holds the file is the service account, from any machine on the internet, until the key is deleted. Compare that with everything else you have used in this course:

  • Your laptop uses your user credentials via gcloud auth application-default login. They are tied to your Google account, MFA and session policies.
  • Cloud Run gets short-lived tokens for sniplink-runtime from the metadata server. No file exists. A token lives about an hour and only works for that identity.

A JSON key breaks that model in four ways. It has no expiry by default. It is portable, so it works from an attacker’s machine as well as yours. It gets copied (CI secrets, .env files, a colleague’s Downloads folder), and you cannot know all the copies. And rotation is manual, so in practice it does not happen.

Google Cloud gives you two keyless ways to run CI:

  1. Run the pipeline inside Google Cloud. Cloud Build executes each build as a service account you choose. The build gets short-lived tokens from the metadata server, exactly like Cloud Run. There is no key to store.
  2. Federate an external CI. GitHub Actions, GitLab and others issue a signed OIDC token for each job. Workload Identity Federation lets Google Cloud trust that token, check its claims (which repository, which branch) and exchange it for a short-lived Google token. Again, no key.

The main path in this course is Cloud Build, because it is Google’s first-party tool and the identity story is the simplest. GitHub Actions with Workload Identity Federation comes as a tab later, and it is a perfectly good choice.

The best defence against the incident above is a policy that refuses to create keys at all. The organisation policy constraint iam.disableServiceAccountKeyCreation does exactly that. Google enforces it by default, in its newer managed form iam.managed.disableServiceAccountKeyCreation, on every organisation created on or after 3 May 2024, as part of its security baseline (some organisations created from February to April 2024 have it too). If your project belongs to an organisation, check what applies to it:

Terminal window
export PROJECT_ID=your-project-id
export REGION=europe-west1
# Shows the effective policy, including inherited rules
gcloud org-policies describe iam.disableServiceAccountKeyCreation \
--project="$PROJECT_ID" --effective
# The managed form that new organisations enforce by default
gcloud org-policies describe iam.managed.disableServiceAccountKeyCreation \
--project="$PROJECT_ID" --effective

If your course project has no organisation (a personal Gmail account), there is nowhere to set org policies, and the commands report that. That is fine for this course: the lab ends with a self-check that no user-managed keys exist. In a company, I treat this constraint as non-negotiable and handle the rare exception with a time-boxed policy override on one project.

When you create a Cloud Build trigger you choose the service account its builds run as. Older projects had a “legacy Cloud Build service account” with broad project-level roles; do not use it, and do not use the default compute SA either. Create a dedicated one: sniplink-build.

The rule is the same as for sniplink-runtime in module 3: grant each role on the narrowest resource that supports it. Here is what the pipeline needs, and why.

Role Scope Why
roles/artifactregistry.writer the sniplink repository Push images. Includes read, so the build can also resolve digests. Nothing outside this repo.
roles/run.developer the sniplink-api service Update the service to a new revision. Not run.admin: developer cannot change the service’s IAM policy, so CI cannot make a private service public.
roles/run.viewer project Wait for the update to finish. Pulumi polls the long-running operation, and Cloud Run operations (run.operations.get) live at project level, outside the service’s IAM policy. Read-only: CI can list services and revisions but change none of them.
roles/iam.serviceAccountUser the sniplink-runtime SA Deploying a revision that runs as sniplink-runtime requires permission to “act as” it. Granted on that one SA, so CI cannot deploy something that runs as any other identity.
roles/storage.objectAdmin the PROJECT_ID-pulumi-state bucket Read and write Pulumi state and its lock files. Object-level only; CI cannot delete the bucket or change its settings.
roles/cloudkms.cryptoKeyEncrypterDecrypter the pulumi/state key Pulumi decrypts config secrets (the lab key) and encrypts secret outputs in state.
roles/secretmanager.admin the sniplink-lab-key secret The stack manages this secret: adding versions on rotation, destroying old ones, and the runtime SA’s binding on it. Admin on one secret, not on Secret Manager.
roles/logging.logWriter project Builds with a user-specified SA write their logs to Cloud Logging. This role has no narrower scope.

Note what is not there: no Editor, no Owner, no project-level run.admin, no iam.serviceAccountAdmin, no access to the GitHub token secret you create later.

Add the build identity in a new file. The state bucket and the KMS key come from the bootstrap script in module 0, so the stack references them by name instead of creating them.

infra/BuildIdentity.cs
using Pulumi;
using Gcp = Pulumi.Gcp;
namespace Sniplink.Infra;
public static class BuildIdentity
{
public static Gcp.ServiceAccount.Account Create(
string project,
string region,
Gcp.ArtifactRegistry.Repository repo,
Gcp.CloudRunV2.Service service,
Gcp.ServiceAccount.Account runtimeSa,
Gcp.SecretManager.Secret labSecret)
{
var buildSa = new Gcp.ServiceAccount.Account("sniplink-build", new()
{
AccountId = "sniplink-build",
DisplayName = "Sniplink CI/CD (Cloud Build)",
});
var member = buildSa.Email.Apply(e => $"serviceAccount:{e}");
new Gcp.ArtifactRegistry.RepositoryIamMember("build-ar-writer", new()
{
Location = repo.Location,
Repository = repo.Name,
Role = "roles/artifactregistry.writer",
Member = member,
});
new Gcp.CloudRunV2.ServiceIamMember("build-run-developer", new()
{
Location = service.Location,
Name = service.Name,
Role = "roles/run.developer",
Member = member,
});
new Gcp.Projects.IAMMember("build-run-viewer", new()
{
Project = project,
Role = "roles/run.viewer",
Member = member,
});
new Gcp.ServiceAccount.IAMMember("build-actas-runtime", new()
{
ServiceAccountId = runtimeSa.Name,
Role = "roles/iam.serviceAccountUser",
Member = member,
});
new Gcp.Storage.BucketIAMMember("build-state-bucket", new()
{
Bucket = $"{project}-pulumi-state",
Role = "roles/storage.objectAdmin",
Member = member,
});
new Gcp.Kms.CryptoKeyIAMMember("build-state-key", new()
{
CryptoKeyId = $"projects/{project}/locations/{region}/keyRings/pulumi/cryptoKeys/state",
Role = "roles/cloudkms.cryptoKeyEncrypterDecrypter",
Member = member,
});
new Gcp.SecretManager.SecretIamMember("build-labkey-admin", new()
{
SecretId = labSecret.SecretId,
Role = "roles/secretmanager.admin",
Member = member,
});
new Gcp.Projects.IAMMember("build-log-writer", new()
{
Project = project,
Role = "roles/logging.logWriter",
Member = member,
});
return buildSa;
}
}
infra/Program.cs (excerpt)
using Sniplink.Infra;
// …
var gcpConfig = new Config("gcp");
var project = gcpConfig.Require("project");
var region = gcpConfig.Require("region");
// … after repo, service, runtimeSa and labSecret are created
var buildSa = BuildIdentity.Create(project, region, repo, service, runtimeSa, labSecret);

BuildIdentity lives in the Sniplink.Infra namespace and the top-level statements in Program.cs do not, hence the using. The module 5 program only read gcp:region; the build identity also needs the project ID, so the program now reads gcp:project too. Both keys have been in Pulumi.dev.yaml since module 1.

Read the table again with an attacker’s eyes. pulumi up in CI applies whatever the code says. If the code in main adds a binding that grants someone Owner, CI will try to apply it. Least privilege on the build SA is what stops that: sniplink-build has no permission to change project IAM, its own bindings, or the service’s IAM policy, so such a change fails with 403 instead of succeeding.

This works because Pulumi only calls the APIs for resources that changed. A normal commit changes the image, the revision name and a few env vars, so CI touches the Cloud Run service and nothing else. When a commit changes BuildIdentity.cs or the trigger, CI fails and a human with more rights has to apply it from their machine. I consider that failure a feature: changes to who-can-do-what get a second pair of eyes by construction.

The price is that one stack now has two kinds of resources: app resources that CI applies, and platform resources (the build SA, its bindings, the GitHub connection, the trigger) that only a human applies. There is a second, quieter cost. CI holds cryptoKeyEncrypterDecrypter on the key that encrypts every config secret in this stack, so anything you store in Pulumi.dev.yaml is readable by a build.

The pattern that removes both problems is two stacks: a bootstrap or platform stack, owned by an admin group and applied from a protected environment, that holds identities, triggers, the connection and its own KMS key; and an app stack that CI applies, which only receives references (the build SA email, the runtime SA) as stack references. For a course with one student and one service I keep a single stack and accept the trade-off. On a team I would split on day one. Be aware which side of the line you are on.

Connecting GitHub: the second deliberate exception

Section titled “Connecting GitHub: the second deliberate exception”

Cloud Build’s 2nd-gen repositories connect a Cloud Build connection (one per GitHub account or organisation, per region) to repositories inside it. Triggers then point at a repository resource. All of that is Pulumi code, except one step that cannot be: GitHub has to agree to let Google read your repository. That consent happens by installing the Google Cloud Build GitHub App on your account, in GitHub’s UI. This is the second and last exception to the zero-ClickOps rule, and it is acceptable for the same reason as the bootstrap script: it happens once, it is in GitHub rather than in your cloud project, and everything after it is code.

  1. Install the Google Cloud Build app from the GitHub Marketplace on your GitHub account. When GitHub asks which repositories, choose Only select repositories and pick your Sniplink repository. The app should see one repo, not all of them.

  2. Find the installation ID. Open GitHub Settings → Applications → Installed GitHub Apps → Google Cloud Build → Configure. The URL ends with /installations/12345678; that number is the ID.

  3. Create a GitHub personal access token that Cloud Build uses to read the repository. Use a classic token with the repo and read:user scopes (add read:org if the app is installed in an organisation), and give it an expiry date. Cloud Build does not accept fine-grained tokens here.

  4. Store the token and the installation ID in stack config. The token is encrypted with your KMS key, like the lab key in module 3.

Terminal window
cd infra
pulumi config set --secret sniplink:githubToken ghp_xxx # paste your token
pulumi config set sniplink:githubAppInstallationId 12345678
pulumi config set sniplink:githubRepo YOUR_GITHUB_USER/sniplink
pulumi config set sniplink:projectNumber "$(gcloud projects describe "$PROJECT_ID" --format='value(projectNumber)')"

Yes, the token now has an expiry, which means the connection breaks when it expires. That is the right failure: you rotate it with pulumi config set --secret and pulumi up, and Secret Manager gets a new version.

Connection, repository and triggers in code

Section titled “Connection, repository and triggers in code”

The token goes into Secret Manager, because the connection references a secret version, not a raw string. The Cloud Build service agent (service-PROJECT_NUMBER@gcp-sa-cloudbuild.iam.gserviceaccount.com, a Google-managed identity that Cloud Build uses for its own plumbing) reads it. That is a different identity from sniplink-build, which never gets access to the token.

Two triggers point at the same repository:

  • Push to main runs cloudbuild.yaml: test, build, push, deploy.
  • Pull request against main runs cloudbuild.pr.yaml: test and pulumi preview, no deploy.

Both run as sniplink-build, set through ServiceAccount.

infra/GitHubPipeline.cs
using Pulumi;
using Gcp = Pulumi.Gcp;
namespace Sniplink.Infra;
public static class GitHubPipeline
{
public static void Create(string project, string region, Config config,
Gcp.ServiceAccount.Account buildSa)
{
// From config, not a GetProject invoke: invokes run on every `pulumi up`,
// and sniplink-build has no permission to read project metadata.
var projectNumber = config.Require("projectNumber");
var cloudBuildAgent =
$"serviceAccount:service-{projectNumber}@gcp-sa-cloudbuild.iam.gserviceaccount.com";
// 1. GitHub token in Secret Manager
var tokenSecret = new Gcp.SecretManager.Secret("github-token", new()
{
SecretId = "sniplink-github-token",
Replication = new Gcp.SecretManager.Inputs.SecretReplicationArgs
{
Auto = new Gcp.SecretManager.Inputs.SecretReplicationAutoArgs(),
},
});
var tokenVersion = new Gcp.SecretManager.SecretVersion("github-token-v", new()
{
Secret = tokenSecret.Id,
SecretData = config.RequireSecret("githubToken"),
});
var agentAccess = new Gcp.SecretManager.SecretIamMember("cb-agent-token", new()
{
SecretId = tokenSecret.SecretId,
Role = "roles/secretmanager.secretAccessor",
Member = cloudBuildAgent,
});
// 2. Connection (one per GitHub account, per region)
var connection = new Gcp.CloudBuildV2.Connection("github", new()
{
Location = region,
Name = "github",
GithubConfig = new Gcp.CloudBuildV2.Inputs.ConnectionGithubConfigArgs
{
AppInstallationId = config.RequireInt32("githubAppInstallationId"),
AuthorizerCredential = new Gcp.CloudBuildV2.Inputs.ConnectionGithubConfigAuthorizerCredentialArgs
{
OauthTokenSecretVersion = tokenVersion.Id,
},
},
}, new CustomResourceOptions { DependsOn = { agentAccess } });
// 3. Repository inside the connection
var githubRepo = config.Require("githubRepo"); // "owner/sniplink"
var repository = new Gcp.CloudBuildV2.Repository("sniplink", new()
{
Location = region,
Name = "sniplink",
ParentConnection = connection.Name,
RemoteUri = $"https://github.com/{githubRepo}.git",
});
var substitutions = new InputMap<string>
{
["_REGION"] = region,
};
// 4a. Push to main → deploy
new Gcp.CloudBuild.Trigger("deploy-main", new()
{
Location = region,
Name = "sniplink-deploy-main",
RepositoryEventConfig = new Gcp.CloudBuild.Inputs.TriggerRepositoryEventConfigArgs
{
Repository = repository.Id,
Push = new Gcp.CloudBuild.Inputs.TriggerRepositoryEventConfigPushArgs
{
Branch = "^main$",
},
},
Filename = "cloudbuild.yaml",
ServiceAccount = buildSa.Id, // projects/PROJECT/serviceAccounts/EMAIL
Substitutions = substitutions,
});
// 4b. Pull request → test + preview
new Gcp.CloudBuild.Trigger("preview-pr", new()
{
Location = region,
Name = "sniplink-preview-pr",
RepositoryEventConfig = new Gcp.CloudBuild.Inputs.TriggerRepositoryEventConfigArgs
{
Repository = repository.Id,
PullRequest = new Gcp.CloudBuild.Inputs.TriggerRepositoryEventConfigPullRequestArgs
{
Branch = "^main$",
// Forks need a maintainer's /gcbrun comment before a build runs
CommentControl = "COMMENTS_ENABLED_FOR_EXTERNAL_CONTRIBUTORS_ONLY",
},
},
Filename = "cloudbuild.pr.yaml",
ServiceAccount = buildSa.Id,
Substitutions = substitutions,
});
}
}
infra/Program.cs
// …
var buildSa = BuildIdentity.Create(project, region, repo, service, runtimeSa, labSecret);
GitHubPipeline.Create(project, region, config, buildSa);

A few values deserve a sentence each:

  • Location = region on the connection and triggers: 2nd-gen repositories are regional, and a trigger must live in the same region as its connection.
  • Branch = "^main$" is a regular expression. Without the anchors, main also matches not-main-yet.
  • CommentControl matters because your repository is public. A pull request from a fork brings its own cloudbuild.pr.yaml, and that file runs as sniplink-build, which can decrypt your stack secrets. With this setting, a fork’s PR builds only after an owner or collaborator comments /gcbrun. Read the diff before you type it.
  • ServiceAccount takes the full resource name projects/PROJECT_ID/serviceAccounts/EMAIL, which is what Account.Id returns.

Apply this once from your laptop. You are an Owner of the project, so you have the iam.serviceAccounts.actAs permission that creating a trigger with a custom service account requires.

Terminal window
cd infra
pulumi up

At this point the service code is still the module 5 version, and the stack config still holds your module 5 values, so the preview must show only creates (build SA, bindings, token secret, connection, repository, two triggers) and no change to sniplink-api. If it wants to update the service, stop and find out why before you continue.

Cloud Build reads a YAML file with a list of steps. Each step is a container image plus arguments. Steps run in order by default, share the /workspace directory (your checked-out repo), and share the Docker daemon, so an image built in one step is visible in the next.

Before the file, the Pulumi program has to change in three ways: revision names come from the build, the service learns which commit and build produced it, and the stack exports what it deployed.

From hand-named releases to one revision per build

Section titled “From hand-named releases to one revision per build”

In module 5 you named every revision yourself (sniplink-api-v1, sniplink-api-v2) through sniplink:release, and drove traffic with three config keys. CI keeps the second half and replaces the first:

  • Revision name. Every build creates a revision named sniplink-api-<shortSha>, for example sniplink-api-3f9c2ab. The build passes $SHORT_SHA as sniplink:revisionSuffix, so you never type a name, two builds never collide, and every revision name points at a commit. Short SHAs are hex, so they can never clash with the v1/v2 names from module 5, which stay in the revision list and remain valid rollback targets.
  • Traffic. Still stack config, committed in git, changed by pull request. One new key, sniplink:trafficMode, chooses between two shapes:
    • promote: the revision this build creates gets 100 %. This is the everyday mode.
    • canary: sniplink:stableRevision (a name you commit) keeps 100 - canaryPercent %, and the revision this build creates gets canaryPercent % and the canary tag.
Key Set by Example Meaning
sniplink:image CI, per build …/api@sha256:… Image digest for this build’s revision
sniplink:revisionSuffix CI, per build 3f9c2ab Revision name becomes sniplink-api-3f9c2ab
sniplink:commitSha, sniplink:buildId CI, per build full SHA, UUID COMMIT_SHA and BUILD_ID env vars
sniplink:version you, by PR 2.1.0 APP_VERSION env var
sniplink:trafficMode you, by PR promote promote or canary
sniplink:stableRevision you, by PR sniplink-api-3f9c2ab Keeps the rest of the traffic in canary mode
sniplink:canaryPercent you, by PR 0 Share for this build’s revision in canary mode
sniplink:studentId you, once (module 1) cd_… LABKIT_STUDENT_ID env var

A release now looks like this. Merge a change with trafficMode: canary, stableRevision set to the revision that serves today, and canaryPercent: 0; the build deploys the new code behind the canary tag. Raise canaryPercent by PR. Promote by PR with trafficMode: promote. Roll back by PR with trafficMode: canary, stableRevision set to the old revision and canaryPercent: 0. Each of those merges runs a build, and the build gets a new revision name even when only config changed: the revision behind the tag is renamed, its code is the same. Routing changes the moment pulumi up runs; the minutes before that are the build. For an emergency, the gcloud run services update-traffic escape hatch from module 5 still works, followed by the PR that makes git agree.

Switch the committed config over once, before you push the new program:

Terminal window
cd infra
pulumi config rm sniplink:release # replaced by revisionSuffix from CI
pulumi config rm sniplink:canaryRevision # in canary mode the canary is "this build"
pulumi config rm sniplink:image # CI passes the digest on every run
pulumi config set sniplink:trafficMode promote
pulumi config set sniplink:canaryPercent 0
# sniplink:version and sniplink:stableRevision stay; from now on you change them by PR

Removing sniplink:image is deliberate: a pulumi up from your laptop now fails with a missing-config error instead of quietly redeploying an old image.

sniplink:studentId from module 1 stays as well. It is committed stack config and not a secret, so the build reads it from Pulumi.dev.yaml like any other key, and cloudbuild.yaml needs nothing extra for LABKIT_STUDENT_ID.

infra/Program.cs (excerpt)
using Pulumi.Gcp.CloudRunV2; // same usings as in module 5
using Pulumi.Gcp.CloudRunV2.Inputs;
using Sniplink.Infra;
// …
// config is Config("sniplink"); repo, runtimeSa, labSecret, labKeyAccess,
// labKeyVersion from modules 1 and 3
const string serviceName = "sniplink-api";
var gcpConfig = new Config("gcp");
var project = gcpConfig.Require("project");
var region = gcpConfig.Require("region");
var image = config.Require("image"); // repo@sha256:… from CI
var revisionSuffix = config.Require("revisionSuffix"); // $SHORT_SHA from CI
var version = config.Require("version"); // APP_VERSION, committed
var commitSha = config.Get("commitSha") ?? "local";
var buildId = config.Get("buildId") ?? "local";
var trafficMode = config.Get("trafficMode") ?? "promote";
var stableRevision = config.Get("stableRevision");
var canaryPercent = config.GetInt32("canaryPercent") ?? 0;
var studentId = config.Require("studentId");
var revisionName = $"{serviceName}-{revisionSuffix}";
// Fail at preview time, not halfway through an update.
if (trafficMode is not ("promote" or "canary"))
throw new ArgumentException("trafficMode must be 'promote' or 'canary'");
if (canaryPercent is < 0 or > 100)
throw new ArgumentException("canaryPercent must be 0..100");
if (trafficMode == "canary" && (stableRevision is null || stableRevision == revisionName))
throw new ArgumentException("canary mode needs a stableRevision other than this build's revision");
var traffics = new InputList<ServiceTrafficArgs>();
if (trafficMode == "promote")
{
traffics.Add(new ServiceTrafficArgs
{
Type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION",
Revision = revisionName,
Percent = 100,
});
}
else
{
traffics.Add(new ServiceTrafficArgs
{
Type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION",
Revision = stableRevision!,
Percent = 100 - canaryPercent,
});
traffics.Add(new ServiceTrafficArgs
{
Type = "TRAFFIC_TARGET_ALLOCATION_TYPE_REVISION",
Revision = revisionName,
Percent = canaryPercent,
Tag = "canary",
});
}
var service = new Service(serviceName, new ServiceArgs
{
Name = serviceName,
Location = region,
DeletionProtection = false,
Ingress = "INGRESS_TRAFFIC_ALL",
Template = new ServiceTemplateArgs
{
Revision = revisionName,
ServiceAccount = runtimeSa.Email,
// … MaxInstanceRequestConcurrency, Scaling, Volumes from modules 3 and 4
Containers =
{
new ServiceTemplateContainerArgs
{
Image = image,
Envs =
{
new ServiceTemplateContainerEnvArgs { Name = "LABKIT_STUDENT_ID", Value = studentId },
new ServiceTemplateContainerEnvArgs { Name = "APP_VERSION", Value = version },
new ServiceTemplateContainerEnvArgs { Name = "COMMIT_SHA", Value = commitSha },
new ServiceTemplateContainerEnvArgs { Name = "BUILD_ID", Value = buildId },
},
// … ports, resources, probes and volume mounts from modules 1, 3 and 4
},
},
},
Traffics = traffics,
}, new CustomResourceOptions { DependsOn = { repo, labKeyAccess, labKeyVersion } });
// … public invoker binding from module 1
var buildSa = BuildIdentity.Create(project, region, repo, service, runtimeSa, labSecret);
GitHubPipeline.Create(project, region, config, buildSa);
return new Dictionary<string, object?>
{
["url"] = service.Uri,
["canaryUrl"] = service.TrafficStatuses.Apply(statuses =>
statuses.FirstOrDefault(s => s.Tag == "canary")?.Uri ?? ""),
["deployedRevision"] = revisionName,
["deployedImage"] = image,
["deployedRevisionSuffix"] = revisionSuffix,
["deployedCommitSha"] = commitSha,
["deployedBuildId"] = buildId,
};

What changed compared with module 5, line by line:

  • Revision = $"{serviceName}-{revisionSuffix}". Same prefix rule as module 5, but the suffix now comes from the build. The release key is gone.
  • trafficMode replaces canaryRevision. In module 5 the canary was a revision you named. In CI the canary is always “the revision this build creates”, so there is nothing to name. Only the stable side is named in config, because it is the one revision you deliberately keep.
  • Both targets stay pinned by name. Nothing points at “latest”, so a build only gets traffic because the config committed with it says so.
  • COMMIT_SHA and BUILD_ID. LabKit’s /_lab/info reads them from the environment. Locally they fall back to local, so a laptop deploy is visibly not a CI deploy.
  • The deployed* outputs let the PR pipeline (below) and your laptop preview against what CI actually deployed.

For the lab, keep trafficMode: promote, so the main URL serves the revision of the commit you just pushed. The final project uses canary mode with exactly this program.

Commit this file at the repository root.

cloudbuild.yaml
# Runs on push to main as sniplink-build. No keys anywhere.
substitutions:
_REGION: europe-west1
_PULUMI_IMAGE: pulumi/pulumi-dotnet-10.0:3.267.0 # pinned, see the note below
steps:
# 1. Unit tests. A red test stops the build here.
- id: test
name: mcr.microsoft.com/dotnet/sdk:10.0
entrypoint: dotnet
args: ['test', 'Sniplink.sln', '--configuration', 'Release']
# 2. Build the image, tagged with the short commit SHA (readable in the registry)
- id: build
name: gcr.io/cloud-builders/docker
env: ['DOCKER_BUILDKIT=1']
args:
- build
- --tag=$_REGION-docker.pkg.dev/$PROJECT_ID/sniplink/api:$SHORT_SHA
- --file=Dockerfile
- .
# 3. Push to Artifact Registry as sniplink-build
- id: push
name: gcr.io/cloud-builders/docker
args: ['push', '$_REGION-docker.pkg.dev/$PROJECT_ID/sniplink/api:$SHORT_SHA']
# 4. Resolve the immutable digest; tags can be moved, digests cannot
- id: digest
name: gcr.io/cloud-builders/docker
entrypoint: bash
args:
- -c
- |
set -euo pipefail
docker inspect --format='{{index .RepoDigests 0}}' \
$_REGION-docker.pkg.dev/$PROJECT_ID/sniplink/api:$SHORT_SHA > /workspace/image-ref.txt
echo "Deploying $$(cat /workspace/image-ref.txt)"
# 5. Deploy: the same Pulumi program you ran locally, now with CI's identity
- id: deploy
name: $_PULUMI_IMAGE
dir: infra
entrypoint: bash
args:
- -c
- |
set -euo pipefail
pulumi login gs://$PROJECT_ID-pulumi-state
pulumi up --yes --non-interactive --stack dev \
--config sniplink:image="$$(cat /workspace/image-ref.txt)" \
--config sniplink:revisionSuffix=$SHORT_SHA \
--config sniplink:commitSha=$COMMIT_SHA \
--config sniplink:buildId=$BUILD_ID
options:
# Required when a build runs as a user-specified service account
# (alternatively: a logs bucket you own).
logging: CLOUD_LOGGING_ONLY
timeout: 1200s

Walk through it once:

  • Substitutions. Cloud Build replaces $NAME and ${NAME} in the file before running anything. PROJECT_ID, BUILD_ID, COMMIT_SHA and SHORT_SHA are built in; COMMIT_SHA and SHORT_SHA are filled only for builds started by a trigger, not for a manual gcloud builds submit. Names starting with _ are yours. The substitutions: block gives defaults, and the trigger’s Substitutions override them, which is how the region in your stack config reaches the build.
  • $$. Because Cloud Build owns $, a dollar sign meant for bash must be written $$. $$(cat …) reaches bash as $(cat …). Forget this and Cloud Build rejects the file for an unknown substitution, which is at least a loud failure.
  • Step 1 runs tests in the same SDK image your Dockerfile uses. The bin/ and obj/ folders it leaves in /workspace are kept out of step 2 by the .dockerignore from module 2.
  • Steps 2-3 build and push with the tag $SHORT_SHA, so a human looking at the registry can match images to commits. DOCKER_BUILDKIT=1 matters: the cloud-builders/docker image still defaults to the legacy builder, which does not set $BUILDPLATFORM and $TARGETARCH, so the module 2 Dockerfile fails at its first line with failed to parse platform.
  • Step 4 resolves the digest. RepoDigests is filled in only after a push. Deploying api@sha256:… instead of api:abc1234 means the revision runs exactly the bytes that were tested, even if someone later pushes a different image under the same tag.
  • Step 5 runs Pulumi. Both the GCS state backend and the Google provider use Application Default Credentials, and in Cloud Build those come from the metadata server as sniplink-build. Nothing to configure. --config writes the four per-build values into the build’s copy of Pulumi.dev.yaml for this run only; your committed file is not changed. Traffic comes from the committed keys, so what a build does to routing is whatever the merged PR said.

Steps 1 and 2 are independent, so you could start both at once with waitFor: ['-'] on each and waitFor: ['test', 'build'] on the push. I keep the file sequential in the course because it is easier to read in logs, and a failed test then saves the image build time.

A PR build answers two questions: do the tests pass, and what would this change do to production? It never deploys.

cloudbuild.pr.yaml
# Runs on pull requests to main as sniplink-build. Never deploys.
substitutions:
_PULUMI_IMAGE: pulumi/pulumi-dotnet-10.0:3.267.0 # same pinned tag as cloudbuild.yaml
steps:
- id: test
name: mcr.microsoft.com/dotnet/sdk:10.0
entrypoint: dotnet
args: ['test', 'Sniplink.sln', '--configuration', 'Release']
- id: preview
name: $_PULUMI_IMAGE
dir: infra
entrypoint: bash
args:
- -c
- |
set -euo pipefail
pulumi login gs://$PROJECT_ID-pulumi-state
# Compare against what CI deployed, not against the committed config,
# otherwise every preview shows the image and env vars as "changed".
pulumi preview --non-interactive --diff --stack dev \
--config sniplink:image="$$(pulumi stack output deployedImage --stack dev)" \
--config sniplink:revisionSuffix="$$(pulumi stack output deployedRevisionSuffix --stack dev)" \
--config sniplink:commitSha="$$(pulumi stack output deployedCommitSha --stack dev)" \
--config sniplink:buildId="$$(pulumi stack output deployedBuildId --stack dev)"
options:
logging: CLOUD_LOGGING_ONLY

The preview output in the build log is what a reviewer reads next to the code diff. If a PR that “only renames a variable” shows a Cloud Run replacement, you find out before merging, not after. For a traffic PR (say canaryPercent 0 → 10) the preview shows the traffic change against the revision that is live now; after the merge, the build applies the same split with its own, newly named revision in the canary slot.

Your laptop uses the same trick. To preview locally after module 6, pass the deployed values exactly as cloudbuild.pr.yaml does.

I use the same service account for both triggers to keep the module short. A stricter setup gives PR builds their own SA with read-only roles (state bucket objectViewer, KMS decrypter, and viewer roles on the resources), so a malicious PR cannot write anything even after a careless /gcbrun. Read access to the state bucket is enough because pulumi preview takes no lock on a self-managed backend such as a GCS bucket; pulumi up writes a lock file and needs write access.

Commit both YAML files and the infra changes, push to main, and watch the build:

Terminal window
git add cloudbuild.yaml cloudbuild.pr.yaml infra/
git commit -m "Deploy from Cloud Build"
git push origin main
# Latest builds in your region, then stream one build's log
gcloud builds list --region="$REGION" --limit=5
gcloud builds log BUILD_ID --region="$REGION" --stream

The logs are in Cloud Logging as well, because of CLOUD_LOGGING_ONLY, so the read-only console view is Cloud Build → History for your region. When the build is green:

Terminal window
curl -s "$(cd infra && pulumi stack output url)/_lab/info" | jq '{commitSha, buildId, revision, version}'

commitSha should be the full SHA of your commit, buildId a UUID, and revision sniplink-api- followed by the short SHA.

GitHub Actions is the CI most .NET teams already use. The pipeline is the same; what changes is how the job becomes sniplink-build. Here GitHub signs an OIDC token for each job, and Google Cloud trusts it through Workload Identity Federation.

Nothing more to do. The trigger runs cloudbuild.yaml as sniplink-build, and the identity comes from the metadata server inside your project. The build definition is in your repo, the trigger and its identity are in your Pulumi stack, and the logs are in your project’s Cloud Logging.

The identity side: pool, provider, binding

Section titled “The identity side: pool, provider, binding”

Three resources make the GitHub token acceptable to Google Cloud:

  • A workload identity pool: a container for external identities in your project.
  • A provider in that pool: “trust tokens issued by https://token.actions.githubusercontent.com”, plus an attribute mapping that copies claims from GitHub’s token into Google attributes, and an attribute condition that rejects every token not from your repository. Without the condition, any GitHub repository in the world could present a valid token to your pool; Google now requires a condition for GitHub providers for that reason.
  • A binding of roles/iam.workloadIdentityUser on sniplink-build to the principal set “all identities in this pool whose repository attribute is owner/sniplink”. That is what lets a job impersonate the SA.

The Security Token Service API (sts.googleapis.com) does the token exchange. It was not in the bootstrap list, so the stack enables it.

infra/GitHubActionsIdentity.cs
using Pulumi;
using Gcp = Pulumi.Gcp;
namespace Sniplink.Infra;
public static class GitHubActionsIdentity
{
public static void Create(string githubRepo, Gcp.ServiceAccount.Account buildSa)
{
var sts = new Gcp.Projects.Service("sts", new()
{
ServiceName = "sts.googleapis.com",
DisableOnDestroy = false,
});
var pool = new Gcp.Iam.WorkloadIdentityPool("github", new()
{
WorkloadIdentityPoolId = "github",
DisplayName = "GitHub Actions",
}, new CustomResourceOptions { DependsOn = { sts } });
new Gcp.Iam.WorkloadIdentityPoolProvider("github-oidc", new()
{
WorkloadIdentityPoolId = pool.WorkloadIdentityPoolId,
WorkloadIdentityPoolProviderId = "github-oidc",
DisplayName = "GitHub OIDC",
Oidc = new Gcp.Iam.Inputs.WorkloadIdentityPoolProviderOidcArgs
{
IssuerUri = "https://token.actions.githubusercontent.com",
},
AttributeMapping =
{
["google.subject"] = "assertion.sub",
["attribute.repository"] = "assertion.repository",
["attribute.ref"] = "assertion.ref",
},
// Only tokens from this repository are accepted at all
AttributeCondition = $"assertion.repository == '{githubRepo}'",
});
// pool.Name = projects/NUMBER/locations/global/workloadIdentityPools/github
new Gcp.ServiceAccount.IAMMember("github-impersonates-build", new()
{
ServiceAccountId = buildSa.Name,
Role = "roles/iam.workloadIdentityUser",
Member = pool.Name.Apply(n =>
$"principalSet://iam.googleapis.com/{n}/attribute.repository/{githubRepo}"),
});
}
}

Two refinements I use at work. Condition on assertion.repository_id rather than the name, because a deleted and re-created repository (or a renamed one) can reuse the name, but never the numeric ID. And for deploys, add assertion.ref == 'refs/heads/main' to the condition (or a separate provider), so a job on a feature branch cannot impersonate the deploy SA at all.

Cloud Build GitHub Actions + WIF
Identity SA attached to the build; nothing to federate OIDC federation: pool, provider, condition to get right
Where things live Trigger and identity in your project and your IaC; logs in Cloud Logging Workflow in the repo; logs and run history in GitHub
Ecosystem Any container as a step; no marketplace Huge marketplace of actions, matrix builds, environments with approvals
Developer experience Logs and status reach the PR via the GitHub app, but the UI is the Cloud console PR checks, annotations and re-runs where developers already are
Private networking Private pools can reach resources in your VPC Needs self-hosted runners for that
Free tier See below Free on standard runners for public repos; 2,000 minutes a month for private repos on GitHub Free

My honest take: if your team lives in GitHub and does not need VPC access from CI, Actions with WIF is a fine default and the extra identity setup is a one-time cost. Cloud Build wins when you want the build identity, its logs and its audit trail inside the project, private pools, or zero third-party trust. The skill this module teaches, a keyless identity with a short list of scoped roles, is the same in both.

Cloud Build’s free tier is 2,500 build-minutes per month per billing account, on the default e2-standard-2 machine type in the default pool. Beyond that, the same machine costs $0.006 per build-minute. Google states that the free tier is promotional and may change, so check the pricing page before you rely on it.

A Sniplink build takes a few minutes: tests, an image build, and a dotnet build of the infra project before pulumi up. At around six minutes per push you get roughly 400 builds a month inside the free tier, which a course project will not come near. The other costs from this module are cents: one more secret (the GitHub token) and an image in Artifact Registry per commit. That last one grows quietly; an Artifact Registry cleanup policy that keeps the last N versions is worth adding once you are past the course.

Logs go only to Cloud Logging (CLOUD_LOGGING_ONLY), which has its own free ingestion allowance. Treat build logs as you would treat the job log in the incident: anyone who can read them sees everything a step prints. The pipeline above never prints a secret, and pulumi up masks secret config values in its output.

Goal: a push to main on your public GitHub repository is tested, built and deployed by Cloud Build as sniplink-build, and the running service can prove which commit and build it came from.

Requirements:

  1. Your Sniplink repository is public on GitHub and contains cloudbuild.yaml and cloudbuild.pr.yaml at the root.
  2. The build SA, connection, repository and both triggers are created by Pulumi, with the roles from this module and nothing broader.
  3. The service’s current revision was deployed by the push trigger with trafficMode: promote, so the main URL serves sniplink-api-<shortSha>, COMMIT_SHA is the full 40-character SHA and BUILD_ID is the Cloud Build UUID.
  4. No user-managed service-account keys exist in the project.

Check your lab

Loading your lab…

Sign in to verify this lab and save your progress. Everything above works without an account.

  • Your Cloud Run service URL (pulumi stack output url)
  • Your public GitHub repository URL (https://github.com/<owner>/<repo>)

5 checks run against the service you deployed.

Some checks are self-reported: the checker trusts what your service says about itself and verifies the rest from the outside.

What each check verifies:

  • Token and health: the baseline from every lab. The service runs on Cloud Run in your project.
  • Commit SHA: /_lab/info reports a commitSha of 40 lowercase hex characters. Self-reported: the service returns what its env says.
  • Build ID: buildId is a UUID, the format Cloud Build uses. Self-reported as well.
  • GitHub commit: the platform asks GitHub’s public API whether that commit exists in the repository you entered, and whether cloudbuild.yaml exists at that commit. This ties the running revision to real code in a real repo.

Be clear about the limit: a determined student could set those env vars by hand from a laptop and pass. A stricter check would read Cloud Build history and the service’s IAM directly, and that is what the optional Verified tier will do later. The self-check below cannot be automated through the URL at all, so it is on you.

Self-check: no keys. List user-managed keys for every service account in the project. Google-managed keys are rotated by Google and are not shown with this filter. Empty output is the pass.

Terminal window
for sa in $(gcloud iam service-accounts list --project="$PROJECT_ID" --format='value(email)'); do
gcloud iam service-accounts keys list \
--iam-account="$sa" --managed-by=user --format='value(name)'
done

If anything prints, delete it with gcloud iam service-accounts keys delete KEY_ID --iam-account=SA_EMAIL, then find out why it existed.

The final project in module 7 is a new service, but it reuses the same patterns: a build SA, a GitHub connection, a trigger. You have two options.

Keep everything (recommended if you continue soon). The connection and triggers cost nothing while idle, and the module 7 pipeline can reuse the connection. Cloud Run scales to zero. The only steady costs are cents for the KMS key version, the secrets and stored images.

Destroy the stack (if you are pausing). Run it from your laptop, because only you have the rights to delete IAM bindings and the trigger:

Terminal window
cd infra
pulumi destroy

This removes the service, the registry with its images, the secrets, the build SA, the WIF pool (if you created it), the connection and both triggers. pulumi destroy needs the same per-build config as a preview, so pass the deployed values with --config as in cloudbuild.pr.yaml if it complains about sniplink:image. Two things stay and are yours to remove by hand: the Google Cloud Build app installation on GitHub (Settings → Applications), and the GitHub token, which you should revoke in Settings → Developer settings. Deleted workload identity pools stay in a soft-deleted state for 30 days and block the same pool ID until then.

  • A service-account key is a non-expiring, copyable credential; keyless CI uses an attached identity (Cloud Build) or federation (Workload Identity Federation), and the iam.disableServiceAccountKeyCreation policy makes keys impossible by default.
  • A dedicated build SA gets each role on the narrowest resource: repo, service, runtime SA, bucket, key, secret. When CI manages its own IAM, least privilege is what turns a dangerous commit into a failed build, and a separate platform stack is the clean long-term split.
  • A 2nd-gen GitHub connection needs one GitHub app install; the token secret, connection, repository and triggers are code, and PR builds from forks need a comment gate.
  • cloudbuild.yaml tests, builds, pushes and deploys by digest, naming each revision after its commit and passing COMMIT_SHA and BUILD_ID through Pulumi config into the service, while PR builds only pulumi preview. Traffic stays in committed config, so promotion and rollback are pull requests.
  • GitHub Actions does the same job with an OIDC pool, a provider with an attribute condition, and a workloadIdentityUser binding scoped to one repository.