Skip to content

Signed in as

Sign in to Curious Dev Learn

Sign in to verify labs, save your progress and get a certificate. Lessons stay open without an account.

Module 032 h 15 minLab

Config, secrets and a service identity

Move a secret out of your code and into Secret Manager, read it at runtime from a mounted volume, and run Sniplink as a dedicated service account that can read exactly one secret.

Two separate mistakes are stacked in that log. The first is a secret living in a file that is versioned, copied and forked like any other source file. Deleting it in the next commit does nothing: the key is in git history, in every clone and in the fork that a stranger now owns. The only fix is to rotate the key, and the team discovers they have no idea how, because nobody ever has.

The second mistake is quieter and worse. The service ran as the default compute service account, which in many projects holds roles/editor on everything. Anyone who gets code execution in that container, or steals a token from it, can delete databases, read every bucket and create VMs to mine crypto. The leaked API key was one secret. The identity was the keys to the building.

This module answers two questions. Where does a secret live so that it never touches your image, your env files or git, and can be rotated without a deploy? And which identity should Sniplink run as so that a compromise of the service is a compromise of almost nothing? By the end, Sniplink reads a lab key from Secret Manager at runtime, runs as sniplink-runtime with permission to read that one secret and nothing else, and you will prove both from outside.

Terminal window
export PROJECT_ID=your-course-project-id
export REGION=europe-west1

Configuration in .NET, and where secrets don’t go

Section titled “Configuration in .NET, and where secrets don’t go”

ASP.NET Core builds one IConfiguration from a stack of sources. With the default WebApplication.CreateBuilder, the order is roughly: appsettings.json, appsettings.{Environment}.json, user secrets (Development only), environment variables, then command-line arguments. Later sources override earlier ones, and __ in an environment variable name maps to : in a key, so LabSecret__Path becomes LabSecret:Path.

That model is good for configuration: log levels, feature flags, URLs of other services, timeouts. Values that are fine to read in a pull request. A secret is different. My working definition: a value that grants access to something, so that anyone who reads it can act as you. API keys, database passwords, signing keys, webhook secrets and OAuth client secrets all qualify. A connection string with a password in it is a secret even though it looks like configuration.

Here is why each common hiding place fails:

  • git (appsettings.json, .env committed by accident). History is forever, forks are public, and secret scanners run by both defenders and attackers find keys within minutes of a push.
  • The container image. Anything you COPY in or bake as ENV is readable by anyone who can pull the image, and layers keep deleted files. Your registry permissions become your secret permissions, and they are usually much wider.
  • Env files and plain environment variables in your deploy config. Better than the image, but the value is now in your IaC code or CI variables, visible in the Cloud Run console to anyone with viewer access, and printed by the first debug endpoint that dumps the environment.
  • User secrets (dotnet user-secrets). Fine for local development, which is exactly what they are for. They are a JSON file in your home directory, not a production store.

What you want instead is a store that holds the secret encrypted, controls who may read it per secret, logs every access, keeps versions so rotation is a first-class operation, and hands the value to your running service without it passing through git, the image or the deploy config. On Google Cloud that store is Secret Manager.

  1. Secret. A named container, for example sniplink-lab-key. It holds metadata (labels, replication policy, IAM policy) but no value by itself.

  2. Version. The actual bytes. Versions are numbered 1, 2, 3 and immutable: you never edit a value, you add a new version. Each version is ENABLED, DISABLED or DESTROYED. Disabled can be re-enabled; destroyed cannot.

  3. latest. An alias that resolves to the most recently created version. Consumers that ask for latest follow rotations automatically. Consumers that ask for 3 stay on 3 until someone changes their config.

  4. IAM per secret. You can grant roles/secretmanager.secretAccessor on the project (every secret) or on a single secret. This course always binds on the single secret. The accessor role can read values; it cannot list other secrets, change IAM or add versions.

  5. Replication. Automatic replication lets Google choose where to store the encrypted data. User-managed replication pins it to regions you pick, which matters if you have data residency rules. For Sniplink, automatic is the right choice.

Cost: you pay per active secret version per month and per access operation, with a small free tier. For this module the bill is a few cents at most; check current pricing on the Secret Manager pricing page before you assume that for your production workload, especially if a hot path reads a secret on every request (which is one more reason to cache, see below).

Pulumi config secrets: getting the value in without writing it down

Section titled “Pulumi config secrets: getting the value in without writing it down”

You need the lab key in Secret Manager, created by Pulumi, without the key ever appearing in infra/Program.cs or in a plain-text config file. Pulumi has exactly the tool for this: config secrets. In module 0 you created a Cloud KMS key and configured the stack with the gcpkms:// secrets provider. Every value you set with --secret is encrypted with that KMS key before it is written to Pulumi.<stack>.yaml, and stays encrypted in the state bucket.

Open the lab page for m03-secrets on the course site. It shows a lab key generated for you. Copy it, then:

Terminal window
cd infra
# Omitting the value makes Pulumi prompt for it, so the key never lands in shell history
pulumi config set --secret sniplink:labKey
# Paste the lab key at the prompt

Look at the stack file. You will see something like sniplink:labKey: secure: v1:AAAA.... That ciphertext is safe to commit: decrypting it requires cloudkms.cryptoKeyVersions.useToDecrypt on your KMS key. When Pulumi runs, it decrypts the value in memory and treats it as a secret output, so it is masked in previews, logs and stack outputs.

Now create the secret and its first version. Add this to your Pulumi program, next to the Artifact Registry repository from module 1.

infra/Program.cs (excerpt)
// … config is the Config("sniplink") from module 1
var labKeyValue = config.RequireSecret("labKey"); // Output<string>, marked secret
var labSecret = new Gcp.SecretManager.Secret("lab-key", new()
{
SecretId = "sniplink-lab-key",
Replication = new Gcp.SecretManager.Inputs.SecretReplicationArgs
{
Auto = new Gcp.SecretManager.Inputs.SecretReplicationAutoArgs(),
},
});
var labKeyVersion = new Gcp.SecretManager.SecretVersion("lab-key-version", new()
{
Secret = labSecret.Id, // full name: projects/…/secrets/sniplink-lab-key
SecretData = labKeyValue,
// When the value changes, Pulumi replaces this resource.
// DISABLE keeps the old version recoverable instead of destroying it.
DeletionPolicy = "DISABLE",
});
// …

An honest note on the Terraform tab. sensitive = true hides the value in plan output, but the value is still stored in plain text in Terraform state, so your state backend becomes a secret store and needs the same protection. Recent Terraform and Google provider versions add write-only arguments that avoid storing the value in state; check the provider docs for your version. Pulumi encrypts secret values in state with your KMS key by default, which is one of the reasons I use it here.

Every Cloud Run revision runs as a service account. Code in the container gets that account’s credentials from the metadata server, and every Google client library uses them automatically. If you set nothing, Cloud Run uses the default compute service account, PROJECT_NUMBER-compute@developer.gserviceaccount.com.

The problem with the default account is that it is shared and over-privileged. It is shared because every Cloud Run service, Compute Engine VM and some other products in the project use it unless told otherwise, so you cannot grant it something for one service without granting it to all. It is over-privileged because, in many projects, it was automatically granted roles/editor on the project when the Compute API was enabled. Newer organisations enforce a policy that stops this automatic grant, so your course project may not have it. Check before you relax: either way, a shared identity makes least privilege impossible.

The fix is one service account per workload, granted only what that workload needs. Sniplink needs exactly one thing from Google Cloud: read the lab key. So:

infra/Program.cs (excerpt)
// …
var runtimeSa = new Gcp.ServiceAccount.Account("sniplink-runtime", new()
{
AccountId = "sniplink-runtime",
DisplayName = "Sniplink API runtime identity",
});
// Resource-level binding: this SA can read this one secret, nothing else.
var labKeyAccess = new Gcp.SecretManager.SecretIamMember("runtime-reads-lab-key", new()
{
SecretId = labSecret.SecretId,
Role = "roles/secretmanager.secretAccessor",
Member = runtimeSa.Email.Apply(email => $"serviceAccount:{email}"),
});
// …

Use SecretIamMember, not SecretIamBinding or SecretIamPolicy. A member resource adds one principal to one role and leaves everything else alone. A binding resource is authoritative for the role on that secret, and a policy resource is authoritative for the whole IAM policy of the secret; both will happily remove grants that something else created. Additive is the safe default.

Notice what the runtime account does not get. It does not need a role to write logs: Cloud Run collects stdout and stderr itself. It does not need Artifact Registry access: the image is pulled by the Cloud Run service agent, a Google-managed account, not by your runtime identity. It does not need anything to mint the ID token LabKit uses for /_lab/verify: a service account can always get identity tokens for itself from the metadata server. One role, on one resource.

Cloud Run can expose a Secret Manager secret to your container in two ways. Both are configured on the service, and both require the runtime service account to have secretAccessor on the secret.

As an environment variable. Cloud Run reads the secret version when an instance starts and injects the value as an env var. Your code reads it like any other config value. The value never changes for the life of that instance.

As a volume. Cloud Run mounts the secret as a file in a directory you choose. The value is fetched from Secret Manager when your code reads the file. If the volume references latest, a new secret version becomes visible on the next read, in the same revision, on the same instance.

Environment variable Volume mount
When the value is read Once, at instance start When your code reads the file
New version with latest Only instances started after the change see it; running ones keep the old value, so one revision can serve two values at once Next file read sees it, on all instances
Recommended version reference A pinned number (3), changed by a new revision latest is fine
Rotation needs a new revision Yes, to be consistent No
Code changes None: it is just IConfiguration Read a file (or add a key-per-file config source)
Failure mode if access is missing Instance fails to start Read fails at runtime; startup may succeed
Leaks via env dumps and crash reports Yes, it is in the process environment No

Google’s own guidance is to pin a specific version when you use an env var, precisely because latest makes the value depend on when an instance happened to start. That makes env vars predictable but turns every rotation into a deploy.

I use a volume for Sniplink for three reasons. Rotation becomes a Secret Manager operation instead of a release. The value stays out of the process environment, which is where debug endpoints and crash dumps look first. And it makes runtime reads provable, which the lab relies on. The cost is a few lines of code and a failure mode you have to handle at read time instead of at startup.

Here is the Cloud Run service from module 1 with the identity and the volume added. Only the template changes. From here on an image always exists, so the service is shown without the “no image yet” branch from module 1 (and the Terraform tab drops its count).

infra/Program.cs (excerpt)
// …
var service = new Gcp.CloudRunV2.Service("sniplink-api", new()
{
Name = "sniplink-api",
Location = region,
DeletionProtection = false,
Ingress = "INGRESS_TRAFFIC_ALL",
Template = new Gcp.CloudRunV2.Inputs.ServiceTemplateArgs
{
ServiceAccount = runtimeSa.Email,
Volumes =
{
new Gcp.CloudRunV2.Inputs.ServiceTemplateVolumeArgs
{
Name = "lab-key",
Secret = new Gcp.CloudRunV2.Inputs.ServiceTemplateVolumeSecretArgs
{
Secret = labSecret.SecretId,
Items =
{
new Gcp.CloudRunV2.Inputs.ServiceTemplateVolumeSecretItemArgs
{
Version = "latest",
Path = "lab-key", // file name inside the mount
},
},
},
},
},
Containers =
{
new Gcp.CloudRunV2.Inputs.ServiceTemplateContainerArgs
{
Image = image,
Ports = new Gcp.CloudRunV2.Inputs.ServiceTemplateContainerPortsArgs
{
ContainerPort = 8080,
},
Envs =
{
new Gcp.CloudRunV2.Inputs.ServiceTemplateContainerEnvArgs
{
Name = "LABKIT_STUDENT_ID",
Value = studentId,
},
},
VolumeMounts =
{
new Gcp.CloudRunV2.Inputs.ServiceTemplateContainerVolumeMountArgs
{
Name = "lab-key",
MountPath = "/secrets", // file appears at /secrets/lab-key
},
},
},
},
},
}, new CustomResourceOptions
{
// Cloud Run checks secret access when it creates the revision.
DependsOn = { repo, labKeyAccess, labKeyVersion },
});
// … public invoker binding and outputs from module 1, unchanged

The DependsOn matters. Without it, Pulumi may try to create the revision in parallel with the IAM binding, Cloud Run checks whether the runtime account can access the secret, and the revision fails with a permission error that disappears on the next pulumi up. Flaky infrastructure is worse than broken infrastructure, so make the order explicit.

Run pulumi up. The preview should show four creates (secret, version, service account, IAM member) and one update to the service. The update creates a new revision, because the template changed. Then confirm from the CLI, read-only:

Terminal window
gcloud secrets versions list sniplink-lab-key --project "$PROJECT_ID"
gcloud run services describe sniplink-api \
--region "$REGION" --project "$PROJECT_ID" \
--format 'value(spec.template.spec.serviceAccountName)'

The second command should print sniplink-runtime@PROJECT_ID.iam.gserviceaccount.com.

LabKit’s POST /_lab/secret-proof needs the lab key to compute a proof, and by the plumbing-versus-skill rule it does not read the key for you. It asks the DI container for an ILabSecretSource and calls GetKeyAsync(). Reading a secret correctly at runtime is the skill this module teaches, so that part is yours.

Requirements for your implementation:

  • Read the key from the file Cloud Run mounts, /secrets/lab-key. Make the path configurable so you can point it at a local file during development.
  • Trim whitespace. A trailing newline is the most common reason a correct key produces a wrong HMAC.
  • Fail loudly if the file is missing or empty, with a message that names the path but never the content.
  • Do not read the file on every call forever, but do not cache it for the life of the process either. More on that below.

Try writing it yourself before you look. Here is my reference implementation.

src/Sniplink.Api/FileLabSecretSource.cs
using CuriousDev.LabKit;
namespace Sniplink.Api;
public sealed class FileLabSecretSource(IConfiguration config, TimeProvider time) : ILabSecretSource
{
private static readonly TimeSpan CacheFor = TimeSpan.FromSeconds(30);
private readonly string _path = config["LabSecret:Path"] ?? "/secrets/lab-key";
private readonly SemaphoreSlim _gate = new(1, 1);
private volatile CachedKey? _cached;
public async ValueTask<string> GetKeyAsync(CancellationToken cancellationToken = default)
{
var current = _cached;
if (current is not null && time.GetUtcNow() - current.ReadAt < CacheFor)
{
return current.Value;
}
await _gate.WaitAsync(cancellationToken);
try
{
current = _cached; // another caller may have refreshed it while we waited
if (current is not null && time.GetUtcNow() - current.ReadAt < CacheFor)
{
return current.Value;
}
if (!File.Exists(_path))
{
throw new InvalidOperationException($"Lab secret file not found at '{_path}'.");
}
var value = (await File.ReadAllTextAsync(_path, cancellationToken)).Trim();
if (value.Length == 0)
{
throw new InvalidOperationException($"Lab secret file at '{_path}' is empty.");
}
_cached = new CachedKey(value, time.GetUtcNow());
return value;
}
finally
{
_gate.Release();
}
}
private sealed record CachedKey(string Value, DateTimeOffset ReadAt);
}

Register it in Program.cs, before AddLabKit():

src/Sniplink.Api/Program.cs
using Sniplink.Api; // top-level statements live in the global namespace
// …
builder.Services.AddSingleton(TimeProvider.System);
builder.Services.AddSingleton<ILabSecretSource, FileLabSecretSource>();
builder.Services.AddLabKit();
// …

For local runs, put a dummy key in a file outside the repo and point the config at it with LabSecret__Path=/tmp/lab-key dotnet run. Never put the real lab key in appsettings.Development.json; the whole module is about not doing that.

Every read of /secrets/lab-key on Cloud Run is a Secret Manager access: it costs a little money, adds latency and counts against quotas. Reading on every request is correct but wasteful on a hot path. Caching forever is cheap but defeats the volume: you would rotate the secret and the running process would keep using the old value until the instance is replaced, which is the env-var behaviour with extra steps.

A short cache is the compromise. With 30 seconds, a rotated key is picked up within 30 seconds on every instance, and a busy service makes at most two reads per minute per instance. Pick the window from your rotation story: how long can the old and new value both be in use? If the answer is “the old one must stop working immediately”, you need a different mechanism (revoke at the provider, short-lived tokens), not a shorter cache.

The volatile reference to an immutable record is deliberate: readers either see the old cached pair or the new one, never a torn mix of a new value with an old timestamp. The semaphore stops a burst of requests at expiry from all reading the file at once.

A secret you have never rotated is a secret you cannot rotate in an incident. The team in the incident log learned that at 10:06. Rotation should be a boring, rehearsed operation, and with a volume plus latest it is:

  1. Add a new version. The new value becomes latest. Consumers reading latest pick it up on their next read; with the 30-second cache, within half a minute.

  2. Wait for the rollout. Confirm consumers are using the new value: logs, a health signal, or in this lab, a passing proof. Anything that pins the old version number needs a new revision first.

  3. Disable the old version. Not destroy: disable. If something you forgot still depends on it, it fails visibly and you can re-enable the version in seconds.

  4. Destroy later. After a quiet period, destroy the old version, or let a policy do it.

In Pulumi, the simplest form of this is what you already wrote: change the config value and run pulumi up. Because SecretData cannot be updated in place, Pulumi replaces the SecretVersion resource. It creates the new version first and then removes the old resource, and DeletionPolicy = "DISABLE" turns that removal into a disable instead of a destroy. One command does steps 1 and 3.

That is fine for Sniplink, where the only consumer reads latest and caches for 30 seconds. For a secret with consumers you don’t control, split it into two pulumi up runs: add a second SecretVersion resource for the new value and keep the old one, then set Enabled = false on the old resource once the rollout is done. The code is slightly longer; the incident is much shorter.

The lab has to prove three things from outside your project, without ever seeing your key or your configuration: you read the key from Secret Manager, you read it at runtime, and the service runs as a dedicated identity.

secretProof. The platform sends a random nonce to POST /_lab/secret-proof. LabKit calls your ILabSecretSource, computes HMAC-SHA256(key, nonce) and returns the HMAC as hex, plus the current K_REVISION. The platform knows your lab key v1, computes the same HMAC and compares. Only the HMAC leaves your service. An HMAC does not reveal the key, and because the nonce is fresh every time, an old response cannot be replayed. The platform also records the revision.

secretRotation. When the run reaches this check, it pauses and the lab page shows your lab key v2. You set it in config and run pulumi up:

Terminal window
pulumi config set --secret sniplink:labKey # paste lab key v2
pulumi preview # expect: SecretVersion replaced, service unchanged
pulumi up

Read the preview. You should see the secret version being replaced and no change to sniplink-api. If the service shows an update, something in your template depends on the secret value, and the check will fail. Wait a minute for your cache to expire, then press Continue. The platform sends a new nonce and expects an HMAC computed with key v2 and the same revision it recorded in the first check.

Why that combination is convincing:

  • Hardcoding the key in code or appsettings.json cannot pass: key v2 did not exist when you built the image, and rebuilding with it creates a new revision.
  • Putting the key in a plain env var cannot pass: changing it creates a new revision.
  • A volume reading latest passes naturally: new version, same revision, new value on the next read.

Honesty about the edges, as promised in module 0: an env var referencing a secret with latest could pass if Cloud Run happens to start a fresh instance between your rotation and the check, because new instances resolve latest again. The lab does not try to catch that case. It is the inconsistency the comparison table warns about, and it is not the design this module teaches.

tokenClaim. The email claim in the Google ID token from /_lab/verify is the service account the revision runs as. The check requires that it does not end with -compute@developer.gserviceaccount.com, the suffix of the default compute account. With sniplink-runtime it will be sniplink-runtime@PROJECT_ID.iam.gserviceaccount.com. The token is signed by Google, so the service cannot claim an identity it does not have. What the check cannot see is which roles the account holds; that is the Verified tier’s job, when it arrives.

Goal: Sniplink runs as sniplink-runtime, reads the lab key from Secret Manager through a volume at runtime, and survives a key rotation without a new revision.

  1. Open the m03-secrets lab page and copy lab key v1. Store it with pulumi config set --secret sniplink:labKey.
  2. Add the secret, the secret version, the sniplink-runtime service account and the secretAccessor binding on that secret to your Pulumi program.
  3. Update the Cloud Run service: ServiceAccount = runtimeSa.Email, a secret volume for sniplink-lab-key with item path lab-key and version latest, mounted at /secrets.
  4. Implement and register ILabSecretSource reading /secrets/lab-key. Build and push a new image, then pulumi up.
  5. Press Verify. When the run reaches secret-rotation, it pauses and the page shows lab key v2.
  6. Set lab key v2 in config, confirm with pulumi preview that the service is unchanged, run pulumi up, wait a minute for your cache to expire, then press Continue.

Check your lab

Loading your lab…

Sign in to verify this lab and save your progress. Everything above works without an account.

  • Your Cloud Run service URL (pulumi stack output url)
  • Lab key v1

5 checks run against the service you deployed.

What each check verifies:

  • token (baseline): a valid Google ID token for the course audience with your student id, issued to a service in your project.
  • health: GET /health still returns 200.
  • runtime-identity: the token’s email does not end with -compute@developer.gserviceaccount.com, so the revision is not running as the default compute account.
  • secret-proof: POST /_lab/secret-proof returns HMAC-SHA256(lab key v1, nonce); the platform records the revision.
  • secret-rotation: after rotation, the HMAC matches lab key v2 and the revision is the same as in secret-proof.

None of these are self-reported: the HMACs can only be produced with the key, and the token is signed by Google. The checks do not prove the runtime account has no other roles; a stricter check would read the project’s IAM policy, which needs the Verified tier.

Keep the stack. Module 4 builds on the same service, identity and secret.

If you are pausing the course for more than a few days, destroy everything. To come back, run pulumi up (it recreates the repository and fails on the service, because the image went with the old repository), push the image again with the same tag, and run pulumi up once more:

Terminal window
cd infra
pulumi destroy

DeletionPolicy = "DISABLE" protects individual versions during rotation, but pulumi destroy also deletes the secret itself, and a deleted secret takes all its versions with it. That is what you want for a course project. Your lab keys stay available on the lab page.

  • Secrets don’t belong in git, images, env files or plain env vars; they belong in a store with per-secret access control, versions and audit logs, which on Google Cloud is Secret Manager.
  • Pulumi config secrets, encrypted with your Cloud KMS key, let you get a value into infrastructure code without writing it down anywhere in plain text.
  • A dedicated service account with one resource-level role replaces the shared, often over-privileged default compute account.
  • A secret mounted as a volume and read at latest rotates without a new revision; a short cache balances cost against how quickly a rotation takes effect.
  • Rotation is add, wait, disable, destroy later, and it should be rehearsed before you need it.