Skip to content

Signed in as

Sign in to Curious Dev Learn

Sign in to verify labs, save your progress and get a certificate. Lessons stay open without an account.

Module 022 hLab

A container worth shipping

Replace the default image with a small, non-root, shell-free multi-stage build and push it to Artifact Registry.

Nothing is down. The service answers, latency is fine, and yet a security team has just handed the owners 88 findings to patch or justify by Friday. Almost none of them are in the code they wrote. They are in Perl, Git, an SSH client and Python, which shipped because the Dockerfile started FROM the .NET SDK image and ran dotnet run inside it. The published application is 4 MB. The image around it is 917 MB, runs as root and has a shell that an attacker would be very happy to find.

This module answers one question: what should a production .NET container contain, and how do you make sure it contains nothing else? You will look at the image you shipped in module 1, write a multi-stage Dockerfile that produces a small, non-root, shell-free image, push it to Artifact Registry, and point the Cloud Run service at it with Pulumi.

In module 1 you built the image with dotnet publish /t:PublishContainer and let the SDK make every decision. Before replacing it, let me show you what those decisions were. Set the variables you will use for the rest of this lesson:

Terminal window
export PROJECT_ID=your-project-id
export REGION=europe-west1
export IMAGE=$REGION-docker.pkg.dev/$PROJECT_ID/sniplink/api

The module-1 image went straight from the SDK to Artifact Registry, so it is not in your local Docker cache yet. Its full reference, with the m01-<sha> tag, is in your stack config, so read it from there instead of retyping it:

Terminal window
# the image the service runs right now, e.g. …/sniplink/api:m01-3f9c2ab
export M01_IMAGE=$(cd infra && pulumi config get sniplink:image)
# the SDK pushed this image directly; pull it so Docker can inspect it
docker pull $M01_IMAGE

Now ask Docker what is in it:

Terminal window
# size in bytes, the user the process runs as, and the base OS
docker image inspect $M01_IMAGE --format '{{.Size}} bytes, user={{.Config.User}}'
docker image inspect $M01_IMAGE --format '{{json .Config.Env}}'
docker history $M01_IMAGE

You should see something like this:

  • Size in the low 200s of megabytes.
  • User set to app (or its numeric ID). Since .NET 8, the SDK container tooling runs the app as the non-root user from the base image by default.
  • Env containing APP_UID=1654, ASPNETCORE_HTTP_PORTS=8080 and DOTNET_RUNNING_IN_CONTAINER=true. These come from the Microsoft base image, not from your project.
  • History with a dozen or so layers from the base image (shown as <missing>, which is normal for pulled images) and one layer on top that holds your published app.

The SDK picked mcr.microsoft.com/dotnet/aspnet:10.0 as the base, which for .NET 10 is Ubuntu 24.04 “noble”. That is a full distribution: it has bash, apt, dpkg, ICU, time zone data and a few hundred other files your API never touches. Prove it:

Terminal window
# the full image has a shell, so this works
docker run --rm --entrypoint sh $M01_IMAGE -c 'id && ls /usr/bin | wc -l'

So the module-1 image is far better than the one in the incident log: it is not built on the SDK and it does not run as root. But every binary in /usr/bin is a package a scanner will report, a file you are responsible for patching, and a tool available to anyone who gets code execution in your container.

The single-stage Dockerfile in the incident looked roughly like this: FROM sdk, copy everything, dotnet run. The SDK image has to contain compilers, NuGet, MSBuild and their dependencies, which is why it is close to a gigabyte. You need all of that to build the app and none of it to run it.

A multi-stage build uses one image to build and a different, much smaller one to run. Only what you explicitly COPY --from the build stage ends up in the final image. Put this at the repository root, next to Sniplink.sln:

Dockerfile
# syntax=docker/dockerfile:1
# ---- build stage: full SDK, runs on the machine doing the build ----
FROM --platform=$BUILDPLATFORM mcr.microsoft.com/dotnet/sdk:10.0 AS build
ARG TARGETARCH
WORKDIR /src
# 1. Copy only the project file and restore. This layer is cached
# until the .csproj changes, so code edits don't re-download packages.
COPY src/Sniplink.Api/Sniplink.Api.csproj src/Sniplink.Api/
RUN dotnet restore src/Sniplink.Api/Sniplink.Api.csproj -a $TARGETARCH
# 2. Copy the source and publish without restoring again.
COPY src/Sniplink.Api/ src/Sniplink.Api/
RUN dotnet publish src/Sniplink.Api/Sniplink.Api.csproj \
-a $TARGETARCH \
--configuration Release \
--no-restore \
--output /app/publish \
/p:UseAppHost=false
# ---- runtime stage: chiseled ASP.NET Core runtime, no shell ----
FROM mcr.microsoft.com/dotnet/aspnet:10.0-noble-chiseled AS final
WORKDIR /app
COPY --from=build /app/publish .
# Run as the non-root 'app' user defined by the base image (UID 1654).
USER $APP_UID
# Documentation only: the port ASP.NET Core and Cloud Run agree on.
EXPOSE 8080
ENTRYPOINT ["dotnet", "Sniplink.Api.dll"]

Each line has a reason:

  • --platform=$BUILDPLATFORM and -a $TARGETARCH. Cloud Run runs linux/amd64. If you are on an Apple silicon Mac, a plain docker build produces an arm64 image that Cloud Run cannot start. These two lines run the SDK natively on your machine and cross-compile the app for the target architecture, which is much faster than emulating the whole SDK. Docker fills in both variables; the .NET CLI accepts amd64 and arm64 as architecture names.
  • Copy the .csproj first, then restore. Docker caches each layer and reuses it until the files that went into it change. Package references change rarely; code changes all the time. Splitting the two means a code edit re-runs publish but not restore. If your repository has a Directory.Build.props, Directory.Packages.props or nuget.config at the root, copy those before the restore too, or restore will not see them.
  • --no-restore. Publish reuses the restore from the cached layer instead of doing it again.
  • /p:UseAppHost=false. Skips the native launcher executable; we start the app with dotnet Sniplink.Api.dll, so the launcher would be dead weight.
  • USER $APP_UID. Covered in the next section.
  • EXPOSE 8080. Does not open anything. It documents that the image listens on 8080, which matches both ASPNETCORE_HTTP_PORTS in the base image and Cloud Run’s default PORT. Your module-1 code still reads PORT and wins if Cloud Run sends something else.
  • Exec-form ENTRYPOINT. The JSON array form runs dotnet as PID 1 with no shell wrapper, so it receives SIGTERM directly. That matters for graceful shutdown in module 4. The shell form would not work here anyway: there is no shell.

You will notice there is no RUN instruction in the final stage. There cannot be: RUN needs a shell, and the chiseled image does not have one. Anything you need to prepare happens in the build stage and gets copied over.

Microsoft publishes the .NET runtime images in several families. They differ in what operating system files sit under your app.

Family Example tag What is inside Default user When I use it
Full aspnet:10.0, aspnet:10.0-noble Full Ubuntu userland: bash, apt, ICU, tzdata, CA certificates root (the app user exists, you opt in) Apps that need extra OS packages installed with apt; local debugging
Alpine aspnet:10.0-alpine Alpine Linux with musl libc, BusyBox shell, apk; invariant globalization by default root Small images when you also want a shell; watch for musl differences in native libraries
Chiseled aspnet:10.0-noble-chiseled Only the files the .NET runtime needs: libc, OpenSSL, CA certificates. No shell, no package manager, no ICU, no tzdata app (UID 1654) The default for ASP.NET Core APIs. This course uses it
Chiseled extra aspnet:10.0-noble-chiseled-extra Chiseled plus ICU and tzdata app (UID 1654) Apps that need culture-aware formatting or named time zones
Runtime deps runtime-deps:10.0-noble-chiseled Chiseled OS without the .NET runtime app (UID 1654) Self-contained or Native AOT apps that carry their own runtime

As a rough guide, the chiseled ASP.NET Core image is about half the size of the full one, and the runtime-deps chiseled image is an order of magnitude smaller again. The size matters less than the count: a chiseled image contains a handful of packages, so a scanner has a handful of things to report.

“Chiseled” comes from Canonical’s chisel tool, which installs slices of Ubuntu packages, only the files a given workload needs, instead of whole packages. You get Ubuntu’s security updates for those files without the rest of the distribution.

Every .NET image since .NET 8 defines a user named app with UID 1654 and exports that number as the APP_UID environment variable. The full and Alpine images still default to root for compatibility; the chiseled images already default to app.

I write USER $APP_UID anyway, even though the chiseled base already sets it. The security property of my image should be visible in my Dockerfile, not inherited silently from someone else’s. If a teammate later swaps the base for the full image to debug something, the user stays non-root. Using the numeric ID rather than the name also lets Kubernetes-style runAsNonRoot checks verify it without reading /etc/passwd.

What non-root buys you: the published files in /app are owned by root and only readable by app, so a compromised process cannot rewrite its own binaries. It cannot bind ports below 1024, which is why .NET 8 moved the default from 80 to 8080. And a container escape starts from an unprivileged account.

The plain chiseled image has no ICU, the library .NET uses for culture-specific formatting, sorting and casing. Without ICU, .NET must run in globalization-invariant mode: every culture behaves like the invariant culture. The chiseled base sets DOTNET_SYSTEM_GLOBALIZATION_INVARIANT=true so the runtime does not fail at startup looking for ICU.

Again, I put the decision in the project rather than relying on the base image:

src/Sniplink.Api/Sniplink.Api.csproj
<Project Sdk="Microsoft.NET.Sdk.Web">
<PropertyGroup>
<TargetFramework>net10.0</TargetFramework>
<Nullable>enable</Nullable>
<ImplicitUsings>enable</ImplicitUsings>
<!-- No ICU in the chiseled image: make that a build-time decision. -->
<InvariantGlobalization>true</InvariantGlobalization>
</PropertyGroup>
<!-- … package references unchanged … -->
</Project>

This also makes your tests and local runs behave the same way as production, which an environment variable in a base image does not. Sniplink stores URLs and random slugs, so it loses nothing.

You need -extra (and InvariantGlobalization set to false) when the app:

  • formats dates, numbers or currency for specific cultures, or sorts text by language rules;
  • converts times with TimeZoneInfo.FindSystemTimeZoneById("Europe/Kyiv"); without tzdata only UTC exists;
  • depends on a library that needs ICU (some database drivers have historically refused to run in invariant mode).

If you are unsure, run your test suite with InvariantGlobalization on. The failures will tell you.

Now build the image locally and compare it to module 1:

Terminal window
docker build -t sniplink-api:local .
docker images | grep -E 'sniplink|api'

The first time you try docker exec -it <container> bash on a chiseled image you get executable file not found in $PATH. That is the feature working. But you still need to debug things, so here is what I use instead, from most to least frequent.

Logs and the app itself. Most production questions are answered by structured logs and by endpoints you control. On Cloud Run there is no exec into an instance anyway, whatever the base image, so your debugging plan already has to work without a shell. LabKit’s /_lab/info is a small example of the pattern: the process reports facts about itself.

Terminal window
gcloud run services logs read sniplink-api --region $REGION --limit 50

A debug container next to the running one. Locally, you can start a throwaway container that shares the app container’s process and network namespaces. It brings its own shell and tools; your production image stays untouched.

Terminal window
docker run -d --name sniplink -p 8080:8080 sniplink-api:local
# busybox brings the shell; --pid and --network join the app's namespaces
docker run --rm -it --pid=container:sniplink --network=container:sniplink busybox sh
# inside: ps, wget -qO- localhost:8080/health, ls /proc/1/root/app

The app’s filesystem is visible under /proc/1/root. Docker Desktop also offers docker debug, which does the same thing with a richer toolbox; it has been free for all users since Docker Desktop 4.49.

A debug build target. When you really need a shell inside the same image layout, add a stage that uses the full image and the same published output, and build it only on demand:

Dockerfile (optional extra stage)
# docker build --target debug -t sniplink-api:debug .
FROM mcr.microsoft.com/dotnet/aspnet:10.0 AS debug
WORKDIR /app
COPY --from=build /app/publish .
USER $APP_UID
ENTRYPOINT ["dotnet", "Sniplink.Api.dll"]

Put it before the final stage. A plain docker build builds the last stage in the file, so production stays chiseled. Never push the debug tag to the registry your service deploys from.

A chiseled, non-root container nudges you toward treating the filesystem as read-only. The app user cannot write to /app. Cloud Run’s filesystem is writable, but it is in memory and counts against the instance’s memory limit, and it disappears with the instance. Anything that matters belongs in a service (the next course covers persistence), and anything temporary goes to /tmp.

You can check that Sniplink does not depend on a writable filesystem locally:

Terminal window
docker rm -f sniplink
docker run --rm -p 8080:8080 --read-only --tmpfs /tmp sniplink-api:local

If it starts and /health returns 200, the app makes no hidden writes outside /tmp.

The image above is framework-dependent: the .NET runtime comes from the base image and your publish output contains only your assemblies. There are three ways to go further. None of them is required for the lab, and for most APIs I would not start with them.

Self-contained. Publish with --self-contained and the runtime ships inside your output. You then use runtime-deps as the base. The total size is roughly the same, since the runtime just moved from one layer to another. What you gain is independence from the shared runtime version; what you lose is base-layer sharing between services and the habit of getting runtime patches by rebuilding on a newer base tag.

Trimming. PublishTrimmed=true removes unused code from your app and the runtime. It only works for self-contained apps, and it relies on static analysis: anything reached only by reflection can be trimmed away and fail at run time. The trimmer warns you (IL2026 and friends), and those warnings are not optional reading.

Native AOT. PublishAot=true compiles the app to a single native executable with no JIT. For a web API on Cloud Run this is the most interesting option, because it cuts startup time and memory, and Cloud Run starts instances on demand. The costs are real:

  • ASP.NET Core supports AOT for minimal APIs, not MVC controllers or Razor. You use WebApplication.CreateSlimBuilder and source-generated JSON (JsonSerializerContext) for every type you serialize.
  • Libraries that rely on runtime reflection or code generation may not work or may need configuration; check each dependency before committing.
  • The build needs a native toolchain (clang and friends) in the build stage and cannot cross-compile between operating systems. Microsoft publishes SDK images with that toolchain preinstalled, such as mcr.microsoft.com/dotnet/sdk:10.0-noble-aot.
  • Builds get noticeably slower.

My rule: start framework-dependent on chiseled. Measure cold starts on Cloud Run (module 4 shows how). If they matter for your traffic pattern and your dependencies are AOT-compatible, move to Native AOT on runtime-deps:10.0-noble-chiseled. Otherwise, MinInstanceCount = 1 is often the cheaper fix in engineering time, if not in money.

When you run docker build ., Docker first sends the whole directory, the build context, to the builder. Without a .dockerignore, that includes bin/ and obj/ from your machine, the .git folder and your Pulumi stack files. That is slow, and the obj/ folders are actively harmful: they contain restore results for your OS that can confuse dotnet publish inside Linux. Add this at the repository root:

.dockerignore
# build output and IDE state from the host
**/bin/
**/obj/
**/.vs/
**/.vscode/
**/.idea/
**/*.user
**/TestResults/
# not needed to build the image
.git/
.github/
infra/
tests/
**/*.md
# never send these to a builder
**/.env
**/*.pfx
**/secrets/
# the build files themselves
Dockerfile*
.dockerignore

infra/ deserves a note: your Pulumi stack files will hold encrypted secrets from module 3 onward. They are encrypted, but they still have no business in an image build. Tests are excluded because they run outside the image (in module 6 Cloud Build runs dotnet test as its own step).

Rebuild and look at the first lines of the output; the context transfer should now be measured in kilobytes:

Terminal window
docker build --progress=plain -t sniplink-api:local . 2>&1 | grep -i 'transferring context'

You need a tag that changes on every build. Pulumi only creates a new Cloud Run revision when the configuration changes, so reusing a tag such as latest means pulumi up reports no changes and nothing deploys. I use the module name plus the short commit hash:

Terminal window
export TAG=m02-$(git rev-parse --short HEAD)
Terminal window
# Artifact Registry credentials for Docker (already done in module 1)
gcloud auth configure-docker $REGION-docker.pkg.dev
# build for Cloud Run's architecture, whatever your laptop is
docker build --platform linux/amd64 -t $IMAGE:$TAG .
# smoke test locally before pushing
docker run --rm -d --name sniplink -p 8080:8080 -e PORT=8080 $IMAGE:$TAG
curl -s localhost:8080/health
curl -s localhost:8080/_lab/info
docker rm -f sniplink
docker push $IMAGE:$TAG

Locally, /_lab/info shows empty service and revision values because there is no Cloud Run environment, but user.uid, shellPresent, invariantGlobalization and dotnet are already real. Check them now; it is the cheapest place to find a mistake.

Prefer deploying by digest rather than tag. A tag can be moved to another image; a digest cannot. Cloud Run resolves a tag to a digest when it creates the revision anyway, but putting the digest in your config makes the git history say exactly what ran:

Terminal window
# after the push, Docker knows the registry digest
export IMAGE_REF=$(docker inspect --format '{{index .RepoDigests 0}}' $IMAGE:$TAG)
echo $IMAGE_REF # europe-west1-docker.pkg.dev/…/sniplink/api@sha256:…

With Cloud Build, copy the digest from the end of the build output instead.

The infrastructure code does not change: in module 1 the Cloud Run service already reads the image from stack config. You only change the value.

Terminal window
cd infra
pulumi config set sniplink:image $IMAGE_REF
pulumi preview # expect one update: the service's container image
pulumi up

The preview should show a single in-place update on the Cloud Run service. If it also wants to replace the Artifact Registry repository or the IAM binding, stop and read the diff; something else drifted.

When the update finishes, call the service:

Terminal window
export URL=$(pulumi stack output url)
curl -s $URL/health
curl -s $URL/_lab/info

You should see user.uid 1654, shellPresent false, invariantGlobalization true, a dotnet version starting with 10., and a new revision.

Vulnerability scanning in Artifact Registry

Section titled “Vulnerability scanning in Artifact Registry”

A smaller image has fewer findings, but “fewer” is not “none”, and new CVEs appear in packages after you ship. Artifact Registry can scan images automatically through Artifact Analysis: enable the Container Scanning API and every image pushed to the repository is scanned for known vulnerabilities in its OS packages and in application language packages, including the NuGet packages of a .NET app. After the first scan, results keep updating as new vulnerabilities are published, for images pulled within the last 30 days; after that they go stale.

Scanning is billed per scanned image, so check current pricing before turning it on in a project that pushes on every commit. Following the course rule, enabling an API is still infrastructure, so it goes in code:

infra/Program.cs (optional)
// … inside the stack, next to the other resources
// Optional: on-push vulnerability scanning for Artifact Registry. Billed per image.
var containerScanning = new Gcp.Projects.Service("container-scanning", new()
{
ServiceName = "containerscanning.googleapis.com",
DisableOnDestroy = false, // don't switch the API off for the whole project on destroy
});

Then push an image (scanning applies to pushes after the API is on) and read the results from the CLI:

Terminal window
gcloud artifacts docker images describe $IMAGE_REF --show-package-vulnerability

If you push the module-1 image and the module-2 image to a scanning-enabled repository, the difference in findings is the most convincing argument for this module I know.

Goal: Sniplink runs on Cloud Run from your own multi-stage image on aspnet:10.0-noble-chiseled, as a non-root user, with invariant globalization declared in the project.

  1. Add the Dockerfile and .dockerignore from this lesson at the repository root.
  2. Set <InvariantGlobalization>true</InvariantGlobalization> in Sniplink.Api.csproj.
  3. Build the image with docker build (or gcloud builds submit) and push it to $REGION-docker.pkg.dev/$PROJECT_ID/sniplink/api with a unique tag.
  4. Set sniplink:image to the new image (digest preferred) and run pulumi up.
  5. Check $URL/_lab/info yourself, then submit your service URL below.

Check your lab

Loading your lab…

Sign in to verify this lab and save your progress. Everything above works without an account.

  • Your Cloud Run service URL (pulumi stack output url)

6 checks run against the service you deployed.

Some checks are self-reported: the checker trusts what your service says about itself and verifies the rest from the outside.

What the checks verify:

  • Token (baseline): POST /_lab/verify returns a valid Google ID token for the platform audience with your student id, with your nonce, from a real Cloud Run service in your project.
  • Health: GET /health returns 200, so the new image starts and serves.
  • Non-root: user.uid in /_lab/info is not 0.
  • No shell: shellPresent is false.
  • Invariant globalization: invariantGlobalization is true.
  • Runtime: dotnet starts with 10..

The last four checks are self-reported. The running process tells the platform about itself, and the platform cannot see your image. A determined student could patch LabKit to say anything, and even an honest LabKit only sees the running process, not the size or origin of the image. That is acceptable for a learning certificate, and I would rather tell you than pretend otherwise. The optional Verified tier, coming later, will check strictly: with read access granted to the platform, it reads the image manifest from Artifact Registry, compares the base layers against the published chiseled image, checks the configured user and the image size, and confirms the Cloud Run revision runs that exact digest.

Keep the stack running: module 3 adds a secret and a service identity to the same service.

The module-1 image is no longer used. Artifact Registry storage costs cents, but there is no reason to keep it. M01_IMAGE is the variable you exported at the start of this lesson; sniplink:image now points at the module-2 image, so do not read it from config again:

Terminal window
gcloud artifacts docker images delete $M01_IMAGE --delete-tags --quiet
docker image prune -f

If you are pausing the course for a while, destroy the stack instead. Destroying it deletes the repository and its images, so when you return, run pulumi up (it recreates the repository; the service fails until the image exists), push the image again with the module 2 commands, set sniplink:image to the new digest and run pulumi up once more:

Terminal window
cd infra
pulumi destroy
  • How to inspect an image’s size, layers, user and environment with docker image inspect and docker history, and what PublishContainer chooses for you.
  • How a multi-stage Dockerfile keeps the SDK out of production, and how copying the project file first makes restores cacheable.
  • The .NET image families, why chiseled is a good default for APIs, what the non-root app user (UID 1654) gives you, and when you need -extra instead of InvariantGlobalization.
  • How to debug without a shell, and where self-contained, trimming and Native AOT fit for a web API.
  • How to push to Artifact Registry with Docker or Cloud Build, deploy by digest through Pulumi config, and turn on vulnerability scanning.