Lifecycle: start, scale, stop
Make Sniplink start cleanly, scale within a budget and stop without dropping requests on Cloud Run.
What breaks
Section titled “What breaks”The service in this log is fine until a deploy. Then about 2 % of requests fail for twenty seconds: some with 503 because the new revision got traffic before it was ready, some with a connection reset because the old instance was killed mid-request. Later a marketing email lands, the autoscaler does exactly what it was allowed to do, and the bill follows.
Neither problem is in the business logic. Both are about the lifecycle of an instance: how it starts, when it is ready, how many requests it takes at once, how many copies may exist, and how it stops. Cloud Run has a setting for each, and the defaults do not know when your app is warm. In this module you make Sniplink tell Cloud Run the truth about itself, and you put a ceiling on how far it can scale.
The life of an instance
Section titled “The life of an instance”Cloud Run runs your revision as a set of instances: copies of your container, each with its own CPU and memory. You never create or delete them yourself. The autoscaler does, based on incoming requests. Every instance goes through the same stages.
1. Start Cloud Run pulls the image, starts the container, sets PORT. 2. Startup probe Cloud Run polls your startup probe. No traffic is sent yet. Probe succeeds -> instance is ready. Probe fails failureThreshold times -> instance is killed. 3. Serving Requests arrive, up to the concurrency limit at once. The liveness probe runs in the background. 4. Idle No requests in flight. The instance may be kept around for a while so the next request avoids a cold start. 5. SIGTERM Cloud Run decides to stop the instance (scale-in, old revision drained after a deploy, maintenance). 6. Grace period 10 seconds to finish in-flight work and exit. 7. SIGKILL If the process is still alive after 10 s, it is killed.A few details matter for .NET.
Scale to zero. With MinInstanceCount = 0, once all instances have been idle long enough, Cloud Run stops them all. You pay nothing for CPU and memory while nothing runs. The next request pays the price instead: it waits for a cold start.
What makes a cold start slow. A cold start is the time from “a request needs an instance” to “an instance is ready”. For a .NET API the main contributors are:
- Image pull. Bigger images take longer. The chiseled image from module 2 helps here.
- Runtime start and JIT. The .NET runtime loads, and your code is compiled to machine code on first use. ReadyToRun and Native AOT reduce this; we covered the trade-offs in module 2.
- Host build. Configuration, DI container, Kestrel startup. Usually fast unless you do I/O in constructors.
- Your warm-up. Anything you do before you can honestly answer requests: reading a secret, loading a lookup table, opening a connection pool.
- CPU during startup. All of the above is CPU-bound, and a 1 vCPU instance does it on one core. This is what startup CPU boost addresses (more on that below).
Cold starts do not only happen after scale to zero. Every scale-out creates a cold instance, and every deploy creates a whole new set of them. That is why the incident above was tied to deploys: the new revision was cold, and something told Cloud Run it was ready too early.
The 10-second grace period is fixed. You do not configure it on Cloud Run. Your job is to make sure the process uses those 10 seconds well: stop taking new work, finish what is in flight, flush what needs flushing, exit. Anything still running at second 10 is gone, mid-response.
Health: live versus ready
Section titled “Health: live versus ready”Since module 1 Sniplink has had a single GET /health that returns 200. That endpoint answers “is anything listening?”, which is the same question Cloud Run’s default TCP startup probe already asks. It says nothing about whether the app is ready to serve.
Two questions are worth separating:
- Live: is the process working at all? If not, the only fix is a restart. This must not depend on anything outside the process.
- Ready: has the instance finished its startup work, so the first real request will be served properly and quickly?
On Cloud Run you map these to two probes. The startup probe runs only while the instance starts, and it is what gates traffic: no request is routed to an instance until it passes. The liveness probe runs for the rest of the instance’s life, and repeated failures restart it.
The readiness flag
Section titled “The readiness flag”Readiness needs something to flip. I use a tiny singleton with a flag, set by a background service when startup work is done. It is deliberately boring.
namespace Sniplink.Api.Lifecycle;
public sealed class StartupState{ private volatile bool _ready;
public bool IsReady => _ready;
public void MarkReady() => _ready = true;}The warm-up runs in a BackgroundService. That matters: hosted services are started before Kestrel begins listening, so if you do slow work in StartAsync, the port does not open until the work is done. A background service lets Kestrel start immediately while the readiness flag stays false. The startup probe then tells Cloud Run exactly when the instance is warm, instead of “when the port opened”.
using System.Text.Json;using CuriousDev.LabKit;
namespace Sniplink.Api.Lifecycle;
public sealed class WarmupService( StartupState state, ILabSecretSource secrets, ILogger<WarmupService> logger) : BackgroundService{ protected override async Task ExecuteAsync(CancellationToken stoppingToken) { await Task.Yield(); // never block host startup, whatever the .NET version
// 1. The secret volume from module 3 must be mounted and readable. // This also fills the short cache in FileLabSecretSource. _ = await secrets.GetKeyAsync(stoppingToken);
// 2. Touch the code paths the first request will use, so JIT and // serializer metadata are paid for here, not by a user. _ = JsonSerializer.Serialize(new { slug = "warmup", url = "https://example.com" });
state.MarkReady(); logger.LogInformation("Warm-up complete, instance is ready"); }}If warm-up throws, the instance does not stay up in a half-ready state. The default BackgroundServiceExceptionBehavior is StopHost: the host logs the exception and stops the application, and the process exits. On Cloud Run the instance fails to start, so it never passes the startup probe and never gets traffic. If it is a new revision, that revision does not become ready, the deploy fails, and the previous revision keeps serving. That is the behaviour you want: a broken instance never receives traffic, and the reason is in the logs.
The endpoints
Section titled “The endpoints”ASP.NET Core has a health checks library built in. It gives you the right status codes for free (200 for healthy, 503 for unhealthy) and a clean way to choose which checks run on which endpoint, using tags.
using Microsoft.Extensions.Diagnostics.HealthChecks;
namespace Sniplink.Api.Lifecycle;
public sealed class StartupHealthCheck(StartupState state) : IHealthCheck{ public Task<HealthCheckResult> CheckHealthAsync( HealthCheckContext context, CancellationToken cancellationToken = default) => Task.FromResult(state.IsReady ? HealthCheckResult.Healthy() : HealthCheckResult.Unhealthy("warming up"));}And in Program.cs:
using Microsoft.AspNetCore.Diagnostics.HealthChecks;using Sniplink.Api.Lifecycle;// …
builder.Services.AddSingleton<StartupState>();builder.Services.AddHostedService<WarmupService>();builder.Services.AddHealthChecks() .AddCheck<StartupHealthCheck>("startup", tags: ["ready"]);
// …var app = builder.Build();
app.MapGet("/health", () => Results.Ok(new { status = "ok" })); // module 1, kept: labs 1-3 and 6 call it
// Live: no checks at all. If the pipeline can answer, the process is alive.app.MapHealthChecks("/health/live", new HealthCheckOptions{ Predicate = _ => false,});
// Ready: only checks tagged "ready".app.MapHealthChecks("/health/ready", new HealthCheckOptions{ Predicate = check => check.Tags.Contains("ready"),});// …Predicate = _ => false looks odd but it is the documented pattern for liveness: run zero checks and report healthy. If the request pipeline can produce that response, the process is alive.
Run it locally and call both endpoints right after start. The warm-up reads the lab key, and there is no /secrets/lab-key on your machine, so point LabSecret:Path at a dummy key file first, as in module 3. Without it the warm-up throws and the app stops, exactly as described above. If your warm-up is fast you will only ever see 200; add a await Task.Delay(5000, stoppingToken) to the warm-up temporarily and you will see /health/ready return 503 for five seconds while /health/live returns 200.
printf 'local-dummy-key' > /tmp/lab-keyLabSecret__Path=/tmp/lab-key dotnet run --project src/Sniplink.Api# in a second terminalcurl -i http://localhost:8080/health/livecurl -i http://localhost:8080/health/readyWhat not to put in health checks
Section titled “What not to put in health checks”Health checks are called often, by machines, with no human looking. A few rules I follow:
- No external dependencies in liveness. If liveness checks the database and the database blips, every instance fails liveness at the same moment and Cloud Run restarts all of them. You turned a database hiccup into a full outage.
- Be deliberate about dependencies in readiness. On Cloud Run the startup probe only runs at start. A dependency check there means “do not start new instances while the dependency is down”, which also blocks deploys and scale-out. Sometimes right, often not.
- No expensive work. A health check that takes 800 ms of CPU is a self-inflicted load test.
- No secrets, versions or internals in the response. The default health checks output is just
HealthyorUnhealthy. Keep it that way;/healthis public on a public service. - No authentication. The probes come from Cloud Run and do not carry your auth headers.
- No Information-level logging per probe. You will pay to store thousands of “healthy” lines per day.
Wiring the probes and limits in Pulumi
Section titled “Wiring the probes and limits in Pulumi”Now tell Cloud Run about the endpoints and set the scaling limits for this module. All of these live in the revision template, so changing them creates a new revision.
The values, and why:
- Startup probe on
/health/ready, every 2 s, up to 15 failures. That is a 30-second budget, ten times what a warm chiseled .NET API with boost usually needs. A 1-second timeout is enough, because the endpoint does no work. - Liveness probe on
/health/live, every 15 s, 3 failures. Restarting a healthy but busy instance under load makes overload worse, so liveness should be slow to give up. - Concurrency 10 and max instances 3: the lab requirements, chosen so the numbers are small enough to observe. We discuss real values below.
- Min instances 0: scale to zero, no idle cost.
- CPU only during requests and startup CPU boost: explained after the code.
LABKIT_WORK_PROBE=trueswitches on LabKit’s/_lab/workendpoint, which the lab uses to measure in-flight requests per instance. It is for this lab only; you remove it again in the clean-up.
using Pulumi.Gcp.CloudRunV2.Inputs;// …
var service = new Gcp.CloudRunV2.Service("sniplink-api", new(){ Name = "sniplink-api", Location = region, // … ingress, deletion protection from module 1 Template = new ServiceTemplateArgs { ServiceAccount = runtimeSa.Email, // module 3
MaxInstanceRequestConcurrency = 10, Scaling = new ServiceTemplateScalingArgs { MinInstanceCount = 0, MaxInstanceCount = 3, },
Containers = { new ServiceTemplateContainerArgs { Image = image, // … ports and the /secrets volume mount from modules 1 and 3
Resources = new ServiceTemplateContainerResourcesArgs { Limits = { { "cpu", "1" }, { "memory", "512Mi" } }, CpuIdle = true, // request-based billing StartupCpuBoost = true, },
Envs = { new ServiceTemplateContainerEnvArgs { Name = "LABKIT_STUDENT_ID", Value = studentId, }, new ServiceTemplateContainerEnvArgs { Name = "LABKIT_WORK_PROBE", Value = "true", }, },
StartupProbe = new ServiceTemplateContainerStartupProbeArgs { HttpGet = new ServiceTemplateContainerStartupProbeHttpGetArgs { Path = "/health/ready", }, InitialDelaySeconds = 0, PeriodSeconds = 2, TimeoutSeconds = 1, FailureThreshold = 15, },
LivenessProbe = new ServiceTemplateContainerLivenessProbeArgs { HttpGet = new ServiceTemplateContainerLivenessProbeHttpGetArgs { Path = "/health/live", }, PeriodSeconds = 15, TimeoutSeconds = 2, FailureThreshold = 3, }, }, }, // … Volumes from module 3 },});resource "google_cloud_run_v2_service" "sniplink_api" { name = "sniplink-api" location = var.region # … ingress, deletion_protection from module 1
template { service_account = google_service_account.runtime.email max_instance_request_concurrency = 10
scaling { min_instance_count = 0 max_instance_count = 3 }
containers { image = var.image # … ports and the /secrets volume mount from modules 1 and 3
resources { limits = { cpu = "1" memory = "512Mi" } cpu_idle = true # request-based billing startup_cpu_boost = true }
env { name = "LABKIT_STUDENT_ID" value = var.student_id }
env { name = "LABKIT_WORK_PROBE" value = "true" }
startup_probe { http_get { path = "/health/ready" } initial_delay_seconds = 0 period_seconds = 2 timeout_seconds = 1 failure_threshold = 15 }
liveness_probe { http_get { path = "/health/live" } period_seconds = 15 timeout_seconds = 2 failure_threshold = 3 } } # … volumes from module 3 }}The module 3 service already has one entry in Envs, LABKIT_STUDENT_ID from module 1, and LABKIT_WORK_PROBE goes next to it; module 5 adds APP_VERSION and module 6 adds COMMIT_SHA and BUILD_ID. Run pulumi preview first and read the diff: you should see changes only in the template.
A useful side effect: once a startup probe exists, a revision whose instances never become ready fails to deploy, and pulumi up fails with it. Traffic stays on the previous revision. A broken build now costs you a red pipeline, not an outage.
CPU allocation and startup boost
Section titled “CPU allocation and startup boost”Cloud Run has two billing modes, and they map to how CPU is allocated.
- Request-based billing (
CpuIdle = true). CPU is allocated only while the instance is handling requests, and you are billed for that time, rounded to small increments. Between requests the CPU is throttled close to zero. This is the right default for an API with spiky or low traffic, which is most course projects and a lot of real ones. - Instance-based billing (
CpuIdle = false). CPU is always allocated for the whole lifetime of the instance, and you pay for that whole lifetime, at a lower unit price. This makes sense when traffic is steady enough that instances are rarely idle, or when the app does real work outside requests: background queues, timers, a hosted service that keeps processing after the response is sent.
The trap is the second bullet in reverse. With request-based billing, a BackgroundService that keeps working after the response is sent gets almost no CPU. It does not crash; it just becomes very slow. Sniplink’s warm-up is fine because it runs during startup, which is CPU-allocated.
Startup CPU boost gives the instance extra CPU while it starts and for a short time after, then returns to the configured limit. .NET startup is CPU-bound (runtime load, JIT, DI), so boost shortens cold starts noticeably, and you pay for the extra CPU only during those seconds. For a .NET service on Cloud Run I turn it on by default.
Graceful shutdown in ASP.NET Core
Section titled “Graceful shutdown in ASP.NET Core”Now the other end of the lifecycle: the connection resets in the incident.
When Cloud Run sends SIGTERM, the .NET host catches it and starts an orderly shutdown. In order:
IHostApplicationLifetime.ApplicationStoppingfires. Callbacks registered on it run.- The host stops hosted services in reverse order, including the web server. Kestrel stops accepting new connections, closes idle keep-alive connections, and lets in-flight requests run to completion.
- If in-flight requests are still running when
HostOptions.ShutdownTimeoutelapses, the stop token is cancelled and Kestrel aborts the remaining connections. ApplicationStoppedfires, the host is disposed (which flushes the console logger), and the process exits.
This already works out of the box. The problem is one number: the default ShutdownTimeout is 30 seconds, and Cloud Run gives you 10. With the default, a slow request keeps the host waiting past second 10, SIGKILL arrives, and everything still in memory is lost: the response, the log lines buffered for it, the ApplicationStopped callbacks. The client sees a reset connection.
So set the timeout below the grace period, with headroom for what comes after draining:
// …// Cloud Run sends SIGKILL 10 s after SIGTERM. Drain for up to 8 s,// keep 2 s for ApplicationStopped callbacks, log flush and process exit.builder.Services.Configure<HostOptions>(o => o.ShutdownTimeout = TimeSpan.FromSeconds(8));
// …var app = builder.Build();
app.Lifetime.ApplicationStopping.Register(() => app.Logger.LogInformation("SIGTERM received, draining in-flight requests"));
app.Lifetime.ApplicationStopped.Register(() => app.Logger.LogInformation("Shutdown complete"));// …Why 8 and not 9.9: the timeout covers only the hosted services stop phase. After it, the host still runs ApplicationStopped callbacks, disposes services and flushes logs. Two seconds is a comfortable margin; I have never needed more for an API this size.
Three rules for the callbacks:
- Keep them short and synchronous. They run on the shutdown path. A callback that makes an HTTP call has just spent your grace period.
- Use
ApplicationStoppingto stop taking work (stop pulling from a queue, flip a flag) and to log that shutdown started. That log line is the most useful one when you investigate a deploy incident. - Use
ApplicationStoppedto flush anything with a buffer: a telemetry exporter, a batch writer. Built-in console logging to stdout is flushed on host disposal; Cloud Run collects stdout into Cloud Logging.
The other half of graceful shutdown is not code: requests must be shorter than the drain window. If an endpoint can run for 20 seconds, no setting will drain it in 8. For Sniplink every endpoint is milliseconds, and LabKit caps /_lab/work at 3 seconds for exactly this reason. Long work belongs in a job or a queue, not in a request.
Concurrency: how many requests per instance
Section titled “Concurrency: how many requests per instance”Concurrency on Cloud Run is the maximum number of requests a single instance handles at the same time. The default is 80. When all instances are at their limit, the autoscaler starts another one, up to the maximum instance count.
What 80 means for .NET
Section titled “What 80 means for .NET”For an ASP.NET Core API that is mostly I/O-bound, 80 is usually fine and often conservative. A request waiting on await httpClient.GetAsync(…) or a database call does not hold a thread; it is a small state machine on the heap. One vCPU can juggle hundreds of such requests, because most of the time nobody is using the CPU.
Two things change the picture:
- CPU-bound work. Image resizing, hashing big payloads, PDF generation, heavy JSON transformations. With 1 vCPU and 80 such requests at once, each gets roughly 1/80 of a core, and every one of them is slow. Latency for everyone goes up together. Lower concurrency makes Cloud Run spread the work across more instances instead, so each request gets a real share of CPU.
- Blocking code.
.Result,.Wait(), synchronous file or network I/O. Each blocked request holds a thread-pool thread, and the pool grows slowly under pressure. With high concurrency and sync-over-async, you get thread-pool starvation: low CPU, high latency, and timeouts that look like the network’s fault. The fix is async all the way down, not lower concurrency, though lowering it hides the symptom.
Memory scales with concurrency too: 5 MB per request at peak times 80 is 400 MB.
My rule of thumb: keep 80 for I/O-bound .NET APIs, measure, and go lower only when CPU per request is significant or a dependency (a connection pool, a rate-limited API) cannot take that many parallel calls. Concurrency 1 is almost never right for .NET.
Capacity and what happens on overflow
Section titled “Capacity and what happens on overflow”Two settings define the total ceiling:
max concurrent requests = concurrency × max instances = 10 × 3 = 30 (this module) = 80 × 100 = 8,000 (Cloud Run defaults)Combine that with request latency and you get throughput. With the lab’s /_lab/work?ms=1500, each request holds a slot for 1.5 s, so 30 slots give about 20 requests per second at most. For Sniplink’s real endpoints, which take a few milliseconds, the same 30 slots go a long way.
When every slot is busy and no new instance can be started, Cloud Run does not fail the request immediately. It holds it in a pending queue, hoping a slot frees up or an instance becomes ready, for up to 10 seconds or 3.5 times the service’s average instance startup time, whichever is longer. If that does not happen in time, the request is rejected with 429 Too Many Requests (“no available instance” in the docs). If instances fail to start at all you see 500 instead, and a spike that Cloud Run cannot scale for in time can also produce 500s. Your clients should treat these as retryable with backoff.
Max instances is a strong guard, not an exact guarantee: the docs note it can be briefly exceeded in some situations.
Max instances as a cost guard
Section titled “Max instances as a cost guard”The incident log scaled to 58 instances because nothing told it to stop. On a public URL, that is not only a marketing email: it is also a crawler, a misbehaving client retry loop, or someone testing a load tool on your service.
MaxInstanceCount turns “the bill can grow with traffic” into “the bill has a ceiling”. Worst case with this module’s settings: 3 instances × 1 vCPU × 512 MiB, all busy, all the time. You can put a number on that from the pricing page, and it fits under the budget you created in module 0. The cost of the guard is that beyond the ceiling, users see 429s instead of a bigger bill. For a course project that is clearly the right trade. For a production service, set the limit from expected peak times a safety factor, and alert when you get close to it.
Min instances and their price
Section titled “Min instances and their price”MinInstanceCount = 1 keeps one instance warm at all times, so the first request after a quiet period does not pay a cold start. It also means you pay for that instance 24 hours a day. With request-based billing an idle min instance is billed at a reduced idle rate; with instance-based billing it is billed like any other instance. Either way it is a constant line on the bill that does not care whether anyone used the service.
I keep MinInstanceCount = 0 for course projects and internal tools, and set it to 1 or more only when a cold start would break a latency target someone actually measures.
Measuring it
Section titled “Measuring it”Settings you have not watched in action are guesses. You need a load generator on your laptop and the Cloud Run metrics page, read-only.
Install either oha or hey:
# oha (Rust): brew install oha, orcargo install oha# hey (Go)go install github.com/rakyll/hey@latestDeploy the changes and get the URL:
export PROJECT_ID=your-course-project-idexport REGION=europe-west1cd infrapulumi upexport SERVICE_URL=$(pulumi stack output url) # the output name from module 1cd ..Now reproduce what the lab does: 30 concurrent clients, each asking the work probe to hold a request for 1.5 s.
# 300 requests, 30 at a timeoha -n 300 -c 30 "$SERVICE_URL/_lab/work?ms=1500"# or with heyhey -n 300 -c 30 "$SERVICE_URL/_lab/work?ms=1500"Check the status code distribution (mostly 200, maybe some 429 while instances are cold) and the latency histogram (about 1.5 s plus network, with a cold-start tail).
Then push past the ceiling and watch the guard work:
# 60 concurrent clients against 30 slotsoha -n 300 -c 60 "$SERVICE_URL/_lab/work?ms=1500"You should see a mix of 200 and 429 responses. That is the cost guard doing its job.
Finally, look at a single response to see what LabKit reports:
curl -s "$SERVICE_URL/_lab/work?ms=200"# {"instanceId":"…","inFlight":1,"maxInFlightSeen":10}maxInFlightSeen is the highest number of concurrent work requests that instance has seen in the last 60 seconds. With concurrency 10 it should never go above 10, however hard you push.
In the console, open Cloud Run → sniplink-api → Metrics and set the range to the last hour. The useful charts:
- Container instance count, split into active and idle. You should see it climb from 0 to 3 under load and fall back to 0 some minutes after you stop.
- Max concurrent requests per instance. It flattens at 10.
- Request count by response code class. Your 429s show up here.
- Request latencies and container startup latency. The second one is your cold start, and the number to compare when you toggle startup boost.
If you want to see graceful shutdown pay off, run a long, gentle load while you deploy a trivial change (for example, add an env var DEPLOY_MARK=1 in Pulumi):
oha -z 90s -c 5 "$SERVICE_URL/_lab/work?ms=500"# in another terminal, during the run:cd infra && pulumi up --yesWith the probes and the 8-second shutdown timeout in place, you should see no errors, or a handful at most. Then search the logs for SIGTERM received and Shutdown complete, and check that the gap between them is well under 10 seconds.
Goal: Sniplink runs with explicit probes, a concurrency of 10, at most 3 instances, and the LabKit work probe enabled, and it proves the limits hold under concurrent load.
- Add
StartupState,WarmupServiceandStartupHealthCheck, and map/health/liveand/health/readyas shown. - Set
HostOptions.ShutdownTimeoutto 8 seconds and log fromApplicationStoppingandApplicationStopped. - Rebuild and push the image (module 2 commands), then update
infra/with the startup probe, liveness probe,MaxInstanceRequestConcurrency = 10,MinInstanceCount = 0,MaxInstanceCount = 3, explicit resources withCpuIdleandStartupCpuBoost, andLABKIT_WORK_PROBE=true. - Run
pulumi upand confirm both health endpoints return 200 withcurl. - Run the
ohaorheytest above and checkmaxInFlightSeenstays at or below 10. - Paste your service URL below and run the checks.
Check your lab
What the checks verify:
- Token (baseline).
POST /_lab/verifyreturns a valid Google ID token for your project, as in every lab. - Liveness and readiness.
GET /health/liveandGET /health/readyboth return 200. - Work probe.
GET /_lab/work?ms=100returns 200, which meansLABKIT_WORK_PROBE=truereached the container. - Concurrency. The platform first sends one warm-up request, then 30 concurrent
GET /_lab/work?ms=1500. Every successful response must reportmaxInFlightSeenof 10 or less, the responses must come from at most 3 distinctinstanceIdvalues, and at least 20 of the 30 must succeed. Overflow requests answered with 429 or 503 are accepted; they are the limit working, not a failure. If a rare autoscaler overshoot gives you a fourth instance, run the check again.
The concurrency check observes behaviour, not configuration: concurrency 80 would show maxInFlightSeen far above 10, and no instance limit could spread over more than 3 instances. The counters come from LabKit inside your container, so they are self-reported in the narrow sense, but hard to fake by accident.
Self-check (not automated). Confirm these yourself before you move on:
- The Pulumi program defines a startup probe on
/health/readyand a liveness probe on/health/live. -
HostOptions.ShutdownTimeoutis set and is below 10 seconds. - Logs from a deploy show
SIGTERM receivedfollowed byShutdown completeon the old revision.
These are not automated because nothing outside your project can see them. Probe configuration lives in the Cloud Run service spec, which only the Verified tier will be able to read. The shutdown timeout is a value inside your process, and the only external symptom of getting it wrong is a few failed requests during a deploy, which is probabilistic and would make the check flaky. An honest self-check is better than a check that passes or fails by luck.
Clean up
Section titled “Clean up”Turn the work probe off once the lab has passed. /_lab/work lets anyone on the internet hold a request open for up to 3 seconds, and nothing after this module needs it; the final project switches it on again for its own lab. Delete the LABKIT_WORK_PROBE entry from Envs (the env block in Terraform) and run pulumi up. The module 5 and 6 programs no longer set it.
Keep the stack. Module 5 builds releases and traffic splitting on top of this exact service, and the probes you added here are what make a canary revision safe to route traffic to.
With MinInstanceCount = 0, the service costs nothing while idle; what remains are small storage charges for Artifact Registry, Secret Manager versions and the KMS key. If you are pausing the course for a while, destroy the stack. To continue with module 5, run pulumi up (it recreates the repository and the secret; the service fails because the images went with the repository), push the image again with the module 2 commands, set sniplink:image to the new digest and run pulumi up again:
cd infrapulumi destroyWhat you learned
Section titled “What you learned”- A Cloud Run instance goes through start, startup probe, serving, idle, SIGTERM, a fixed 10-second grace period and SIGKILL, and cold starts happen on every scale-out and deploy, not only after scale to zero.
/health/liveshould depend on nothing but the process;/health/readyshould reflect real warm-up, and the startup probe on it keeps cold instances and broken revisions away from traffic.HostOptions.ShutdownTimeoutmust be below Cloud Run’s grace period, so Kestrel can drain in-flight requests and the host can flush before SIGKILL.- Concurrency × max instances is your capacity ceiling; async .NET handles the default 80 well, CPU-bound or blocking code does not, and
MaxInstanceCountturns traffic spikes into 429s instead of surprise bills. - Request-based billing, startup CPU boost and
MinInstanceCountare cost decisions with latency trade-offs, and you can measure each one withohaand the Cloud Run metrics page.