Keyless Google Cloud access from a Kubernetes cluster Google can't reach
Workload Identity Federation for a private cluster in a Saudi sovereign cloud, with the OIDC issuer served from a Cloudflare Worker. No service account keys, and nothing on the cluster exposed.
Earlier this year I ran the platform for a B2B SaaS startup that needed a production copy of its app inside Saudi Arabia, for data residency. The cluster lived in a Saudi sovereign cloud. Everything it depended on (Secret Manager, Artifact Registry, Pub/Sub, backups) lived in Google Cloud, so pods in Saudi Arabia had to call Google APIs.
Two rules took the easy answers away:
- The organization enforced
iam.disableServiceAccountKeyCreation. No service account JSON keys, anywhere. A good rule, and it removed the lazy option. - The cluster was not reachable from the internet. Nothing could call into it.
This is how pods in that cluster got short-lived Google credentials anyway, with no key stored anywhere.
How the pieces fit
Kubernetes already hands out identities. Every pod can get a projected service account token: a JWT signed by the cluster, with an issuer (iss), a subject such as system:serviceaccount:backend:api, and an audience.
Workload Identity Federation lets Google trust that token. The pod sends it to Google's Security Token Service (STS), which checks the signature against the cluster's public keys and returns a short-lived federated token. The pod can then act as a Google service account that holds the actual permissions.
pod in the private cluster Google Cloud
-------------------------- ------------
projected token (iss=https://oidc.example.com,
sub=system:serviceaccount:backend:api,
aud=<provider>)
|
+--> STS sts.googleapis.com ------------> GET oidc.example.com/.well-known/...
| checks signature GET oidc.example.com/openid/v1/jwks
| <-- federated token (Cloudflare Worker, public)
|
+--> IAM Credentials: generateAccessToken
| <-- 1h access token for api@my-prod-project
|
+--> Secret Manager, Pub/Sub, Artifact Registry, ...There is one hard requirement: Google has to get the cluster's public keys, the JWKS, from somewhere.
Three ways to give Google the keys
- A public issuer URL. Managed clusters such as EKS and AKS publish one and Google fetches the keys itself. Ours was private.
- Upload the JWKS into the provider. This is Google's documented path for self-hosted clusters (
--jwk-json-path, orjwks_jsonin OpenTofu). It works. Google's docs warn that the command doesn't validate the JWKS, so a bad file only shows up later as failed logins. - Host the discovery document yourself at a public URL, and point the cluster's issuer at it. The common recipe is a public Cloud Storage bucket. This organization also enforced
storage.publicAccessPrevention, so that was out.
I went with the third option and served it from a Cloudflare Worker instead of a bucket, with growth in mind: any other service that ends up running outside Google Cloud can get its own issuer on the same Worker, and anything that speaks OIDC can trust it, not only Google. Nothing on the cluster is exposed either: Google fetches the keys from Cloudflare, never from the cluster.
The Worker
It serves two static JSON documents and nothing else. The field names are the ones Kubernetes itself puts in its discovery document.
// Public OIDC discovery for a Kubernetes cluster nobody can reach.
// Both documents are copies from the cluster; only these URLs are public.
const ISSUER = 'https://oidc.example.com';
// From: kubectl get --raw /.well-known/openid-configuration
// with issuer and jwks_uri pointed at this Worker.
const discovery = {
issuer: ISSUER,
jwks_uri: `${ISSUER}/openid/v1/jwks`,
response_types_supported: ['id_token'],
subject_types_supported: ['public'],
id_token_signing_alg_values_supported: ['RS256'], // copy from your cluster
};
// From: kubectl get --raw /openid/v1/jwks
const jwks = { keys: [ /* the cluster's public signing keys */ ] };
export default {
fetch(request) {
const { pathname } = new URL(request.url);
if (pathname === '/.well-known/openid-configuration') return json(discovery);
if (pathname === '/openid/v1/jwks') return json(jwks);
return new Response('not found', { status: 404 });
},
};
const json = (body) => new Response(JSON.stringify(body), {
headers: { 'content-type': 'application/json', 'cache-control': 'public, max-age=300' },
});On the cluster, the API server's issuer has to match the Worker's URL, or the tokens it signs will name the wrong issuer. --service-account-jwks-uri makes the cluster's own discovery document point at the public key URL too.
--service-account-issuer=https://oidc.example.com # signs new tokens
--service-account-issuer=https://old-issuer.internal # still accepted
--service-account-jwks-uri=https://oidc.example.com/openid/v1/jwksBe careful switching the issuer on a running cluster: every token already issued names the old one. The flag can be given more than once. The first value signs new tokens and all of them are accepted, so list the new URL first and keep the old one until existing tokens have expired.
The org policy nobody mentions
This organization denied Workload Identity providers by default: iam.workloadIdentityPoolProviders set to deny all, with per-project allowlists of trusted issuers (GitHub Actions, GitLab and Terraform Cloud for CI). Pointing a provider at a new issuer means adding that URL to the project's allowlist first. If your organization was set up with Google's Cloud Foundation Fabric or a similar security baseline, check for this before debugging anything else.
The provider
resource "google_iam_workload_identity_pool" "sov" {
workload_identity_pool_id = "sov-k8s"
}
resource "google_iam_workload_identity_pool_provider" "sov" {
workload_identity_pool_id = google_iam_workload_identity_pool.sov.workload_identity_pool_id
workload_identity_pool_provider_id = "sov-k8s-oidc"
attribute_mapping = {
"google.subject" = "assertion.sub"
"attribute.namespace" = "assertion['kubernetes.io']['namespace']"
"attribute.service_account" = "assertion['kubernetes.io']['serviceaccount']['name']"
}
# Only Kubernetes service accounts, nothing else this issuer might sign.
attribute_condition = "assertion.sub.startsWith('system:serviceaccount:')"
oidc {
issuer_uri = "https://oidc.example.com"
}
}
# The pod's Kubernetes service account may act as this Google one.
resource "google_service_account_iam_member" "api" {
service_account_id = google_service_account.api.name
role = "roles/iam.workloadIdentityUser"
member = "principal://iam.googleapis.com/${google_iam_workload_identity_pool.sov.name}/subject/system:serviceaccount:backend:api"
}Two details matter here. The attribute condition limits the provider to Kubernetes service accounts. And with no allowed_audiences, Google requires the token's audience to be the provider's full resource name, //iam.googleapis.com/projects/…/providers/sov-k8s-oidc. That string shows up everywhere below.
Three consumers, one mechanism
Secrets: External Secrets Operator
ESO mints a token for its own Kubernetes service account, exchanges it, and acts as a Google service account that can read Secret Manager. One trick paid off: the store has the same name as on the GKE clusters, so every Helm chart that references it worked unchanged.
apiVersion: external-secrets.io/v1
kind: ClusterSecretStore
metadata:
name: gcp-secret-manager # same name as on GKE, so no chart changes
spec:
provider:
gcpsm:
projectID: my-prod-project
auth:
workloadIdentityFederation:
audience: //iam.googleapis.com/projects/123456789/locations/global/workloadIdentityPools/sov-k8s/providers/sov-k8s-oidc
gcpServiceAccountEmail: [email protected]
serviceAccountRef:
name: external-secrets
namespace: external-secrets
audiences:
- //iam.googleapis.com/projects/123456789/locations/global/workloadIdentityPools/sov-k8s/providers/sov-k8s-oidcImages: Artifact Registry without node credentials
Outside GKE the kubelet has no Google identity, so image pulls need a pull secret. ESO's GCRAccessToken generator mints short-lived registry tokens, and a ClusterExternalSecret writes them as a dockerconfigjson secret into every application namespace.
That identity can only read Artifact Registry, which is what makes it safe to copy into every namespace. It can't reach Secret Manager or anything else.
The one thing that bit me: the tokens came back with about 30 minutes to live, not the hour I had assumed. With my first refresh interval a pull secret could expire before its replacement arrived, and image pulls would fail with 401. Refreshing every 10 minutes closed that gap.
apiVersion: generators.external-secrets.io/v1alpha1
kind: ClusterGenerator
metadata:
name: registry-token
spec:
kind: GCRAccessToken
generator:
gcrAccessTokenSpec:
projectID: my-registry-project
auth:
workloadIdentityFederation:
audience: //iam.googleapis.com/projects/123456789/locations/global/workloadIdentityPools/sov-k8s/providers/sov-k8s-oidc
serviceAccountRef:
name: default
audiences:
- //iam.googleapis.com/projects/123456789/locations/global/workloadIdentityPools/sov-k8s/providers/sov-k8s-oidc
---
apiVersion: external-secrets.io/v1
kind: ClusterExternalSecret
metadata:
name: registry-pull-secret
spec:
externalSecretName: registry-pull-secret
namespaceSelector:
matchExpressions:
- key: kubernetes.io/metadata.name
operator: NotIn
values: [kube-system, kube-public, kube-node-lease, external-secrets]
refreshTime: "10m" # tokens came back with ~30 min to live, not an hour
externalSecretSpec:
refreshInterval: "10m"
target:
name: registry-pull-secret
template:
type: kubernetes.io/dockerconfigjson
data:
.dockerconfigjson: |
{"auths":{"europe-west1-docker.pkg.dev":{"auth":"{{ printf "%s:%s" .username .password | b64enc }}"}}}
dataFrom:
- sourceRef:
generatorRef:
apiVersion: generators.external-secrets.io/v1alpha1
kind: ClusterGenerator
name: registry-tokenApplication code: plain Google client libraries
For the apps themselves, Google's client libraries already understand an external_account credential file. Mount a projected token with the provider as its audience, mount the credential config, set GOOGLE_APPLICATION_CREDENTIALS, and the same code that runs on GKE authenticates here. In the shared Helm chart this became one flag.
# in the Deployment's pod spec
serviceAccountName: api
containers:
- name: api
env:
- name: GOOGLE_APPLICATION_CREDENTIALS
value: /var/run/secrets/gcp/credential-configuration.json
volumeMounts:
- { name: gcp-token, mountPath: /var/run/secrets/tokens/gcp, readOnly: true }
- { name: gcp-credconfig, mountPath: /var/run/secrets/gcp, readOnly: true }
volumes:
- name: gcp-token
projected:
sources:
- serviceAccountToken:
path: token
audience: //iam.googleapis.com/projects/123456789/locations/global/workloadIdentityPools/sov-k8s/providers/sov-k8s-oidc
expirationSeconds: 3600
- name: gcp-credconfig
configMap:
name: gcp-credconfig{
"type": "external_account",
"audience": "//iam.googleapis.com/projects/123456789/locations/global/workloadIdentityPools/sov-k8s/providers/sov-k8s-oidc",
"subject_token_type": "urn:ietf:params:oauth:token-type:jwt",
"token_url": "https://sts.googleapis.com/v1/token",
"service_account_impersonation_url": "https://iamcredentials.googleapis.com/v1/projects/-/serviceAccounts/[email protected]:generateAccessToken",
"credential_source": {
"file": "/var/run/secrets/tokens/gcp/token",
"format": { "type": "text" }
}
}What did need attention was code that quietly assumed it ran inside Google Cloud, such as shipping logs and traces straight to Google's APIs. Those integrations went behind a separate profile for this cluster.
Debugging the chain by hand
When federation fails, the error rarely tells you which link broke. Walking the chain one step at a time does. Whichever step fails first is the broken link.
One error worth recognising: a typo anywhere in the audience makes STS answer invalid_target, "the pool or provider is disabled or deleted or … doesn't exist", even when the pool is fine.
AUD=//iam.googleapis.com/projects/123456789/locations/global/workloadIdentityPools/sov-k8s/providers/sov-k8s-oidc
# 1. Mint a token exactly as the pod would get it.
TOKEN=$(kubectl create token api -n backend --audience "$AUD" --duration 10m)
# 2. Read its claims: iss must be the Worker URL, aud the provider.
python3 -c 'import sys,json,base64; p=sys.argv[1].split(".")[1]; print(json.dumps(json.loads(base64.urlsafe_b64decode(p+"="*(-len(p)%4))),indent=1))' "$TOKEN"
# 3. Can Google fetch the keys? Same two requests STS makes.
curl -s https://oidc.example.com/.well-known/openid-configuration
curl -s https://oidc.example.com/openid/v1/jwks
# 4. Trade the Kubernetes token for a federated Google token.
FED=$(curl -s https://sts.googleapis.com/v1/token \
-d grant_type=urn:ietf:params:oauth:grant-type:token-exchange \
-d requested_token_type=urn:ietf:params:oauth:token-type:access_token \
-d subject_token_type=urn:ietf:params:oauth:token-type:jwt \
-d scope=https://www.googleapis.com/auth/cloud-platform \
--data-urlencode "audience=$AUD" \
--data-urlencode "subject_token=$TOKEN" | jq -r .access_token)
# 5. Use it to act as the Google service account.
curl -s -X POST -H "Authorization: Bearer $FED" -H 'Content-Type: application/json' \
-d '{"scope":["https://www.googleapis.com/auth/cloud-platform"]}' \
https://iamcredentials.googleapis.com/v1/projects/-/serviceAccounts/[email protected]:generateAccessTokenTrade-offs worth knowing
- The JWKS in the Worker is a copy. If the cluster's signing key rotates, the Worker has to be updated first, serving the old and new keys together until old tokens expire. Put that in the runbook, or let CI copy the keys over.
- The Worker is now in the login path. If it's down, no new tokens: secrets stop refreshing, and image pulls fail once the current pull tokens expire. Monitor it like any other dependency.
- If the cluster is a one-off and Google is its only consumer, uploading the JWKS into the provider is one less thing to run. We expected more services outside Google Cloud, so a real issuer was worth the extra moving part.
- Traffic from the pods to Google APIs crosses the public internet. It's TLS end to end, but check the egress bill and make sure your security review has seen it.