mTLS and Service Identity: Who Actually Proves Who
Mutual TLS answers a question most teams never ask precisely enough, and the gap between 'the certificate was valid' and 'this caller is allowed to do this' is where the interesting failures live.
A service mesh rollout was declared complete: every hop carried mTLS, the policy showed STRICT, and the compliance checklist was signed. Then a job with a valid certificate for the staging namespace read production data, because both certificates were issued by the same authority and the policy checked only that a certificate existed.
Mutual TLS answers one of those. It does not answer the other, and the gap is where the interesting failures live.
The two questions
Authenticity is what the certificate chain establishes: the peer's certificate was issued by a trusted authority, is within its validity window, and has not been revoked. If that passes, the peer is who the certificate says they are.
Authorization is a separate decision about whether that identity may perform this action on this resource. mTLS does not make it; something must read the identity out of the certificate and compare it against a rule.
Meshes blur the two because a single STRICT policy mode feels like both. It is only the first.
What the identity actually contains
The useful part of a peer certificate is not the subject line but the fields the mesh maps into a workload identity:
URI: cluster.local/ns/payments/sa/checkout-worker
SPIFFE ID: spiffe://cluster.local/ns/payments/sa/checkout-worker
That URI, a SPIFFE ID, is the identity the policy engine evaluates. It names the namespace, the service account, and therefore the workload. Two pods using the same service account are, to the mesh, the same identity, which is exactly why choosing service accounts per workload rather than per namespace matters more than most rollout plans assume.
The legacy alternative encodes identity in the subject distinguished name, typically with a spiffe-style string appended as a SAN. It works and is what most people see first, but policy engines match the SAN form, so a certificate that looks right in openssl x509 -text can still fail policy if the URI SAN is missing.
Making the policy mean something
Once the identity is visible, authorization becomes a rule. In a common mesh that looks like:
apiVersion: security.istio.io/v1
kind: AuthorizationPolicy
metadata:
name: checkout-to-payments
namespace: payments
spec:
selector:
matchLabels:
app: payments-api
action: ALLOW
rules:
- from:
- source:
principals:
- "cluster.local/ns/payments/sa/checkout-worker"
to:
- operation:
methods: ["POST"]
paths: ["/charges"]
The important properties: it is scoped to one workload, names a principal, and allows only a specific method on a specific path. A policy with no from block admits everyone who reached the mesh; one with no to block admits every action. Both are common when the goal was "turn on mTLS" rather than "express who may call this."
The default posture matters as much. With no policy defined, most meshes allow everything, so enabling STRICT mTLS without an AuthorizationPolicy gives you encrypted traffic between every service, including the ones that should never reach each other. Identity policy is one layer of a Kubernetes security baseline.
Where identity breaks in practice
Three failure patterns account for most of what goes wrong.
The certificate rotates but the connection pool does not. Long-lived gRPC or HTTP2 connections keep presenting the old identity until recycled, so a revoked workload can continue on an established session. Fix it at the connection layer: short certificate lifetimes plus connection draining, not only at issuance.
The trust domain is shared across environments. One cluster, one signing authority, and staging workloads receive certificates production trusts. Isolation needs separate trust domains or explicit cross-environment policy, neither of which arrives by default.
Identity is asserted but never logged. When a request is denied, the operator sees a 403 with no principal attached, and every authorization failure becomes an investigation instead of an answer.
The test that proves it
End-to-end verification, not config inspection:
# should succeed: identity allowed
curl -s --cert client.pem --key client-key.pem https://payments-api/charges -X POST
# should be rejected: same valid cert, wrong identity
curl -s --cert wrong-sa.pem --key wrong-sa-key.pem https://payments-api/charges -X POST
Both certificates are valid and unexpired. Only the second should fail, and if it does not, the policy is not doing what the checklist claims. Running this pair in CI after every policy change is the difference between a mesh that enforces and a mesh that encrypts.
Summary
Mutual TLS establishes authenticity, not authorization: a valid certificate proves who the caller is, not what they may do. Extract the identity from the URI SAN or SPIFFE ID, then write explicit allow rules scoped by principal, method, and path, remembering that absent policy means allow-all. Verify with a negative test using a valid certificate for the wrong identity; if that succeeds, you have encryption without a security boundary.
SDP Clouds Team
DevOps and cloud engineers writing practical, battle-tested guides on CI/CD, Kubernetes, infrastructure as code, and production operations: every article is based on real incidents and real pipelines, not docs-page rewrites.
More about us →