• Trivy Operator
  • Kubernetes
  • comparisons

Trivy Operator without the etcd problem: where vulnerability findings should live

Trivy Operator stores each report as one object in etcd, and a big image does not fit. How to find the reports that were silently dropped, and what each fix costs.

StackRadar Team

9 min read

If you run Trivy Operator on a cluster of any size, there is a fair chance some of your workloads have no vulnerability report at all, and nothing is telling you. The scan ran. The scan succeeded. Then the operator tried to save the result and the API server refused it:

trivy-operator log
"msg":"Reconciler error","controller":"job", ...
"error":"rpc error: code = ResourceExhausted desc =
  trying to send message larger than max (2633583 vs. 2097152)"

That line is from issue #757, open since December 2022. This post explains why it happens, how to find the reports your cluster has dropped, what each available fix costs, and why we built StackRadar to keep findings out of the cluster altogether.

The short answer: Trivy Operator stores each container's whole finding list as one Kubernetes object, and an object has to fit in a single request to etcd. Images with thousands of findings do not fit. You can trim the report, move reports to a volume, or store findings somewhere built for them. Each has a price, listed below.

One object per container, and a ceiling on the object

Trivy Operator writes its results as custom resources. A VulnerabilityReport holds every finding for one container: ID, package, installed and fixed version, severity, score, title, links. It is a tidy design, because the report is a Kubernetes object like any other: kubectl get reads it, RBAC guards it, and it is garbage-collected with the workload that owns it.

The cost is that Kubernetes objects live in etcd, and a write to etcd has to pass two size limits. etcd rejects any request over 1.5 MiB by default. That is the error in issue #441, etcdserver: request is too large. And the API server's etcd client will not send a message over 2 MiB in the first place. That is the 2097152 in the log line above.

etcd default, 1.5 MiBAPI server to etcd, 2 MiBLargest report, trimmedDescription and Links removed0.35 MBLargest of 104 reportsone cluster, September 20261.07 MBpython:3.5.10-buster1,252 findings, December 20222.63 MB, rejected0123Report size (MB)
Report sizes as measured in trivy-operator issue #757, against the two limits a write to etcd has to pass.

The image in #757 was python:3.5.10-buster, with 1,252 findings when the issue was filed. By January 2026 the same image had 3,752. That is the part that makes this a design problem and not a one-off: the size of a report is set by how many advisories exist for the packages in an image, and that number only goes up. An old image you never touch gets a larger report every month.

It is also not rare. In the last six weeks two people have written in that thread that they keep a blocklist of images the operator cannot store, one of them on a multi-tenant cluster where "a lot of images" hit the limit.

It fails open

The size limit would be an annoyance if it failed loudly. It does not. A report from September 2026, on operator v0.34.0, describes two workloads with no VulnerabilityReport object at all. Their ConfigAuditReport objects were created normally, scan jobs kept running, and nothing was piling up. On every dashboard the two workloads looked scanned and clean. They were the two most internet-exposed services in the cluster.

There is no trivy_* metric for a rejected write. The only counter that moves is the generic controller_runtime_reconcile_errors_total, with no image or workload label, shared with routine errors. So the missing report has to be found by its absence.

You can do that today. Trivy Operator labels every report with the resource that owns it and the container name, so the running containers without a report are one diff away:

Running containers that have no VulnerabilityReport
# Containers running right now, keyed the way Trivy Operator labels its reports
kubectl get pods -A -o json | jq -r '
  .items[] | select(.status.phase=="Running")
  | .metadata.namespace as $ns
  | (.metadata.ownerReferences[0] // {kind:"Pod", name:.metadata.name}) as $o
  | .spec.containers[]
  | "\($ns)/\($o.kind)/\($o.name)/\(.name)"' | sort -u > running.txt

# Containers that have a report
kubectl get vulnerabilityreports -A -o json | jq -r '
  .items[].metadata.labels
  | "\(.["trivy-operator.resource.namespace"])/\(.["trivy-operator.resource.kind"])/\(.["trivy-operator.resource.name"])/\(.["trivy-operator.container.name"])"' \
  | sort -u > reported.txt

comm -23 running.txt reported.txt

Each line it prints is a container that is not yet scanned, whose scan job failed, or whose report was rejected. Only the operator log tells the three apart, and the third one is this issue. Two things to know when reading the output: pods started by a CronJob are reported under the CronJob, not the Job, so they print here even when a report exists; and init containers are left out (add .spec.initContainers[] if you scan them). Run it from a CronJob, alert when the same line shows up two days running, and you have the metric the operator lacks.

The other etcd problem: total volume

The per-object ceiling is the sharp edge. The blunt one is the sum. One report per container, per report type (vulnerabilities, SBOM, config audit, exposed secrets), adds up to a lot of large objects in the store that also holds every Pod, Secret and Lease in the cluster.

The worst account of it is issue #2335, from November 2024: an OKD cluster with about 1,200 pods, with the config, RBAC and compliance scanners already switched off. etcd grew from roughly 380 MB to 1.2 GiB, leader changes started, the API went down, and the cluster was restored from an etcd backup. etcd's default storage quota is 2 GiB. Most clusters will not get near that, but the direction is the point: etcd is the control plane's memory, and scan output is the one thing in it that grows on its own.

The fixes available today, and what each costs

FixWhat it buysWhat it costs
Blocklist the imageThe scan-job queue keeps movingThe images with the most findings are the ones you stop scanning
Filter the report (trivy.ignoreUnfixed, trivy.severity)Fewer findings per report, often enough to fitThe filtered findings are gone from every report, not just the oversized one
Drop fields per finding (PRs #2854, #2860)About 73% smaller with Description and Links removed, by one measurementNot merged: both PRs have been open since January 2026
Raise the limitsHeadroomNot possible on a managed control plane
alternateReportStorageReports leave etcd for JSON files on a volume; no object ceilingNo kubectl, no vulnerability metrics, and files are never cleaned up
webhookBroadcastURLEvery report is POSTed to an endpoint of yoursYou build and run the receiver, the store and the queries

Three of these deserve more than a table cell.

Dropping fields works better than it sounds. The author of the two PRs measured the 3,752-finding image with only the ID and severity kept: 205 kB, a tenth of the limit. But a report with only IDs and severities no longer says which package to upgrade or to what version, so the detail has to be looked up somewhere else. It moves the ceiling a long way off. It does not remove it, and it is not in a release.

Raising the limits is two limits. On a control plane you run yourself, etcd's --max-request-bytes can be raised. The 2 MiB limit is a default of the etcd client the API server is built with, and Kubernetes has no setting for it. On EKS, GKE or AKS neither is yours to change.

Alternate storage is the real fix, half finished. Since 2025 the operator can write reports as JSON files to a persistent volume instead of to etcd. People who hit the etcd problem depend on it. Two open issues come with it. The operator's metrics are built from the CRDs, so with no CRDs trivy_image_vulnerabilities and the other report metrics disappear, and the Grafana dashboard goes blank (#2610, open since June 2025). And Kubernetes garbage collection deleted old CRDs for free, while nothing deletes old files, so the volume grows until it is full (#2791, open since October 2025). The workaround in that thread is a CronJob that deletes files older than the report TTL.

REPORT CRDS IN ETCDOne object per containerFinding list inside it2 MiB ceiling per objectJSON FILES ON A PVCOne file per reportNo per-object ceilingNo metrics, no cleanup yetROWS OFF-CLUSTEROne SBOM per image digestOne row per findingNothing written to etcd
Trivy Operator's two storage modes, and the off-cluster shape StackRadar uses.

A list of findings is a database workload

Step back from the individual issues and they are one issue. etcd is built to hold the desired state of a cluster: small objects, read whole, watched for changes. A finding list is a different kind of data in three ways.

  • It grows without you doing anything. Cluster state changes when you deploy. Findings change when someone publishes an advisory, which is every day.
  • The questions cut across objects. "Which containers run this CVE?" is a query over every report. With CRDs that means fetching all of them and filtering on the client.
  • The old data is the valuable part. "What were we running on the day of the incident?" is about a ReplicaSet that no longer exists. Garbage collection deletes exactly that report, and it has to, because etcd cannot afford to keep it.

None of this is a criticism of Trivy. The scanner is excellent, and storing results as CRDs is what makes the operator install with one chart and no dependencies. It is a trade, and the bill arrives at scale.

How StackRadar stores findings

We build StackRadar, so weigh this section accordingly. It makes the opposite trade: nothing is stored in the cluster, and a hosted service holds the state.

  • The agent writes nothing to your cluster. Its ClusterRole is list and watch on pods, get on the objects that own them, and list on Argo CD applications. It has no create, update, patch or delete on anything, and the chart installs no CRDs. You can read the role in the published source.
  • What leaves is an SBOM, once per image. For each new image digest the agent builds a CycloneDX SBOM with Syft and uploads it. An SBOM lists packages, not findings, so its size depends on what is in the image and not on how many advisories exist.
  • A finding is a row. On the server, an advisory is stored once and a finding is a row linking one package in your image to it. The 3,752-finding image is 3,752 small rows. There is no object for them to outgrow, and "which containers run this CVE?" is an indexed lookup across every cluster.
  • New advisories are matched without a rescan. The server follows OSV.dev's change feed hourly and matches new advisories against the SBOMs it already holds. No scan job runs in your cluster and no vulnerability database is downloaded to it.
  • History is kept on purpose. Findings for an image you stopped running stay for your plan's history window (30 days on Free, up to two years), then are deleted on a schedule.

What you give up

The trade has a price too, and for some teams it is the wrong one.

  • Data leaves the cluster. The SBOM and a workload inventory go to a service hosted in the EU. The architecture reference lists every field. If nothing may leave, Trivy Operator with alternate storage is your answer.
  • It is a hosted service. The agent's source is published. The backend is not, and you cannot self-host it.
  • Vulnerabilities and SBOMs only. No config audit, RBAC assessment, exposed-secret or compliance reports. Trivy Operator does all four.
  • No Prometheus metrics and no kubectl. Findings are in a web dashboard and an API, not in your Grafana.
  • The agent is not weightless. It pulls each new image once and requests 512Mi of memory, more than a single Trivy scan job. The difference is that it does this once per digest and not once per day.

If you stay on Trivy Operator

Most people reading this will, and should. Four things make the etcd problem manageable:

  1. Run the diff above on a schedule and alert on it. The missing report is the dangerous part, and it is the part you can fix this afternoon.
  2. Before you blocklist an oversized image, try trivy.ignoreUnfixed: true. A report of the findings you can act on is better than no report.
  3. If total etcd size is the worry, move to alternateReportStorage, add the cleanup CronJob from #2791, and accept that the vulnerability metrics go away until #2610 is fixed.
  4. Add a thumbs-up to #757 and the two field-removal PRs. Maintainers read those counts.
The two tools run side by side without conflict. A common split is Trivy Operator for config, secrets and RBAC checks inside each cluster, with its vulnerability scanner switched off, and StackRadar as the record of vulnerabilities across all of them. The feature-by-feature comparison is in StackRadar vs Trivy Operator.

See which images would not fit

The images that break report storage are the ones with the longest finding lists, and you can find them without installing anything in the cluster. The free StackRadar CLI scans every running image from your laptop and needs no account. To see the stored, cross-cluster view, open the live demo, or install the agent next to Trivy Operator on a staging cluster and compare the two for a week. The comparison page covers the operating-cost side in detail.

All posts