Capability Matrix · Vendor evidence
Scored candidates, with the evidence graded.
An earlier version of this page described a ninety-vendor database with five evidence tiers mapped one-to-one onto capability scores. That database descended from a retired tool, and the tier-equals-score mapping was wrong, so this page now documents the evidence discipline the live Matrix actually runs: thirteen scoring files across four component families and one gate axis, every cell carrying a 1–5 capability score, a weight, an A–D evidence tier, citations, caveats, and a shipped-versus-claim delta wherever a vendor claim exists. The full scored tables are public per the 2026-07-11 publication ruling; this page is how to read their evidence.
The tier system (MDR-0003)
Four tiers, and the tier grades the citation rather than the tool.
Every citation on every score carries an evidence tier under MDR-0003:
- Tier A. Peer-reviewed research, official standards (NIST, ISO, the OCSF spec body), or production deployment at recognized scale — Netflix on Iceberg, Pinterest, AWS Security Lake are the canonical examples. A named operator at recognized scale is Tier A; an anonymized reference ("a Fortune-100 team") or an NDA-gated result is Tier B at best, because a reader can't check it.
- Tier B. Practitioner engineering blogs from credible sources, conference talks, controlled first-party lab measurements, and multi-source synthesis where two B-tier sources agree. Most of the matrix's cells sit here, and the lab benchmarks that anchor the performance cells are labeled Tier B on their own pages precisely because they are single-host first-party runs rather than named production deployments.
- Tier C. Vendor blog or marketing material, cited only with an explicit bias flag on the cell.
- Tier D. Speculation or hypothesis. A Tier-D citation can never be the sole evidence for a non-null score; the
no_tier_d_uncorroboratedvalidation rule requires at least one A or B citation alongside it.
There is no Tier E. The old page carried one, and mapped each tier to a fixed score (Tier A meant 5, Tier D meant 2); both are retired because MDR-0003 defines four tiers and makes the tier a property of the citation, not of the capability. Where a cell's evidence is too thin to defend a score, the score is null and the weighted total becomes a bounds range — null is honest about evidence absence rather than a hidden zero (MDR-0002).
Tier versus score
A 5 backed by Tier C is less defensible than a 3 backed by Tier A.
The two letters and the number answer different questions: the 1–5 score says what the tool does for a criterion in an archetype, and the A–D tier says how credible the backing for that claim is, so the same score can be strong or weak depending on what stands behind it. The live tables show the distinction working. On the Archetype-A engines table, ClickHouse's raw-performance 5 carries a Tier-B flag because the ~10–11× foil band behind it is a controlled first-party benchmark across three independent draws rather than a named production deployment at scale, while Trino's Iceberg-native 5 carries Tier A because it rests on verifiable spec-level facts; and StarRocks's Iceberg-native cell holds a 3 at Tier C, flagged "less mature," because the only thing standing behind it is the vendor's own documentation, so the score stays conservative rather than borrowing that claim to go higher. Reading the tier column is how you tell which cells would move if better evidence landed, which is exactly what the revalidation triggers in the backing YAML files track.
Shipped versus claim
Where a vendor claim exists, the delta to shipped reality is recorded.
MDR-0003's second validation rule, shipped_vs_claim_when_vendor_claim_exists, requires that whenever a Tier-C vendor claim backs a criterion, the scoring file records a non-null shipped_vs_claim_delta — either the measured gap or "matches measurement" when the claim held up. Two live examples show the shape. Cribl's marketed 70–90% reduction headline is the aggressive-tuned ceiling rather than the out-of-box default; the honest default is 30–50%, the score is set on that adjusted basis, and the delta is documented in the YAML (assumption A-02 on theassumptions register). Tenzir's "OCSF-native" positioning is an availability claim, and the first-party pipeline-fidelity measurement found the shipped mapping can get the OCSF class right while getting the activity classification wrong — so the cell scores availability and fidelity separately instead of taking the claim whole. The point of the rule is that a vendor claim is allowed into the matrix, but it never enters unaccompanied.
The scored roster
Thirteen scoring files, scored per archetype, reconstructable from their cells.
The candidate roster is not a database of ninety vendors; it is a deliberately narrow set of candidates scored deeply, per archetype, in version-controlled YAML files whose weighted totals a linter reconstructs from the score-times-weight cells so no total can drift from its components. Weights sum to 100 per archetype under the anchoring rule (no weight above 30, none below 5, multiples of 5), and where a criterion holds null the total is published as a lower–upper bounds range instead of a point.
| Family | Candidates | Current leaders (A / B / C) |
|---|---|---|
| Engines (C3) · 3 files | ClickHouse, Trino, StarRocks, DuckDB (+ Athena at C) | ClickHouse 3.65–3.85 / Trino 3.90 / Athena 4.29 |
| Formats + catalogs (C1+C2) · 3 files | Iceberg+Polaris, Iceberg+Nessie, Delta+Unity, Iceberg+Glue (scored as pairs) | Polaris 4.25 / Polaris 4.10 / Glue 4.20 |
| Pipelines (C4) · 3 files | Tenzir, Vector, Cribl, Kafka Connect (+ Firehose+Lambda at C) | Tenzir 3.75–3.95 / Tenzir 3.70–3.90 / Vector 3.55 |
| Detection-survivability (C5) · 3 files | The scored engines, as detection-execution targets | Scored per archetype; bands on its page |
| Compliance / WORM axis · 1 file | The C1+C2 pairs, gate-capable (disqualify, not discount) | Archetype-independent; verdicts on its page |
The archetypes are the published engagement shapes: A is the Zeek-heavy SOC (5–50 TB/day, p99 under 5 s on recurring queries), B the multi-source federated lakehouse (10–100 TB/day, federated joins across CMDB, EDR, identity, and Zeek), and C serverless-within-AWS, where Iceberg-plus-Glue is assumed settled and the question is which engine sits on top. Same candidates, same criterion definitions, different weights, different rankings — which is the methodology's main validation, because a matrix whose ranking never moved with the workload would just be a leaderboard with extra steps. Candidates beyond this roster (the wider commercial SIEM and pipeline field) are surveyed in the component pages' candidate catalogs and enter the scored roster when an engagement or the evidence justifies the scoring effort.
What a reader can verify where
The audit surface is deliberately larger than the claims.
- The scored tables, per-cell tiers, and weighted totals are on theengines,formats + catalogs,pipelines,detection-survivability, andcompliance / WORM pages.
- The scoring decisions behind the weights and criteria are theMDR registry; the premises with refutation criteria are theassumptions register.
- The first-party measurements the performance and cost cells cite are on thelab page, most with public code you can rerun.
- Every weighted total reconstructs from its cells: a linter re-derives the totals from score × weight and the weight sums, so a stated total that stopped matching its components would fail the check rather than ship.
- Two things stay withheld, and the cells say so: figures naming an engine covered by a vendor benchmark-restriction clause carry "figure withheld, available under NDA" (a legal constraint, not paid IP), and per-client outputs — your weights, your crossover, your gate verdicts — are the engagement work product.
The tier column is what lets you argue with the matrix.
Disagree with a score and the citation tells you exactly what evidence you'd need to bring.