Security Data Works

Capability Matrix · Scoring decisions

Thirty-seven scoring decisions. On the record.

These are the scoring-layer decisions: criterion taxonomies per component, archetype-conditional weight tables, the evidence-tier system, null-handling, the publication gate, the revalidation cadence, the disclosure discipline, the v1.3 evidence updates, and the later records that turned a shortlisting tool into a migration-decision instrument, from the decision paths and the board read through the two scored security axes to the measured re-scores, ending with the v1.4 engines concurrency cut of 2026-07-20. Each MDR carries its provenance and is open to challenge.

The frame these decisions operate inside lives on the siblingdecisions page, which carries the Matrix's meta-decisions in narrative form: the 1–5 scoring scale, the A–D evidence tiers, the 3–14 dimension band, archetype-conditional weights, null propagation with the 0.75 completeness gate, the two gate-capable security axes, the revalidation cadence, the disclosure discipline, and the 2026-07-11 publication ruling. That URL used to carry six architecture decision records for the retired MCP vendor-lookup tool, which survive in its git history rather than on the live page, so this page is the per-record registry and that one is the frame.

Migration-instrument cut · 2026-06-17

Six new decision records take the matrix from shortlisting to a migration-decision instrument.

MDR-0026 (the Pipelines virtual-view criterion, Proposed and held pending evidence) is the scoring-layer candidate. The other five answer the questions a board actually votes on. MDR-0027 + MDR-0028 (Accepted) make the incumbent and the partial moves scored candidate paths with a crossover computed from the cost-to-serve cells, and require a four-part board-defensibility read per recommendation — surfaced on the decision-path page.

MDR-0029 (Accepted) makes Component 5 scored on detection-survivability — what fraction of the detection corpus survives a move, where survive means executes correctly, scoring silent degradation worst (detection-survivability). MDR-0030 (Accepted) adds the first gate-capable cross-cutting axis: compliance/WORM, where a mandated control disqualifies rather than discounts (compliance / WORM). MDR-0031 (Accepted 2026-06-17) ratified the illustrative-worked-example exception; the 2026-07-11 publication ruling then made the full matrix public, so the public/paid line it refined no longer exists and the record stands as history.

The registry now holds 39 decision records (MDR-0001–0039), and every one of them has a section below. Section 5 collects this cut and everything after it; the six that landed later are the two Move ratifications (MDR-0032, MDR-0033), the measured detection-survivability re-score and its declined follow-on (MDR-0034, MDR-0035), the DSMOS inventory migration (MDR-0036), and the two engines concurrency re-scores (MDR-0037 and MDR-0038) that carry the v1.4 evidence stamp of 2026-07-20, and the evidence-tier clarification the same pass forced (MDR-0039). The label is worth pinning down, because it used to sit on two things at once: v1.4 is the 2026-07-20 evidence-vintage cut for Component 3, so the six records above are dated 2026-06-17 as the migration-instrument additions rather than as that cut.

v1.3 evidence cut · 2026-05-25

Three evidence MDRs landed in the v1.3 cut.

MDR-0025 (Nessie / Polaris metadata-fetch parity at toy scale) is the newest evidence-strengthening addition: a first-party Tier-B measurement against the local docker stack closed a Tier-B/C-only differentiator between two catalog choices that had previously been scored on vendor-published claims. Assumption A-05 (Unity OSS self-hosted immature) also moved to confirmed in the same cut.

MDR-0023 (engines native IP type criterion) was accepted in the v1.3 cut (OG-4, 2026-06-19) on Tier-B doc-review evidence, with the Tier-A storage-size and predicate-latency measurement still owed to the Q3 / cluster benchmark; the native-IP-type criterion lifted Trino to #2 at Archetype A and confirmed Athena at #1 at C. MDR-0024 (F+C storage-layer maintenance flag, anchored on the MinIO archive event) was ratified 2026-06-23 on Tier-A GitHub-API evidence, with its application to the scored YAMLs deferred to the next refresh trigger so the scored matrix is not amended mid-version.

Assumption tracker (A-01 through A-17) →

Section 1 · Foundational MDRs

The structure under everything else.

MDRs 1 through 5 and 12 through 16 define what the matrix is, how it scores, how nulls propagate, what gets published, when scores are revisited, and how partnerships are disclosed. Every per-component MDR operates inside this frame.

MDR-0001 · Accepted 2026-05-25

The Capability Matrix is the practice's central product.

Decision. After retiring the POV-on-customer-data offering on 2026-05-18 (the regulatory-compliance burden of running benchmarks across diverse customer environments was the deciding factor), the practice needs a single product around which engagements, content, and brand organize. The matrix is that product. The lab is its benchmark-evidence layer. The scored matrix is published in full as the practice's flagship evidence asset; what engagements deliver, and what clients pay for, is applying it to their own environment: their sources, their weights, the migration sequencing and reversibility cost, rather than access to the scores.

Why this won. "Recommend a vendor" is not a defensible product shape: best-engine and best-pipeline depend on workload, retention, in-house skill, and lock-in tolerance. A vendor-by-vendor recommendation service collapses fair-broker positioning into vendor-relationship management. Pure benchmark publishing loses the engagement economics and the cross-component bundling customers actually need. The matrix is the framework that forces answers like "X is better than Y for this archetype, here is the evidence per criterion, here are the cross-component constraints."

MDR-0002 · Accepted 2026-05-25

Score on a 1–5 scale, with null permitted.

Decision. Every criterion is scored 1 (architectural mismatch) through 5 (best available for this criterion in this archetype). Null is a legitimate state when evidence does not support a defensible score. A validation rule requires at least one citation on every non-null score.

Why this won. A 0–100 NPS-style scale offers false precision. Practitioners anchor on round numbers, and defending a 73 versus a 76 is harder than defending a 4 versus a 4. Likert scales fit opinion surveys, not capability assessments. Pass / Conditional / Fail collapses "good but not best" and "the best" into one bucket, losing the differentiating information the matrix exists to expose. The 1–5 scale matches the granularity I can actually defend in writing, and null is honest about evidence absence rather than fabricating a 3.

MDR-0003 · Accepted 2026-05-25

Every citation carries an A / B / C / D evidence tier.

Decision. Tier A is peer-reviewed research, official standards (NIST, ISO, the OCSF spec body), and production deployment at recognized scale, where the record's own examples are Netflix, Pinterest, and AWS Security Lake. The qualifier that does the work on the production side is the named operator: a production reference earns A when someone is named to corroborate it and the scale is recognized, so an anonymized reference (a "Fortune-100 deployment") or an NDA-gated result sits at B at best, because a reader cannot check it. Tier B is practitioner engineering writing from credible sources, conference talks (Strata, BSides, Black Hat), controlled first-party benchmarks including the lab's own, and multi-source synthesis where two B-tier sources agree. Tier C is vendor blog or marketing material, cited only with explicit bias-flag annotation. Tier D is speculation; it cannot be the sole citation on a non-null score, since the no_tier_d_uncorroborated rule requires an A or B citation alongside it.

Why this won. Without tiering, a vendor benchmark post and a Netflix production reference would carry the same evidence weight, and the matrix is the practice's public flagship, so its credibility lives in that distinction. Four tiers is the natural split (peer-reviewed and named-operator production / practitioner / vendor / speculation); three would collapse vendor with peer-reviewed, and five-plus would over-divide them. The tier stays independent of the 1–5 score throughout, because the two answer different questions, so a 5 backed by Tier C is less defensible than a 3 backed by Tier A, and the site's own essays are not Tier A for this purpose. When a vendor (Tier C) claim exists for a criterion, a shipped_vs_claim_delta field must be non-null, capturing the gap between claim and measured reality.

MDR-0004 · Accepted 2026-05-25

Score per archetype with conditional weights, not absolute weights.

Decision. Every component is scored separately for each archetype. Criterion definitions stay constant; weights vary by archetype. A candidate's score on a given criterion may also shift across archetypes when the archetype reframes the criterion's relevance (concurrency under federated workload differs from concurrency under single-source). Archetypes are pre-set per component, not negotiated per engagement; engagements are assigned to their closest archetype.

Why this won. This is the single most important methodology choice in the matrix. Without it, the matrix collapses to a global ordering with no defensibility ("ClickHouse wins all workloads" or "Trino wins all workloads", both demonstrably wrong). The same criterion (raw_analytical_query_performance) is the dominant value driver for one workload and an afterthought for another. An engagement that does not fit any published archetype is itself a finding: either I add a new archetype or score the engagement with custom weights and an explicit "deviation from standard archetype" note.

MDR-0005 · Accepted 2026-05-25

Three to fourteen criteria per component. Strict bounds.

Decision. Each component's criterion count must be strictly greater than 2 and strictly less than 15. Engines uses 9 criteria (8 at v1.0, with MDR-0023 adding the ninth); Formats+Catalogs 10; Pipelines 9; Component 5 carries 6 under MDR-0029; all compliant.

Why this won. Too few criteria (1–2) collapses real trade-offs into a fortune-cookie ranking. Too many (15+) dilutes weight discrimination below the noise floor. At 15 criteria, weights cluster around 5–10 each, which is below the threshold where "is this a 3 or a 4?" carries meaning. The bound is compatible with the Cynefin framework's preference for roughly 7 ± 2 distinctions before cognitive load degrades discrimination. A concrete trigger for "too high": any criterion in a published matrix carries weight below 4.

MDR-0012 · Accepted 2026-05-25

Weighted total is null when any criterion is null. Bounds instead.

Decision. If any criterion has a null score, the weighted total is null. Instead of a point estimate, two bounds are computed: a lower bound (treats every null as 1, the scale floor) and an upper bound (treats every null as 5, full credit). Customers see "ClickHouse 3.65 – 3.85," not "3.72 ± something I am not telling you."

Why this won. Skipping null criteria silently re-normalizes weights and breaks comparability across candidates with different null patterns. Collapsing null to a single point estimate distorts either way — 0 over-penalizes, "average" (3) is false precision — since the entire reason for null is that I cannot honestly score. Bounds preserve the unknown explicitly by spanning the full 1-to-5 score range. Sensitivity analysis becomes straightforward: if the ranking changes when nulls fill at the low versus high end, the ranking is not v1-publication-ready.

MDR-0013 · Accepted 2026-05-25

75% evidence completeness is the v1 publication gate.

Decision. A scoring run is v1-publication-ready only when at least 75% of criteria carry non-null scores. Below this threshold, the file may exist as a working draft but cannot be cited from the published matrix or aggregate ranking documents. All three v1 category matrices currently clear the threshold: Engines 0.889, Formats+Catalogs 1.0, Pipelines 0.889.

Why this won. 0.50 is too permissive; bounds spread too wide and the ranking destabilizes under sensitivity analysis. 0.90 is too strict and would block publication for legitimate client-specific nulls (like in-house skill base, structurally null at standing-matrix level) without engagement context. At 0.75, the lower-vs-upper-bound spread stays under 1.0 on a 5-point scale for most realistic null patterns. Failing the threshold is information: it tells me where evidence collection is needed before publication.

MDR-0014 · Accepted 2026-05-25

Public methodology and ordering. Paid scores.

Superseded 2026-07-11. The publication ruling makes the full per-criterion scores, citation tiers, weighted totals, and shipped-vs-claim deltas public; they are the flagship evidence asset, not a withheld work product, and what clients pay for is the services engagement that applies the matrix to their environment. The disclosure-scrub delay and the public-correction discipline below survive; the three-tier withholding of scores does not. The original decision is kept here as the record.

Decision. Three-tier split. Methodology is fully public: criterion set, weight tables per archetype, evidence tiers, validation rules, archetype definitions. Ordering is public with a 4–6 week delay after the engagement-internal v1 completes (the delay exists for disclosure-correctness scrub). Per-criterion 1–5 scores, citation tiers, weighted totals, shipped-vs-claim deltas, and cross-component bundles are the engagement work product. They are not published.

Why this won. Publishing full scores destroys engagement economics; there is nothing left to engage for. Publishing nothing destroys the brand surface; the matrix has nothing to demonstrate. Publishing scores at coarse granularity (Pass / Conditional / Fail) collapses the differentiation the matrix exists to provide. The three-tier split makes the public surface "enough to evaluate the method" without becoming "enough to substitute for the engagement." If disclosure flags an issue post-public-release, the response is a public correction, not a silent edit.

MDR-0015 · Accepted 2026-05-25

6-month default revalidation. Faster on benchmark triggers.

Decision. Every score and every MDR carries an explicit last_reviewed date and arevalidate_by date. The default cadence is 6 months. Faster revalidation triggers on a major vendor release (Iceberg V4, ClickHouse 25.x with new Iceberg parity), a relevant benchmark landing, a partnership forming (which changes the disclosure block), or a practitioner-identified shipped-versus-claim delta that voids a previous score.

Why this won. Scores age. Without a revalidation cadence the matrix drifts into outdated territory silently. 3 months is burn-the-team cadence; most criteria do not move that fast. 12 months is too long; the engine, catalog, and table-format landscape moves faster. 6 months matches the median upstream change rate observed across ClickHouse major releases, catalog GA timelines, and Iceberg v1.X releases. The Q3 2026 catalog benchmark, the still-open cluster and AWS-native slices of the OLAP bake-off (its single-host engine-join slice ran 2026-06-10, Tier B), and the ~Q1 2027 H1-COST-02 pipeline benchmark are the primary refresh anchors.

MDR-0016 · Accepted 2026-05-25

Every scored candidate carries a disclosure block.

Decision. The structure is mandatory; the content is null when nothing exists to disclose. When a relationship exists, the block names the nature (explore / active partnership / former employer), the scope in plain language, and the impact_on_score. One disclosure is currently on file: the Delta+Unity candidate carries nature: explore in all three F+C scoring files for an in-progress conversation with Databricks about Lakewatch go-to-market, recorded with no score-altering effect. No commercial partnership exists with any scored candidate, and no other candidate carries a filled block.

Why this won. Fair-broker positioning is the practice's central differentiator. A customer auditing the matrix has no way to detect partnership bias without explicit disclosure. Hiding partnerships would be the most damaging credibility hit available. Disclosing only when commercial would under-disclose; in-progress conversations are also signal. Optional disclosure invites inconsistency. Mandatory structure with conditional content is the rule that lets a reader trust the score even where partnerships exist.

Section 2 · Engines (Component 3)

Query engines. Nine criteria, three archetypes.

ClickHouse, Trino, StarRocks, DuckDB across Archetypes A and B. Athena joins as a 5th candidate at Archetype C. A 9th criterion, native IP type support, was added in the v1.3 cut (MDR-0023).

MDR-0006 · Accepted 2026-05-25

Engines criterion taxonomy — eight criteria.

Decision. Score query engines on: raw analytical query performance, Iceberg native versus connector, semantic layer / materialized views support, federation breadth, operational complexity, SPL dialect distance, concurrency and multi-tenant behavior, and existing in-house skill base (plus, since the v1.3 MDR-0023 addition, native IP type support — nine criteria current). The skill-base criterion is held null at standing-matrix level (filled at engagement scoping), and that pattern caps standing evidence completeness at 0.889.

Why this won. Cost-per-query as a separate criterion is entangled with operational complexity (managed vs self-hosted) and the downstream storage tier; folded into qualitative notes instead. Latency versus throughput as separate criteria is a premature split; most workloads do not care about the distinction until they care a lot. Iceberg V3 / V4 currency as a separate criterion would be a stand-in for "which engine ships new versions fastest"; captured inside the iceberg_native_vs_connector caveats per candidate; revisit at v2 if V4 ships and engines diverge sharply.

MDR-0007 · Accepted 2026-05-25

Engines Archetype A weights — Zeek-heavy SOC.

Decision. For the most common engagement shape (Zeek-heavy SOC at 5–50 TB/day with a p99 under 5 seconds target on 80% recurring queries), raw analytical query performance carried weight 30 at v1.x; the v1.3 native IP type split (MDR-0023) took it to 22 and added native IP type at 8. Iceberg-native, operational complexity, and concurrency each weight 15; SPL dialect distance weights 10; semantic layer, federation, and in-house skill base each weight 5. Sum is 100. Under the v1.4 cut, ClickHouse ranks #1 (3.65–3.85 bounds); Trino #2 (3.41–3.61); StarRocks #3 (2.87–3.07); DuckDB #4 (2.63–2.83).

Why this won. Equal weights destroy archetype signal and produce nonsense rankings. The raw-performance weight (30 at v1.x, 22 after the v1.3 native IP type split) reflects that p99 under 5 seconds is the analyst-experience floor under Archetype A; H3-PERFORMANCE-01 (ClickHouse 46.8× the schema-on-read SIEM on the five-query average over a 10M-event Zeek corpus — 5–62× on the hunting-shaped queries across the native and Iceberg realizations, the index wins the simple lookups, answer-equality verified, single-node) anchors at Tier B. The anchoring rule (no weight above 30, no weight below 5, multiples of 5) floors precision and prevents weight-tuning to predetermine a winner; the v1.3 native IP type split is the one documented exception, funding an 8-point criterion at Archetype A (4 at C) from the raw-performance budget. The OLAP bake-off is the gating validation, and its single-host slice has now run (the engine-join-specialization bench, 2026-06-10, Tier B): every engine answered the SOC join suite in under 1.5 s, so that slice tightened the spread rather than exposing an over-rewarded candidate, and the cluster and federated-join scenarios that could still move the weight remain open. If a remaining slice shows the weight over-rewards a candidate that fails in practice, this MDR gets superseded.

MDR-0017 · Accepted 2026-05-25

Engines Archetype B weights — multi-source federated lakehouse.

Decision. For the second engagement shape (multi-source federated lakehouse at 10–100 TB/day with federated joins across CMDB, EDR, identity, and Zeek), the weight distribution shifts substantially. Iceberg-native lifts to 25 (from A's 15); raw performance drops to 20 (from 30), re-split at v1.3 into 15 + 5 native-IP per MDR-0023; federation_breadth lifts to 20 (from 5), the largest single weight shift in the matrix. Semantic layer lifts to 15; operational complexity drops to 10; SPL distance drops to 5; concurrency drops to 5. The in-house skill base weight drops to 0; the criterion stays in the set but does not affect scoring under B (the matrix presumes greenfield-class hiring at this engagement scale).

Why this won. The Archetype A weighting produces nonsense rankings under federated workloads. Trino's federation strength scored a 5 at A but only contributed weight 5 to the total. At B, federation IS the archetype's name; Pinterest's Trino + Gravitino + Iceberg at 17k+ nodes anchors at Tier A. The weight-0 convention is itself a decision: criteria can be retained in the set but score-neutral in archetypes where they do not apply. The convention propagates to MDR-0022 (Pipelines Archetype C).

MDR-0021 · Accepted 2026-05-25

Engines Archetype C weights — serverless within AWS.

Decision. Archetype C narrows the question to "which engine plays best with Iceberg-on-Glue under AWS-managed-service expectations?" The implicit reference is Athena (AWS-managed Trino). Iceberg-native weights 25; operational simplicity weights 20; within-AWS federation 15; raw performance 11 (cut from 15 to fund the v1.3 native IP type criterion at 4, and capped anyway by AWS-managed-service constraints like Athena workgroup limits and ClickHouse Cloud quotas); concurrency 10; SPL distance, semantic layer, and in-house skill each weight 5. Athena is added as a 5th candidate under C (it was excluded at A and B as a managed-Trino derivative).

Why this won. The earlier session synthesis claimed engines did not have a clear AWS-native distinction the way Formats+Catalogs did, because "engines query any catalog." That synthesis was wrong under the framing that C constrains Components 1+2 to Iceberg+Glue; once the catalog is fixed, the engine question narrows to AWS-managed-service operational simplicity plus Iceberg-native integration depth as the twin design centers. Under v1.3, Athena leads at 4.29 with Trino a clear second at 3.98, a 0.31 gap driven by Athena's managed-Trino integration depth; ClickHouse is at ~3.65 as the raw-perf weight cut outweighs the iceberg lift; DuckDB-on-Lambda sits at ~2.92 anchored on the Okta 7.5T-record threat-hunting case. The single-host slice of the OLAP bake-off ran 2026-06-10 without Athena, which a lab host cannot run, so the Athena inclusion still anchors at Tier-B doc review until the still-open AWS-native scenario runs it in place.

MDR-0023 · Accepted 2026-06-19

Added a ninth Engines criterion — native IP type support.

Decision. Add native_ip_type_support as a 9th criterion to the Engines taxonomy. Scoring 5 means both IPv4 and IPv6 are first-class types with type-aware functions likeisInSubnet and cidr_match; 4 means native types with limited functions or IPv4-only; 3 means extension-available; 2 means string-only with auxiliary functions; 1 means string-only with no IP-aware functions. ClickHouse and Trino score 5; Athena scores 4 (Presto/Trino IPADDRESS lineage); DuckDB scores 3 (inet extension loadable, not default); StarRocks scores 2 (string-only). The criterion is funded from the raw-performance budget (A 30→22+8, B 20→15+5, C 15→11+4) and re-ranks the field: the native IP type criterion lifted Trino to #2 at Archetype A in the v1.3 cut, which the v1.4 concurrency re-score narrowed but did not overturn on the published table, and confirms Athena at #1 at C.

Why this is now Accepted (v1.3, OG-4). Vern Paxson cited Zeek's choice to treat IP addresses as first-class types (not abstracted to integers) as one of the most important performance optimizations in Zeek's design. A 2026-05-25 technical conversation exposed the gap. The criterion is architectural and binary — an engine either ships native IP types or it does not — so it is scoreable from documentation at Tier B today, which is what the v1.3 cut accepts. The Tier-A storage-size and predicate-latency measurement (the Q3 / cluster benchmark generators currently emit src_ip as a string and need to emit native types) is still owed, and it will revalidate the weight rather than the yes/no score. Folding the dimension into raw query performance would conflate an architectural choice (binary) with a continuous performance measurement (different dimension), which is why it is its own criterion.

Section 3 · Formats + Catalogs (Components 1+2)

Bundled paired choices. Ten criteria, three archetypes, two v1.3 additions.

Iceberg+Polaris, Iceberg+Nessie, Iceberg+Glue, and Delta+Unity scored as paired choices across archetypes A (multi-engine open lakehouse), B (single-vendor managed), and C (AWS-native). MDR-0024 adds a storage-layer maintenance flag; MDR-0025 records a first-party Tier-B parity finding for Nessie versus Polaris.

MDR-0008 · Accepted 2026-05-25

Formats and Catalogs scored together — ten criteria, paired candidates.

Decision. Components 1 (table format) and 2 (catalog) are scored as paired choices (Iceberg+Polaris, Iceberg+Nessie, Iceberg+Glue, Delta+Unity), not independently. Ten criteria: multi-engine query support, multi-engine federation, open governance, catalog ecosystem maturity, RBAC and column-level security, schema evolution and time travel (format), time travel and branching (catalog), cloud-native versus self-hosted, cost model (OSS versus licensed), and cloud / on-prem flexibility. No in-house-skill-base criterion; catalog choice is a platform decision, not a per-analyst skill match.

Why this won. Scoring formats and catalogs independently produces operationally undeployable pairs ("Unity 5/5 + Hudi 4/5"). The choice is structurally coupled: Unity ties to Delta, Polaris is Iceberg-first, Glue is AWS-bound, Nessie is Iceberg-native. The bundled framing is what engagements actually need. Scoring only catalogs with formats as an attribute would collapse Iceberg versus Delta, an essential dimension. Bundling further into Component 3 (engines) would create too many simultaneous decisions and lose the locality that makes scoring defensible.

MDR-0009 · Accepted 2026-05-25

F+C Archetype A weights — multi-engine open lakehouse.

Decision. Multi-engine query support carries weight 20, the single highest, and it is the archetype's name. Federation and RBAC each weight 15. Open governance, catalog ecosystem maturity, and cloud / on-prem flexibility each weight 10. Cost model, schema evolution, branching, and cloud-native each weight 5. Iceberg+Polaris ranks #1 at 4.25; Delta+Unity #2 at 3.30; Iceberg+Nessie #3 at 3.10; Iceberg+Glue #4 at 2.70. The 2.70 score for Glue here reflects archetype mismatch (its proper home is Archetype C), not a defect of Glue.

Why this won. Pushing multi-engine to weight 30 would over-reward Iceberg-first candidates; the 20 weight balances against governance and RBAC. Dropping branching to 0 was tempting (few archetypes truly need it) but eliminating it loses Nessie's principal differentiator from the comparison entirely. Pinterest's Trino + Gravitino + Iceberg multi-engine production reference anchors the federation weight at Tier A; Netflix's 5 PB/day Iceberg anchors multi-engine query support.

MDR-0018 · Accepted 2026-05-25

F+C Archetype B weights — single-vendor managed lakehouse.

Decision. For "already on Databricks or similar; optimize for operational simplicity within the vendor ecosystem; accept some portability tradeoff," RBAC lifts to weight 25 (from A's 15), because the production access-control gate is the design center under a managed single-vendor lakehouse. Cloud-native versus self-hosted lifts to 15 (from 5); multi-engine query support drops to 10 (from 20); cost model lifts to 10 (from 5); federation drops to 5 (from 15); open governance drops to 5 (from 10). The honest finding: Polaris narrowly leads Delta+Unity, 4.10 to 3.95, under B. Iceberg+Glue lifts from 2.70 to 3.45, above Nessie's 3.25 and closer to its AWS-native home but not there yet.

Why this won. Pushing RBAC to 30 would tilt decisively to Delta+Unity and predetermine the winner; 25 is already at the high end of the anchoring rule. Polaris narrowly leading Delta+Unity at B, 4.10 to 3.95, is a finding worth flagging: Delta+Unity's expected dominance at its home archetype is partially offset by Polaris's broad strengths (multi-engine 5, OSS cost 5, on-prem flexibility 5) that hold up even when downweighted, so the gap closes to a 0.15 near-call rather than the wide Delta+Unity win the home archetype predicts. Without per-archetype scoring this nuance would be invisible.

MDR-0020 · Accepted 2026-05-25

F+C Archetype C weights — AWS-native lake.

Decision. Archetype C narrows the vendor constraint to AWS specifically (Lake Formation + IAM + Athena + EMR + Lambda + Glue ETL). Cloud-native and RBAC each carry weight 20, the twin design centers. Catalog ecosystem maturity weights 15 (highest across all three F+C archetypes). Multi-engine within-AWS weights 10; cost model 10; federation, open governance, schema evolution, branching, and cloud/on-prem each weight 5. Iceberg+Glue lifts to ~4.20 from A's 2.70, a +1.50 shift, the largest single-candidate archetype-shift in v1.x. Delta+Unity drops to ~3.60 (Databricks-on-AWS is Databricks-managed, not AWS-native).

Why this won. The v1 cross-archetype synthesis showed Glue was mislabeled at A and B; its proper home is C. Combining cloud-native + RBAC into a single 40-weight criterion was considered and rejected because managed-serverless and access-control are conceptually distinct even when practically intertwined, and the anchoring rule (no weight above 30) is the rule that prevents weight-tuning to predetermine winners. Some scores also shift (not just weights) when context reframes the criterion. Glue's multi-engine score moves from 4 at B to 5 at C because the within-AWS multi-engine depth is the reference, not cross-cloud.

MDR-0024 · Accepted 2026-06-23 · application deferred

F+C storage-layer maintenance status flag. Ratified, not yet applied.

Decision. A storage_layer_maintenance_status metadata flag (not a scored criterion, a per-candidate annotation) attaches to F+C scoring entries where the candidate's default or commonly-paired deployment includes a specific named storage layer. Values: ACTIVE, MAINTENANCE, or ARCHIVED, shown as a Go/No-Go gate ahead of the weighted score rather than buried in qualitative notes. The owner ratified the flag design and its backfill table on 2026-06-23 on Tier-A GitHub-API verification (MinIO ARCHIVED, the repo archived with last push 2026-04-24; SeaweedFS, Ceph, Rook, and Garage ACTIVE). The flag is not yet written into the scored YAMLs: application is deferred to the next refresh trigger per MDR-0015, so the scored matrix is not amended mid-version, and each backend's status gets re-verified at application time.

Why this won. The MinIO archive event made the need concrete: UI removed from OSS (2025-05-25), Docker images stopped (2025-10-23), README to maintenance mode (2025-12-03), repo archived 2026-02-13. A customer reading a v1.2 F+C score for a MinIO-paired deployment in May 2026 would not see the maintenance risk without this flag. Ratifying now while deferring application splits the decision from the amendment deliberately, because the external evidence is Tier-A complete and benchmark-independent (MinIO is archived and that will not reverse, so waiting teaches nothing), while writing the field into scored files mid-version would break the revalidation cadence the matrix committed to. Adding it as a scored criterion would spend a dimensionality slot on a binary state better shown as a Go/No-Go footnote.

MDR-0025 · Proposed 2026-05-25

Proposed: Nessie and Polaris share metadata-fetch latency at toy scale.

Decision (Proposed). At toy 1K-rows-per-table scale with DuckDB as the query engine and byte-identical data across both catalogs, Nessie and Polaris exhibit statistically indistinguishable metadata-fetch latency (Nessie median 98.9 ms / p95 241.1 ms; Polaris median 96.4 ms / p95 236.0 ms; all 9 queries within 5% on median across 3 runs). Use this finding to remove "catalog driver overhead" as a Nessie-versus-Polaris differentiator on the read-path criterion, and to anchor a Tier-B confidence floor under the existing per-archetype scores.

Why this is Proposed. Prior to 2026-05-25 the Iceberg+Nessie versus Iceberg+Polaris spread (Archetype A: Polaris 4.25 vs Nessie 3.10) was scored on Tier B/C vendor signals only. This first-party measurement closes the metadata-fetch dimension at Tier B and confirms the 1.15 spread is NOT driven by catalog driver overhead; it is driven by RBAC posture and ecosystem maturity. Toy-scale does not survive to Tier A: at 1 TB, Polaris's per-call authorization checks may scale differently than Nessie's branch-based access pattern. The MDR moves to Accepted when the Q3 2026 catalog benchmark produces a greater-than-1M-row counterpart. Adding metadata_fetch_latency as a new scored criterion mid-version would violate the dimensionality bound without offsetting removal; deferred to v1.4.

Section 4 · Pipelines (Component 4)

Ingestion. Nine criteria, three archetypes, an evidence-asymmetry caveat.

Cribl, Tenzir, Vector, Kafka Connect scored across Archetypes A (cost-reduction-led), B (schema-normalization-led / OCSF), and C (AWS-native ingest). Kinesis Firehose + Lambda joins as a fifth scored candidate at Archetype C, promoted from the archetype's implicit reference in the v1.2 scoring (weighted total 3.35, #3). Tenzir's weighted-total leads at A and B but the production- evidence asymmetry (Cribl Tier A/B versus Tenzir Tier B/C) means weighted total does not translate to "recommend Tenzir."

MDR-0010 · Accepted 2026-05-25

Pipelines criterion taxonomy — nine criteria.

Decision. Score pipeline platforms on: default reduction ratio, aggressive reduction ratio (the ceiling), OCSF normalization fidelity, cross-source schema correlation, lines of config to express a SOC ruleset, resource consumption (CPU and memory), pricing model, vendor lock-in / portability, and in-house skill base (held null at standing-matrix level). Default reduction and aggressive reduction are split deliberately. Vendor marketing collapses them; the matrix should not.

Why this won. A single reduction-ratio criterion would let vendor marketing inflate scores; Cribl markets 70–90% but A-02 (Cribl 70-90 is tuned, not default) anchors that the marketed ceiling requires ongoing tuning. Latency-to-detection as a separate criterion is Tenzir-favoring (pipeline-based detection at 15–50 ms versus query-based 5–30 s) but generalizes weakly across the four candidates; captured in qualitative notes instead of a weighted criterion. Cross-source schema correlation is a Tenzir differentiator that would be lost if dropped; kept at 5.

MDR-0011 · Accepted 2026-05-25

Pipelines Archetype A weights — cost-reduction-led ingest.

Decision. For 500 GB/day, 80% of value in 20% of events, schema-on-read SIEM downstream, 1–2 engineers: default reduction ratio carries weight 25 (the dominant value driver). Config LOC and resource consumption each weight 15. Aggressive reduction, pricing model, and vendor lock-in each weight 10. OCSF, cross-source correlation, and in-house skill each weight 5. Tenzir ranks #1 (3.75–3.95); Cribl #2; Vector #3; Kafka Connect #4. Tenzir's lead does NOT translate to "recommend Tenzir" given the production-evidence asymmetry.

Why this won. Combining default and aggressive reduction into one 30-weight collapses the vendor-claim versus measured-reality distinction. RiverSafe's UK independent measurement of 40% Cribl SIEM reduction (Tier B) anchors the practical reduction-ratio range. The 25 weight stays under the 30 cap because while cost-reduction is dominant, it is not overwhelming; operational realities (config LOC, resource consumption, lock-in) keep the matrix honest about TCO. The production-evidence asymmetry is captured separately in the aggregate ranking: Tenzir's weighted lead is about scores; Cribl is the production-defensible default.

MDR-0019 · Accepted 2026-05-25

Pipelines Archetype B weights — schema-normalization-led, OCSF target.

Decision. For 200 GB/day from 10+ sources writing OCSF-conformant events to Iceberg with 2–3 engineers: OCSF normalization fidelity carries weight 30, the largest single weight in any archetype in the matrix. Config LOC and cross-source schema correlation each weight 15. Resource consumption and default reduction each weight 10. Aggressive reduction, pricing model, vendor lock-in, and in-house skill each weight 5. Tenzir's OCSF-native architecture wins decisively at B; Cribl holds middle ground via Packs; Vector and Kafka Connect drop further.

Why this won. Under Archetype B, schema normalization is the workload itself. OCSF at 25 instead of 30 would understate the shift; the archetype's name carries authority. Combining OCSF normalization with cross-source correlation would collapse two distinct dimensions (per-source normalization quality versus join enrichment); kept separate. Dropping default reduction to 0 would predetermine that reduction does not matter at B, but Archetype B still ingests 200 GB/day and the downstream SIEM may still benefit; kept at 10.

MDR-0022 · Accepted 2026-05-25

Pipelines Archetype C weights — AWS-native ingest.

Decision. For AWS-native ingest where the implicit reference is Kinesis Firehose + Lambda + Glue ETL writing to Iceberg-on-S3: OCSF fidelity and pricing model defensibility each weight 20, the twin design centers. Vendor lock-in and config LOC each weight 15. Resource efficiency and default reduction each weight 10. Correlation capability, aggressive workload handling each weight 5. In-house skill weights 0 (the weight-0 convention from MDR-0017). Vector lifts to #1 at 3.55; the first time Vector ranks #1 under any archetype, driven by OSS economics and lowest vendor lock-in. Tenzir drops to #2 at 3.40; with Firehose + Lambda scored as the fifth candidate at 3.35 (#3), Cribl lands #4 at 2.95 (per-volume licensing on top of AWS infrastructure becomes a material TCO penalty at the 20-weight on pricing) and Kafka Connect #5 at 2.75.

Why this won. The honest finding the matrix produces: no dominant winner under Archetype C. At ratification the implicit Firehose + Lambda reference was estimated at ~3.50 if added; the v1.2 scoring then promoted it to an explicit fifth candidate and it landed at 3.35, third behind Vector (3.55) and Tenzir (3.40), so the estimate held within 0.15 and the finding stands. The matrix is telling AWS-committed customers that Vector-on-EKS, AWS-native Firehose+Lambda, and Tenzir-on-EKS are roughly equivalent in defensible cost-and-lock-in posture; the call comes down to specific source coverage and team preference. That is the methodology working correctly; not every archetype has a clear winner. Pushing pricing-model to 25 would predetermine Vector as winner; the 20-weight cap matches OCSF and preserves the twin design center.

Section 5 · The migration-instrument cut and after

From shortlisting to a migration-decision instrument.

MDR-0026 through MDR-0037 carry the decision-path and board-read records, the two scored security axes (detection-survivability and the compliance gate), the Move ratifications, the two measured re-scores, the registry's data-custody housekeeping, and the v1.4 engines concurrency cut. Two ratification records are deliberately thin summaries here, because their scored outputs are computed per engagement on a client's own inputs; they say so inline rather than padding.

MDR-0026 · Proposed 2026-05-29 · Hold

Proposed and held: a Pipelines virtual-view criterion, gated on evidence.

Decision (Proposed, not accepted). A tenth Pipelines criterion candidate,metadata_realization, would score whether a pipeline physically moves bytes or synthesizes a virtual view over a live source, because the write-contract axis (file-write for Iceberg, SQL-transaction for DuckLake, never-write for Streambased ISK) is where the streaming economics live and the four scored candidates all physically move bytes. Streambased ISK is the candidate that would enter under it. Both stay out of the scored YAML, this record is the tracked gap rather than an accepted claim, and the scored ranking is unchanged by it.

Why this stays Proposed. Three gates hold it. No independent correctness and performance benchmark of virtual Iceberg exists (the two owner-input-free arms of BENCH-D were measured and independently reproduced on 2026-06-28, but they cover the file-write and SQL-transaction contracts, not the never-write arm). Two DuckLake correctness blockers remain open (#1215 silent row resurrection, #1184 the wall above 1,600 columns on CREATE). And no named production virtual-Iceberg deployment exists, against a materialized-REST baseline (Tableflow, WarpStream) that is production-proven with named consumers. The never-write arm needs Streambased vendor access the lab does not have, so this gate may be unclosable without it, which is stated here rather than papered over.

MDR-0027 · Accepted 2026-06-17

The incumbent and the partial moves become scored candidate paths.

Decision. Four standing paths compete on one surface: P0 Stay (the status-quo SIEM, the do-nothing baseline every other path is scored against), P1 Augment (keep the SIEM for hot and detection, offload cold and retention to the lakehouse), P2 Hybrid-tiered (ClickHouse hot plus Iceberg long), and P3 Full-replace. The breakeven is computed, never asserted:crossover_months = migration_cost ÷ (monthly_cost(stay) − monthly_cost(path)), where monthly cost reconstructs from the cost-to-serve-at-retention cells that already exist and migration cost is the one new input, labeled the softest term (Tier C/D judgment until a client's source inventory is scoped). The output is a breakeven curve over the buyer's horizon, not a bare point estimate.

Why this won. The 2026-06-17 maturity assessment found, across all four expert lenses, that the Matrix answered "which open stack if I move" and underserved "should I move at all," so a CISO could not take it to a board for a go/no-go. Hand-asserting a breakeven would have been faster, but a technically-literate client recomputing from the published cost-to-serve cells would find a number that does not reconstruct, which is the exact integrity failure the audit-not-opinion rule exists to prevent. The risk gate rides on Move #3: a path that fails a detection-survivability or compliance gate is not eligible to win the crossover on cost alone, because a full-replace that drops 30% of detections is not cheaper, it is broken.

MDR-0028 · Accepted 2026-06-17

Every recommendation carries a four-part board-defensibility read.

Decision. A fixed-shape, one-paragraph read ships with every engagement recommendation, four parts in order: the call (which path, at this client's retention horizon and archetype weights); what being wrong costs (the reversibility kill-switch number, cited from the existing Service 1 reversibility output rather than re-derived, so two figures cannot drift); the evidence tier behind the crossover (measured-at-scale, modeled, or judgment, with the migration-cost input flagged as the softest term); and the flip-sensitivity, the single assumption that flips the recommendation if it moves. A recommendation missing any part fails review.

Why this won. A board does not adopt a month-count. It adopts a recommendation and asks what backs it, what it costs to back out, and what would change the answer, and per-engagement prose leaves all three to chance. The pattern follows MDR-0016's disclosure lineage, mandatory fixed structure with per-engagement content, and it forces the single-host Tier-B evidence reality to be stated in front of the board rather than hidden inside a point estimate. A longer sensitivity analysis was considered and rejected because it buries the one assumption that actually flips the call under a table the board will not read.

MDR-0029 · Accepted 2026-06-17

Component 5 scores detection-survivability. Survive means executes correctly.

Decision. Component 5 becomes a scored component, and detection-survivability enters as its sixth criterion (from five, inside the 3–14 bound): what fraction of the client's existing detection corpus survives the move, under a three-band model. Survives means the target executes the rule with correct semantics. Silently degrades means the rule compiles and ports but loses a primitive, so it runs while no longer enforcing its own logic, and that band is scored worse than a clean cannot-port failure because nothing downstream announces the loss and the dashboard stays green. The engagement-internal score runs against the client's actual corpus, so client numbers stay with the client; the public surface carries the method, the three bands, and the silent-degradation evidence.

Why this won. "Ports cleanly" is easy to fake, because a rule that compiles on the target looks ported, and the first-party SIGMA-EXEC evidence (Tier B, single host) shows compiling is not executing: pySigma compiles a correlation rule's count logic to a windowless query that over-fires while the dashboard looks healthy. Scoring silent degradation equal to a loud failure would actively mislead, since a loud refusal gets seen and redesigned while the invisible loss just runs wrong. Losing detection coverage mid-migration is the catastrophic risk for a SOC, the one no cost advantage offsets, so the Matrix has to score it to be a security-data-architecture instrument rather than a lakehouse comparison.

MDR-0030 · Accepted 2026-06-17

Compliance / WORM: the first gate-capable cross-cutting axis.

Decision. Compliance and evidentiary posture is scored 1–5 per candidate stack-wide, the same shape as Practitioner-Ownability and Cost-to-serve, so it costs no per-component dimensionality. Five sub-criteria: WORM immutability (17a-4(f) / SEC / DORA), legal-hold, chain-of-custody and audit-trail completeness, residency enforcement, and retrieval-SLA at cold tier. The new mechanic is the gate: where a client's regime mandates a control, a candidate that cannot satisfy it is disqualified before weighted totals are computed, reusing the Go/No-Go lineage MDR-0024 established, rather than receiving a low score a strong cost figure could outweigh.

Why this won. For a regulated buyer, failing WORM is a disqualifier no cost advantage offsets, so a compliance score that can be outvoted by the weighted total is a wrong recommendation waiting to happen. The honest limit is carried on the surface rather than buried: the audit-trail sub-criterion anchors to Iceberg V3 row-lineage, which is blocked on the OSS pyiceberg-plus-Nessie path today (pyiceberg #1551 open; Nessie 0.107.5 silently downgrades V3 to V2; Tier-B first-party smoke test), so the regulated buyer's exact need is only partially reachable on the open stack, and the Matrix says so instead of implying row lineage is free.

MDR-0031 · Accepted 2026-06-17 · overtaken 2026-07-11

The illustrative-worked-example exception, ratified and then overtaken.

Decision. A scored illustration could appear on the public surface only when five conditions held together: illustrative archetype inputs, never a client's; the vendor-claim-versus-shipped delta redacted; no real-client scored ranking; labeled illustrative at the point of use with evidence tiers carried; and method shown, not verdict substituted. It wrote the boundary the worked scorecard and the Move-2/3 demonstration pages already relied on into the rule, so each new illustration stopped being a fresh judgment call.

Why the record stands. The 2026-07-11 publication ruling then made the full matrix public (scores, reasoning, and deltas), so the public/paid line this exception refined no longer exists and the record stands as history rather than an active rule. What survives it is the discipline it encoded: illustrative inputs are still labeled illustrative, evidence tiers still travel with every grounded cell, and the only figures withheld now are the DeWitt / EULA-bound benchmark measurements, which is a legal constraint on naming vendors in published measurements, not a paid gate.

MDR-0032 · Accepted 2026-06-22

Move #2 ratified: the scored augment-vs-replace crossover.

Decision. The owner ratified Matrix Move #2 as specced: build the scored stay-versus-go crossover and the board-defensibility read (MDR-0027 and MDR-0028) into the Matrix, converting it from a shortlisting tool into the replace-or-augment instrument the framing promises. This record is deliberately summary-level here, because the crossover's outputs are computed per engagement on a client's actual costs, horizon, and weights; the method is public above, and the client numbers are the engagement.

Why this won. It closed the first of the two missing scored decisions from the 2026-06-17 maturity assessment, whose unanimous HIGH gap was that the incumbent appeared only as a foil being replaced, so a CISO could not defend a go/no-go at a board. Selling the crossover as a per-engagement custom output was considered and rejected as the default, because the scored object belongs in the product so the Matrix is self-defensible; the engagement deepens it rather than replacing it.

MDR-0033 · Accepted 2026-06-22

Move #3 ratified: detection and compliance become scored, not prose.

Decision. The owner ratified Matrix Move #3: detection-survivability (MDR-0029) and the compliance/WORM axis (MDR-0030) fold into the scored Matrix as first-class dimensions, weights folded into archetype totals under the standing rubric (dimension band 3–14, weights sum 100). With MDR-0032 this closes the assessment's scored-decision gap. As with Move #2, the record here is summary-level on purpose: the per-client corpus survivability fraction and the per-client gate verdict are engagement work, so this page carries the ratification and the method, not those numbers.

Why this won. A security buyer defends detection and compliance to an auditor or a board, so they have to carry weight in the total rather than sit beside it as context. Scoring them as one combined security-fit dimension was considered and rejected, because detection survivability and compliance WORM gates fail independently and on different evidence, so they score as two axes, not one.

MDR-0034 · Accepted 2026-06-24 · band provisional

The first measured re-score: the count-correlation band moves down.

Decision. The revalidation trigger set when Move #3 scored Component 5 fired, and it fired in the refute direction the pre-registration named. The SIGMA-EXEC lakehouse leg (pre-registered and frozen before the run) executed the canonical Sigma event-count correlation on all five SQL lakehouse engines, and every engine compiled it to windowless SQL with the ten-minute timespan dropped, while every hand-written windowed control on the same engines fired correctly. So the detection-survivability band for the count-correlation family (event_count,value_count) drops one band across the engines, its evidence tier moves C to B (now measured, per MDR-0003), and the band-4 cap comes off because the band is no longer an extrapolation. The relative engine ordering inside the band is preserved, because the band also carries orthogonal legs (operational ceiling, federation reach, serverless cadence) the measurement does not touch, and the engine spread stays where the build already carried it, on compilation fidelity. Applied to the three scored files with the scoring linter passing.

Why the band is labeled provisional. A follow-on conditioning pass corrected the framing without moving a value. The 0.286-precision over-fire figure is the unbounded-scan extreme, not a rate: run under a tumbling scheduler at realistic lookbacks, the verbatim windowless emit fires correctly on this corpus, so the coverage risk is real but deployment-contingent. The window-drop is a conformance gap rather than a design choice (the Sigma Correlation Rules Spec v2.1.0 marks timespan mandatory, Tier A), it is untracked upstream at current versions (pySigma 1.3.3), and the measured evidence is one backend design lineage on a single synthetic corpus. That is why the band carries a PROVISIONAL / Tier-B-single-corpus label, and why H-SIGMA-01 holds at 3/5 with no confidence move banked.

MDR-0035 · Accepted 2026-06-24 · re-score declined

The band is plugin-contingent. The re-score was declined anyway.

Decision. A positive control the original measurement lacked: the SigmaHQpySigma-backend-athena plugin (v0.2.0) emits a real enforced window (a RANGE frame over a 600-second interval) for event_count, and executed over the same planted corpus it fires the 20 true bursts at precision 1.0 with zero decoy false positives, while removing only the window frame reproduces the windowless over-fire. So the window-drop is a property of the backend plugin's templates, not of the engine or of SQL, and a single engine can land in different survivability classes depending only on which plugin compiles its rules. The proposed re-score (raise the Presto-family engines whose available plugin emits the window) was surfaced to the owner and declined; no scored cell moved, and the finding is recorded as a capability note in the scoring files.

Why the owner declined. The positive control is deliberately weighed as too thin to move a paid band: it covers event_count only (the plugin leaves value_count andtemporal_ordered unimplemented), it ran on in-process DuckDB with two disclosed dialect adaptations rather than production Athena or Trino, against one synthetic corpus, on a backend with no PyPI release. The scoping ruling underneath is the durable part: survivability scores configuration-as-deployed, the floor a typical client gets through the generic windowless backend, not capability-available, and the flip-conditions for revisiting are named in the record (a windowed backend covering the full count family, run on a production Presto-family engine, at released maturity, against a real client rule mix).

MDR-0036 · Accepted 2026-07-01 · executed

The DSMOS inventory moves into the Matrix's own repository.

Decision. The Data-Source MCP Ownership Score inventory, the per-producer score components behind Component 8's bands for all 31 scored data-source producers, migrated out of the archived security-architect-mcp-server repository and into the Matrix's own tree asdsmos-inventory.json, ending the scoring's dependency on a repository marked for retirement. This is a data-custody move, not a re-score: no score, band, or published number changed.

Why this won. Component 8 is a scored, canonical axis, and leaving its evidence source in a repository that could be deleted or go inaccessible without warning is exactly the silent single-point-of-failure the audit-not-opinion discipline exists to avoid. The migration was verified before anything was written: the full 31-producer inventory round-tripped against the cited rollup with no drift (grade distribution A15 / B6 / C7 / F3; CrowdStrike at 77/B still the best commercial-SaaS producer; the three Gartner NDR Leaders with no MCP at all still at F), and every score component came across intact, so the bands still reconstruct from their parts.

MDR-0037 · Accepted 2026-07-20 · v1.4 evidence cut

Engines concurrency gets re-scored on a measurement instead of an absence.

Decision. StarRocks concurrency_multi_tenant_behavior moves from 2 at Tier C to4 at Tier B at Archetypes A and B, on the first-party concurrency-multiuser bench that ran 2026-06-15 (QPS ceiling 6.79, tail p95/p50 about 1.16 at N=16, 1.27× headroom, zero errors, the best-behaved arm in the run). The same cut brings every other engine's concurrency cell current to the measured record rather than moving its score: ClickHouse and Trino gain the first-party citation with the single-host caveat stated and their 4s holding, the Dremio cells note that the arm ran with its figures withheld under that vendor's benchmark-publication terms, DuckDB's cells note its architecture-driven exclusion from the concurrency bench, and the Trino Archetype-B cell re-tiers from B to A on the named-operator Pinterest production citation that MDR-0003 treats as Tier A. Component 3's evidence vintage is now 2026-06-15, and the stale engines evidence-completeness figure rides the same cut from 0.875 (7 of 8, pre-v1.3) to 0.889 (8 of 9 after MDR-0023).

What it moves. StarRocks at Archetype A goes from bounds of 2.57–2.77 to2.87–3.07, and because the new lower bound clears DuckDB's upper bound of 2.83 cleanly the published ordering changes rather than merely tightening: ClickHouse (3.65–3.85) leads, Trino (3.41–3.61) follows, and StarRocks now sits above DuckDB (2.63–2.83) instead of below it, which moves StarRocks to third on the four-candidate public table. At Archetype B the total moves from 2.70 to 2.80 with the ordering unchanged, and the Archetype C cells were already 4 at Tier B, so they only gain the first-party citation.

Why this won. The 2 rested on a stated ground ("no published data on concurrent Archetype A workload") that stopped being true the day the bench published, and Archetype C had already carried StarRocks at 4 on vendor production references, so the A and B cells sat two points lower on the claim that the evidence C accepted did not exist. Holding at 2 with a refreshed citation was rejected as indefensible, since the new evidence removes the cell's own justification and points up rather than sideways. A 3 under-credits the arm the bench measured as best-behaved. A 5 was refused because a single-host, single-shape, closed-loop, OSS-edition result with 1.27× headroom does not establish design-center concurrency, and MDR-0002 reserves 5 for that. The transfer assumption, that single-host closed-loop behavior on one scan-aggregation shape is informative for an archetype's mixed scheduled and interactive load, is written into every re-scored cell as a caveat rather than left implicit. The falsifier named here landed the same day: the Trino cell did come under pressure and moved, which MDR-0038 below records.

MDR-0038 · Accepted 2026-07-20 · v1.4 evidence cut

The headroom ratio was flattering the slowest engine.

Decision. Trino concurrency_multi_tenant_behavior moves from 4 to3 at Archetype A, evidence tier B unchanged, which takes its weighted band from 3.56–3.76 to 3.41–3.61. Archetype B stays at 4 and Archetype C stays at 3.

Why this won. MDR-0037 left Trino alone, and that left Archetype A with ClickHouse, Trino and StarRocks all sitting at 4 on the criterion the same bench separates most clearly, so the heaviest-weighted concurrency criterion had stopped discriminating. Trino served about 3 queries per second against roughly 6.7 for the other two, with a p50 of 12.5 seconds and a p95 of 21.1 seconds at 32 clients where both others stayed under 8. Its 1.32× headroom reads best of the columnar arms only because headroom divides by the single-client baseline and Trino starts slowest, so the ratio was measuring the wrong thing. A 3 rather than a 2, because Trino completed every query with zero errors and scaled positively, which is a criterion met competently. A 3 rather than a 4, because 4 claims a differentiating strength and the only first-party measurement runs the other way, while the support for the 4 was architecture documentation plus an operator running 17,000-plus nodes, neither of which speaks to a SOC-scale worker fleet under mixed load.

The same reasoning is why ClickHouse did not move. Its 1.10× headroom is the worst in the run, but saturating at a single client is what the all-cores-per-query design does rather than a failure, and it delivered the highest absolute throughput with the second-tightest tail, so the headroom ratio is the wrong instrument for it in the mirror image of the way it was wrong for Trino.

What it moves. On the published four-candidate table the ordering is unchanged and Trino still holds second behind ClickHouse, though the gap it opened is real rather than the narrow one the v1.3 cut showed. Archetype B keeps its 4 on weight 5 and on the Tier-A named-operator Pinterest citation, with the worker-fleet assumption written into the cell as a caveat, which is archetype-conditional scoring working as MDR-0004 intends rather than the same engine graded two ways by accident. The falsifier is a cluster-slice re-run at Archetype A's envelope: if a real worker fleet puts Trino in the same band as the other two, the cell returns to 4.

MDR-0039 · Accepted 2026-07-20

A fact you can check in one step is not the same as a claim.

Decision. A verifiable product-spec fact counts as Tier A, and the definition is deliberately narrow: it has to state the existence, structure, or licensing of a shipped capability rather than its performance, it has to be checkable at a primary source the reader can reach, checking it has to be one step with a yes-or-no answer, and it has to be version-bound. Trino's connector listing, DuckDB's embedded-library architecture, and Athena's native Glue integration all qualify. Anything about how fast, how large, how cheap, or how reliably that capability performs does not, and falls back to the ordinary rules.

Why this won. The tier-enforcement pass rewrote every legend on the site to state MDR-0003 faithfully, which was right, and it exposed a category the record had never addressed: a number of cells carried Tier A on product-structure facts that are neither peer review, nor a standard, nor a named-operator production reference. The retired vendors-page legend had covered them with a spec-fact clause, and that legend was removed because the rest of it was wrong. Re-tiering all of them to B was the obvious alternative and it was rejected, because it would put a fact anyone can verify in one step in the same tier as a conference talk, which inverts the credibility ordering the scheme exists to express. Adding a fifth tier was rejected too, since the five-tier scheme had just been retired for good reason.

What it moves. No score changes, because tier and score answer different questions. The rule cuts both ways and demoted more cells than it protected on the day it was written: a first-party SDW essay is never the Tier-A basis for a cell, which took three Cribl cells and the Glue Archetype-C cost cell down to B, and an anonymized reference is not a named operator, which took the Dremio Archetype-B raw-performance cell down to B however large the deployment behind it. The falsifier is built in — if a cell tiered A this way rests on a capability that turns out not to exist as described, or existed only in a version the cell does not name, the category should collapse into Tier B.

How the parts fit together.

Foundational MDRs (1–5, 12–16) set the structure. Per-component criterion sets (6, 8, 10, and 29 for Component 5) define what gets scored. Per-archetype weight tables (7, 9, 11, 17, 18, 19, 20, 21, 22) are where the workload-conditional "best" claim lives. The three v1.3 additions close evidence gaps (23 and 24 now Accepted, with 24's YAML application deferred to the refresh trigger; 25 still Proposed). The migration-instrument records (26–31) add the decision instrument, and the records after them (32–37) ratify the two Moves, apply the two measured re-scores, settle the registry's data custody, and stamp the v1.4 engines evidence cut of 2026-07-20.

The Q3 2026 catalog benchmark, the remaining slices of the OLAP bake-off (the single-host engine-join slice ran 2026-06-10; the cluster, federated-join, and AWS-native scenarios are still open), and the ~Q1 2027 H1-COST-02 pipeline benchmark are the primary refresh anchors. If any of those shows a weight is over-rewarding a candidate that fails in practice, the relevant MDR gets superseded, not amended silently.

The rest of the Matrix.

Scoring decisions sit alongside the decision records, the component criteria, the evidence discipline, the decision framework, and the in-flight assumption tracker, each of which names the MDR a cell would have to move through before its score changes.

Back to the Matrix →