Capability Matrix · Decision records
How the Matrix runs. On the record.
The Matrix is itself an engineered artifact, so the decisions about how it runs are recorded with the same discipline as the scores it produces. Each section below is a meta-decision of the live Matrix, dated between 2026-05-25 and 2026-07-11 and cited to its Metric Decision Record (MDR) in the project decision store, a format adapted from Michael Nygard's architecture decision records: numbered sequentially, numbers never reused, supersession tracked in both directions by a validation script. The registry holds 39 records as of 2026-07-20, and each one carries its context, the alternatives it rejected, and its consequences, so if I change my mind the change shows up where the prior decision lives.
The scoring-layer detail per component lives on thescoring decisionspage; this page holds the frame those decisions operate inside, and it closes with the one ruling that changed what you can read here.
MDR-0002 · Accepted 2026-05-25
Score on a 1–5 scale, with null permitted.
Decision. Every criterion is scored 1 through 5, where 5 is best in class for the criterion in the archetype under evaluation, 4 is a strong and production-defensible fit, 3 is a competent default that doesn't differentiate, 2 meets the criterion minimally with significant caveats, and 1 is an architectural mismatch. When the evidence doesn't support a defensible score the cell is null, and a validation rule (citations_required_when_scored) requires at least one citation on every score that isn't null.
Why this won. The rejected alternatives are on the record. A 0–100 scale offers false precision, because practitioners anchor on round numbers and defending a 73 against a 76 is harder than defending a 4 against a 4. Likert agreement scales fit opinion surveys, though the criteria here are operational. Five-star ratings carry the same information with worse null handling. And Pass / Conditional / Fail is too coarse, since it collapses good-but-not-best and best-in-class into one bucket, which is exactly the differentiating information the Matrix exists to surface.
MDR-0003 · Accepted 2026-05-25
Every citation carries an A–D evidence tier.
Decision. Tier A is peer-reviewed research, official standards (NIST, ISO, the OCSF spec body), and production deployment at recognized scale; Netflix, Pinterest, and AWS Security Lake are the record's own examples. Tier B is practitioner engineering writing from credible sources, conference talks (Strata, BSides, Black Hat), and multi-source synthesis where two B-tier sources agree. Tier C is vendor blog or marketing material, cited only with an explicit bias-flag annotation. Tier D is speculation or hypothesis, and it can never be the sole citation behind a non-null score: the no_tier_d_uncorroborated validation rule requires at least one A or B citation alongside it. A production reference earns A when the operator is named and the scale is recognized, though an anonymized reference (a "Fortune-100 deployment") or an NDA-gated result can't be independently checked, so it sits at B at best.
Why this won. Without tiering, a vendor benchmark post and a Netflix production reference would carry the same evidence weight, and the Matrix's credibility lives in that distinction. The tier is also independent of the 1–5 capability score, because the two answer different questions: the tier says how credible the backing is, while the score says what the tool does for the criterion, so a 5 backed by a Tier-C citation is less defensible than a 3 backed by Tier-A. One rule follows from the vendor tier: when a Tier-C vendor claim exists for a criterion, the scoring file must carry a shipped_vs_claim_delta recording the gap between the claim and the measured reality, or stating that measurement confirmed it.
MDR-0005 · Accepted 2026-05-25
Every component carries 3 to 14 criteria.
Decision. Every component's criterion count must be strictly greater than 2 and strictly less than 15, so 3 ≤ N ≤ 14. Too few criteria collapses real trade-offs into a fortune-cookie ranking, and too many dilutes weight discrimination below the noise floor, because at 15-plus criteria the weights cluster around 5–10 each, which is below the threshold where a 3-versus-4 judgment carries meaning. The original cap was under 16 and was tightened to under 15 the same day (2026-05-25), and the record carries a concrete over-dilution trigger: if any criterion in a published matrix has a weight below 4, the dimensionality is probably too high.
Where the counts stand. At acceptance the three scored components complied at 8 (Engines), 10 (Formats + Catalogs), and 9 (Pipelines) criteria. The Engines count has since moved to 9: MDR-0023 added native_ip_type_support as the ninth criterion, created 2026-05-25 and accepted 2026-06-19 on the v1.3 refresh trigger, still inside the band. So where MDR-0005's own text says the engines taxonomy is 8 criteria, it's recording the count on its acceptance date rather than today's.
MDR-0004 · Accepted 2026-05-25
Weights are archetype-conditional. A single global ranking is never computed.
Decision. "Best" is workload-conditional: the same criterion that dominates one archetype (raw analytical performance for a Zeek-heavy SOC with a p99-under-5s target) is an afterthought for another (a multi-source federated lakehouse), so every component is scored separately per archetype. Criterion definitions stay constant across archetypes while the weights vary, engagements are assigned to their closest published archetype rather than negotiating weights per deal, and when an engagement fits no published archetype that mismatch is itself a finding: either the practice adds an archetype, or the engagement gets custom weights with an explicit deviation note in the scoring file. A single weighted total across archetypes is meaningless and is never computed. The record calls this the single most important methodology choice in the Matrix, because without it the Matrix collapses to a global ordering with no defensibility.
The anchoring rule. The weight tables themselves run under a rule that entered the record with the first archetype weight table (MDR-0007, Engines Archetype A, accepted 2026-05-25): weights sum to 100, no single weight above 30, none below 5, all in multiples of 5. That floors false precision and prevents tuning a weight table until a predetermined winner falls out.
MDR-0012 + MDR-0013 · Accepted 2026-05-25
Null propagates to the total. Bounds preserve the unknown.
Decision. When any criterion in a candidate's column is null, the weighted total is null too. Instead of a point estimate the scoring file emits two bounds, a lower bound that treats every null as a 1 and an upper bound that treats every null as a 5, so a reader sees "ClickHouse 3.55–3.75" rather than a single number hiding what wasn't measured. The rejected alternatives all leak information somewhere: skipping null criteria silently re-normalizes the weights, treating null as zero over-penalizes, and treating null as an average 3 fabricates the very score the null exists to refuse.
The publication gate. MDR-0013 sets the threshold on top of that: a scoring run is v1-publication-ready only when at least 75% of its criteria carry non-null scores (at 8 criteria that means 6 or more), because below that the bounds spread wide enough that the ranking doesn't survive sensitivity analysis. At acceptance all three scored components cleared it, at 0.875 (Engines), 1.0 (Formats + Catalogs), and 0.889 (Pipelines).
MDR-0030 + MDR-0029 · Accepted 2026-06-17
Two security axes can gate, because some failures shouldn't be averaged.
Compliance/WORM (MDR-0030). The first gate-capable cross-cutting axis. It scores 1–5 across five sub-criteria: WORM/immutability (17a-4(f) / SEC / DORA), legal-hold, chain-of-custody and audit-trail completeness, residency enforcement, and retrieval-SLA at cold tier. Where the client's regime mandates a control, a candidate that cannot satisfy it is disqualified before weighted totals are computed rather than carried as a low score a strong cost number could outweigh, because for a regulated buyer a mandatory control is a requirement and never a preference to trade off. The gate mechanism reuses the Go/No-Go lineage MDR-0024 established, and the record states its honest limit: the audit-trail sub-criterion anchors to Iceberg V3 row-lineage, which is blocked on the OSS pyiceberg + Nessie path today (pyiceberg #1551 open; Nessie 0.107.5 silently downgrades V3 to V2), so the regulated buyer's exact need is partially unreachable on the open stack and the Matrix says so.
Detection-survivability (MDR-0029). The second axis a recommendation can't win past. Component 5 scores what fraction of the client's existing detection corpus survives a move under a three-band model: survives means the detection executes with correct semantics on the target,silently degrades means it compiles and runs but loses a primitive (a dropped time window that over-fires or under-fires while the dashboard stays green), and cannot port is the loud, manageable failure. The silent band is deliberately scored worse than the clean failure, because a loss nobody can see is the one no cost advantage offsets, and "survives" means executes correctly rather than merely compiles. Both axes feed the augment-vs-replace decision path: a route that silently degrades a meaningful fraction of the detections, or fails a mandatory compliance gate, is ineligible to win the crossover on cost alone.
MDR-0015 · Accepted 2026-05-25
Scores expire on a schedule.
Decision. Scores age, because vendor releases, benchmark publications, and partnership changes all invalidate prior cells silently. Every score carrieslast_reviewed and revalidate_by dates with a default cadence of six months, which matches the median change rate observed across engine major versions, catalog GA timelines, and table-format releases. Four events trigger a faster pass: a candidate ships a major version, a practice partnership forms or changes, a benchmark from the lab plan lands, or a practitioner surfaces a shipped-versus-claim delta that voids a previous score. Structurally stable methodology assumptions may run at twelve months, and stale entries are flagged by the validation script rather than trusted.
MDR-0016 · Accepted 2026-05-25
Every scored candidate carries a disclosure block.
Decision. Fair-broker positioning only survives audit if a reader can check for partnership bias without asking, so every scored candidate carries a disclosure block even when there's nothing to disclose (the block is then explicitly null). When a relationship exists, the block records its nature (an in-progress conversation, an active commercial partnership, or a former employer), its scope in plain language, and its impact on the score; when scores are materially affected, the impact field explains how and an MDR captures the methodology adjustment. A validation rule enforces a non-null block whenever any active partnership exists, and an escalation from conversation to commercial is recorded in the same commit as any scoring change it touches.
Publication ruling · Recorded in the method canon 2026-07-11
The Matrix went public, and the record says so plainly.
What changed. For most of its life the Matrix ran a public-method / paid-scores line: methodology public, per-criterion scores engagement-internal (MDR-0014), with a bounded exception (MDR-0031, accepted 2026-06-17) that let illustrative worked examples carry scored cells in public if five conditions held, namely that the inputs were an archetype's rather than a client's, the vendor-claim-versus-shipped delta column stayed redacted, no real client ranking appeared, the illustration was labeled at the point of use with its evidence tier carried, and it showed how a number is produced rather than substituting for the per-client scoring. On 2026-07-11 I changed the line itself: the full Matrix (methodology, candidate catalog, per-vendor weighted scores, criteria-by-criterion reasoning, and the claim-versus-shipped deltas) is published as the practice's flagship evidence asset. Only benchmark measurements bound by vendor DeWitt/EULA benchmark-restriction clauses stay withheld, which is a legal constraint on naming vendors in published measurements rather than paid IP, and the paid work is the services engagement that applies the Matrix to a client's environment: their weights, their totals, the recommended bundle, sequencing, and reversibility.
Where it's recorded. That ruling lives in the method canon as a dated pivot banner (2026-07-11) rather than as a numbered MDR, and this page says so instead of inventing one, because the whole point of a decision record is that it shows what actually happened. MDR-0031 stands as the accepted precursor whose worked-example boundary the ruling made moot on the score side, and the DeWitt withholding convention it left in place is visible wherever a cell says "figure withheld per vendor benchmark-restriction clause."
Page lineage · MDR-0036 · Accepted 2026-07-01
What this page used to say.
Until 2026-07-19 this URL carried six architecture decision records for a different artifact: the MCP-server vendor-lookup tool that predated the published Matrix, covering its SDK choice, its vanilla-JavaScript web tool, its filtering UX, and its own evidence rubric. That tool was archived on 2026-07-01 as fully abandoned, and MDR-0036 (accepted the same day) records the one dependency that outlived it: the DSMOS data-source inventory (31 producers with their score_components) was migrated out of the archived repo into the project's own data store, verified round-trip-clean against the published rollup before the copy. The six tool ADRs remain in this page's git history; the decisions above are the record of the Matrix that's actually live.
Decisions that change get re-recorded. The history is the audit trail.
The full scoring-layer registry, MDR-0001 through MDR-0037, is detailed on the scoring-decisions page. Component criteria, vendor evaluations, the archetypes, and the cross-archetype synthesis cover the Matrix in working detail, and the hub carries the methodology summary with the worked scorecard.