Capability Matrix · Foundational assumptions
Seventeen foundational assumptions. Each one has a refutation criterion.
Every MDR rests on premises. The dangerous assumptions are the ones nobody wrote down. This page lists mine, with confidence, status, basis, and the concrete observable that would refute each one. If a premise flips, the MDRs that depend on it flag for re-examination. You should be able to audit my premises the same way you audit my scores.
v1.3 cut · 2026-05-25
A-05 moved from open to confirmed.
On 2026-05-25 I formally promoted A-05 (Unity Catalog self-hosted remains operationally immature) from open to confirmed on the basis of two distinct sources of first-party evidence: deployment-maturity friction in the OSS Helm chart and Docker image pipeline, and a documented Iceberg-REST functional asymmetry where the OSS server accepts namespace and table reads but refuses POST writes. The same cut also strengthened A-04 (Polaris) and A-06 (Iceberg V3 engine support) with first-party observations. Confidence levels unchanged, basis citations upgraded.
A-10 · In-house skill base is structurally client-specific
The existing_in_house_skill_base criterion cannot be scored at standing-matrix level. It requires engagement context.
confidence: high · status: open · basis: convention · last reviewed: 2026-05-25 · review cadence: every 12 months
Per the fair-broker positioning, I keep the criterion in the framework at score null rather than removing it. The reasoning: a customer's actual engineering skill mix is unknowable from outside the engagement, but dropping the criterion entirely would create a silent gap during engagement scoping. The rule is "held null but visible" so engagement teams cannot forget to ask. The central claim is that this is structurally unscorable rather than scorable from public signals.
Refutation. A structural pattern emerges where in-house skill base can be inferred from non-engagement signals at standing-matrix level. For example, vendor official-skills certifications mapped to a known practitioner population, or a third-party labor-market dataset with sufficient granularity.
Dependent MDRs: MDR-0002 (scoring scale), MDR-0006 (engines taxonomy), MDR-0010 (pipelines taxonomy).
A-06 · Iceberg V3 engine support rolling through 2026
V3 features (row lineage, deletion vectors, variant type) are rolling across engines unevenly. Trino leads. ClickHouse and DuckDB lag.
confidence: high · status: open · basis: public-data · last reviewed: 2026-05-25 · review cadence: every 3 months
Iceberg v1.8 through v1.10 shipped V3 features through 2025; engine-side adoption is per-engine and per-release. I review this on a 3-month cadence, not 6, because the format-war state changes quickly.
First-party stack smoke (2026-05-25, v1.3 cut). I ran V3 row-lineage end-to-end and observed a sharper picture than the public release notes suggest. The refined per-stack reachability table:
- Spark 3.5.6 + Hadoop catalog + Iceberg ≥1.11: V3 row-lineage reachable end-to-end;
_row_idpopulates correctly across snapshots. - pyiceberg 0.11.1 + Nessie 0.107.5: two stacked blockers. pyiceberg cannot write V3 at all, and Nessie silently downgrades V3 to V2 at the catalog layer.
- DuckDB
iceberg_scan(): does not expose_row_ideven when explicitly requested. Engine-reader-side gap independent of writer or catalog. - Polaris ≥1.5.0 and Databricks Unity: unknown, not yet tested.
Refutation. An independent benchmark or feature-matrix audit shows ClickHouse or DuckDB shipping V3 row-lineage parity with Trino, or Iceberg V4 ships before V3 reaches parity, which would alter the lag landscape entirely.
Dependent MDRs: MDR-0008 (formats + catalogs taxonomy).
A-01 · ClickHouse benchmark generalizes from single-node to cluster + concurrent
The ClickHouse-over-schema-on-read-SIEM Zeek result (a triple-draw band, 2026-06-14: ~10–11× foil-over-Iceberg on the scan-heavy aggregations, a ~4.2–4.6× open-format tax, with native ~47× / Iceberg ~10× as the five-query-average arm-pairings of the same numbers) generalizes from single-node sequential workloads to distributed cluster plus concurrent analyst load.
confidence: medium · status: open · basis: inference · last reviewed: 2026-05-25 · review cadence: every 6 months
This is the most consequential single assumption in the Engines matrix. The H3-PERFORMANCE-01 measurement (ClickHouse 46.8× faster than a schema-on-read SIEM on the five-query average over a 10M-event Zeek corpus, the native arm-pairing of the triple-draw band with the Iceberg arm at 10.1×; 5–62× on the hunting-shaped queries across the native and Iceberg realizations with the index winning the simple lookups; answer-equality verified, single-node Tier B) is the strongest anchor I have on the engine-performance side. The measurement shape is single-node sequential. The inference (that ClickHouse's sharding-plus-replication architecture preserves the per-shard query characteristics under cluster and concurrency) is reasonable but unmeasured. Cluster overhead from network shuffles on cross-shard joins could compress the ratio.
Refutation. The cluster + concurrent-query regime (multi-node, ~30 concurrent queries) is where the two-regime advantage could compress: it might pull the hunting-aggregation multiple over the schema-on-read SIEM baseline (OpenSearch 2.18.0, the canonical foil; the 3.7.0 currency check was performance-neutral) below ~5×, or surface a regime where ClickHouse loses to Trino on the recurring-query envelope of Archetype A. The single-host engine-join bake-off has already run (2026-06-10, Tier B — the open engines answered the SOC suite under 1.5 s); the cluster regime remains the gating event, unmeasured.
Dependent MDRs: MDR-0007 (engines Archetype A weights).
A-02 · Cribl 70-90% is aggressive-tuned, not default
Cribl's marketed 70-90% SIEM-cost reduction is achievable but requires aggressive tuning. Default out-of-box reduction is 30-50%.
confidence: high · status: open · basis: public-data · last reviewed: 2026-05-25 · review cadence: every 6 months
Multiple independent practitioner data points converge on the 30-50% default range. RiverSafe (UK consultancy) independently measured 40% Cribl SIEM-licensing reduction. Cribl's own April 2026 economics analysis shows 1.8% savings at 10 GB/day, 35% at 100 GB/day, and 64% at 1 TB/day; the headline 70-90% reflects aggressive tuning at the high end (firewall denies aggregated 99%, network flows sampled 98%). Workload shape matters: EDR telemetry stays near 90%+ pass-through; network flows reduce 98% by sampling; firewall denies reduce 99% by metric aggregation. The 30-50% default reflects a balanced mixed workload.
Refutation. A head-to-head benchmark shows Cribl default (no custom tuning, out-of-box routing) reduces a representative Zeek plus EDR plus firewall workload by 70%+.
Dependent MDRs: MDR-0010 (pipelines taxonomy), MDR-0011 (pipelines Archetype A weights), MDR-0022 (pipelines Archetype C weights).
A-04 · Polaris ecosystem will mature through 2026
Apache Polaris adoption accelerates as ASF governance and Snowflake commercial momentum compound. The current thin reference base is transitional, not structural.
confidence: medium · status: open · basis: inference · last reviewed: 2026-05-25 · review cadence: every 6 months
Polaris entered the Apache Incubator in August 2024 and graduated to an Apache top-level project on 2026-02-18; Pinterest's production deployment (presented at Iceberg Summit 2026) remains the first named Tier-A reference. The momentum signals are real. Apache published the first official Helm chart apache-polaris-1.5.0 on 2026-05-18, eight days before this review.
First-party friction characterization (2026-05-25, v1.3 cut). The chart ships, but it requires k8s 1.33+ (bleeding-edge), is not yet on any Helm registry (install from cloned source), defaults to in-memory persistence (no documented production Postgres-backed path in 1.5.0), and the docker POLARIS_BOOTSTRAP_CREDENTIALS env-var pattern silently regenerates root credentials when the realm name is wrong. OAuth client-credentials flow requires bootstrapping a service principal via the management API rather than via env. STS-less S3 backends need X-Iceberg-Access-Delegation: "" as a header override. I read these as friction signals along the maturity trajectory, not refutations. They refine the per-engagement diligence list without contradicting "Polaris is maturing."
Refutation. A 12-month-flat named-production-reference count for Polaris (no new Tier-A references between 2026-06 and 2027-06), or Snowflake reorients the Polaris roadmap toward Snowflake-only integration (i.e., governance capture).
Dependent MDRs: MDR-0008 (formats + catalogs taxonomy), MDR-0009 (Archetype A weights).
A-05 · Unity Catalog self-hosted remains operationally immature
Unity Catalog self-hosted remains operationally immature relative to Databricks-hosted through the next revalidation cycle.
confidence: high · status: confirmed · basis: public-data · last reviewed: 2026-05-25 · review cadence: every 6 months
Promoted to confirmed on 2026-05-25 on two distinct surfaces of first-party evidence. Open-source Unity Catalog v2 (June 2024) is real, but the full feature set (RLS, column masking, Delta Sharing) requires the Databricks workspace.
First-party deployment evidence. Server v0.4.1 is pre-1.0; the chart has no semver release (SHA-pin required); the chart is not on any Helm registry; the default metastore is H2 in-memory; the default image tag is main (unstable). The GitHub-tagged v0.4.1 is not pushed to Docker Hub at all — v0.4.0 (2026-02-11) is the highest published image.
First-party Iceberg-REST functional incompleteness. Unity OSS v0.4.0 exposes the Iceberg REST catalog API as a read-only subset. GET /v1/{prefix}/namespaces works; POST against the same path returns NotImplementedError: Server does not support endpoint. Same for POST tables, commits, deletes. Standard tooling that expects symmetric Iceberg REST write semantics cannot treat Unity OSS as a peer write-target to Nessie or Polaris.
Refutation. Databricks publishes operational tooling for self-hosted Unity Catalog including a 24/7 support tier AND names an external case study on self-hosted deployment at production scale. Either condition alone is insufficient; both are required.
Dependent MDRs: MDR-0008 (formats + catalogs taxonomy), MDR-0009 (Archetype A weights), MDR-0018 (Archetype B weights).
A-13 · Trino federation latency floor is structural
Trino's ~2.67s floor on the Zeek conn.log workload is coordinator and multi-source planning overhead. Not tunable below 1s even on single-source workloads.
confidence: medium · status: open · basis: public-data · last reviewed: 2026-05-25 · review cadence: every 6 months
The measurement is from my own ENGINE-06-trino-federation post on Zeek conn.log. It is consistent with Trino's coordinator-plus-worker architecture imposing a planning floor regardless of source count. The honest caveat: the measurement may be conflating federation overhead with engine overhead. A single-source isolation test against an Iceberg-only Trino deployment would disambiguate. I haven't run it yet.
Refutation. An independent measurement on a single-source Iceberg-only Trino deployment with no federation shows p50 below 1s on an equivalent Zeek conn.log workload. If this refutes, Trino's lower-bound weighted-total moves from 3.23-3.48 toward 3.50, which cements Trino's #2 slot in the Engines matrix.
Dependent MDRs: MDR-0007 (engines Archetype A weights), MDR-0017 (engines Archetype B weights).
A-14 · DuckDB ~10-concurrent-analyst ceiling
DuckDB single-node ceiling for Archetype A is roughly 10 concurrent analysts before S3 read-quota saturation. Beyond that, MotherDuck or Lambda-per-query patterns are required, which breaks the embedded promise.
confidence: high · status: open · basis: public-data · last reviewed: 2026-05-25 · review cadence: every 6 months
Practitioner reports converge on the ceiling. Okta's published Lambda + DuckDB pattern explicitly works around it by using function-level isolation (1M Lambdas/day) rather than connection pooling. The single-process architecture is fundamental, so the ceiling is not tunable, only workaroundable. DuckDB belongs in the analyst-overlay tier in my scoring, not the primary-engine slot.
Refutation. A DuckDB or DuckLake change ships that enables true multi-process concurrent reads against S3-backed Iceberg without sidestepping the single-process model, or MotherDuck publishes an architecture pattern at 50+ concurrent analysts that does not require a per-analyst dedicated process. DuckLake's streaming performance gains may eventually change this story.
Dependent MDRs: MDR-0006 (engines taxonomy), MDR-0021 (engines Archetype C weights).
A-03 · Tenzir OCSF fidelity unaudited at scale
Tenzir OCSF-native normalization fidelity is architecturally genuine but unaudited at Archetype-A-scale security workloads (5+ TB/day).
confidence: medium · status: open · basis: inference · last reviewed: 2026-05-25 · review cadence: every 6 months
Tenzir's documentation describes OCSF-native normalization without separate packs, and the architecture inspection (Rust, Arrow-first, schema-aware operators) supports the claim. But the public customer count is under 100, and no named Fortune-500 reference at 5+ TB/day has published a fidelity audit. Architectural soundness is high-confidence; production-scale validation is low-confidence. The gap between those two is the central concern.
Refutation. A named Fortune-500 production reference at 5+ TB/day publishes Tenzir OCSF lossiness measurement, or an independent practitioner audit on at least three source classes (Zeek + CloudTrail + EDR) shows lossiness greater than 5%. The Asmin Aktaş WiCyS 2026 talk on resilient log-ingestion is a track-post-conference data point.
Dependent MDRs: MDR-0010 (pipelines taxonomy), MDR-0019 (pipelines Archetype B weights), MDR-0022 (pipelines Archetype C weights).
A-07 · Three scored archetypes cover engagement workload distribution
The three scored archetypes cover the practice's engagement workload distribution. Misfits are signals to add an archetype, not to bend the existing weights.
confidence: medium · status: open · basis: convention · last reviewed: 2026-05-25 · review cadence: every 6 months
The three scored archetypes (Zeek-heavy SOC, federated lakehouse, analyst-led shop) are based on early engagement experience and the workload shapes most common in practitioner conversations. A fourth candidate shape (streaming-fan-out / schema-normalization-led) has not yet been scored; if engagements arrive that don't fit the three, that is the signal to build a fourth archetype weight table, not to bend the existing ones. The assumption is that the three scored archetypes cover the distribution well enough that engagement weights are reusable. This has not yet been validated against a portfolio of completed engagements. The first three to five engagements are the test. If an engagement does not fit, the correct response is to add a new archetype with its own weight tables, not to bend the existing archetypes.
Refutation. Three or more consecutive engagements arrive with no archetype fit within the published taxonomy.
Dependent MDRs: MDR-0004 (archetype-conditional weights), MDR-0017, MDR-0018, MDR-0020, MDR-0021, MDR-0022 (all Archetype-B/C weight files).
A-08 · 4-6 week public-summary scrub is sufficient
A 4-6 week disclosure-correctness scrub between an internal scoring cut and its public release is enough time to catch disclosure issues, sensitive-information leaks, and vendor-relationship surfacing.
confidence: medium · status: open · basis: convention · last reviewed: 2026-05-25 · review cadence: every 12 months
This window was written for MDR-0014's original three-tier line, under which the scored ordering was to reach the public surface only after a delayed summary. The 2026-07-11 publication ruling dissolved that line: the full scored matrix (methodology, candidate catalog, weighted scores, reasoning, claim-vs-shipped deltas) is public, with only DeWitt/EULA-bound benchmark figures withheld, and the illustrative scored surfaces governed by MDR-0031's five conditions. What survives of MDR-0014, and what this assumption now underwrites, is the disclosure-correctness scrub and the public-correction discipline: every scoring cut still gets its scrub window between the internal cut and the public refresh that carries it.
The convention is no longer fully untested, because the 2026-07-11 publication put real scored material on the public surface and started the refutation clock below against it. If the window proves insufficient, I'll revise upward to 8-10 weeks or add a structured pre-release review step.
Refutation. A disclosure-related correction is required within 60 days of the 2026-07-11 publication, or of any subsequent public scoring refresh.
Dependent MDRs: MDR-0014 (public vs. paid line — superseded 2026-07-11; its disclosure-scrub and correction-discipline clauses survive), MDR-0031 (illustrative worked-example exception).
A-09 · 6-month default revalidation cadence
A 6-month default revalidation cadence matches the median upstream-evidence change rate across engines, catalogs, and table formats.
confidence: medium · status: open · basis: convention · last reviewed: 2026-05-25 · review cadence: every 12 months
Calibrated against observed release cadences across ClickHouse, Trino, Polaris, and Iceberg. Iceberg-related entries already sit on a 3-month cadence per A-06; the default 6-month cadence covers most criteria, with faster cadences triggered per MDR-0015's trigger logic. "Calibrated, not measured" is the honest framing.
Refutation. Two or more MDRs or scoring runs are found materially stale within 3 months of their revalidate_by date in a single quarter.
Dependent MDRs: MDR-0015 (revalidation cadence).
A-11 · ClickHouse Cloud oversells turnkey
ClickHouse Cloud's "turnkey" marketing oversells the SRE-comfort requirement for distributed-cluster scale. Mid-market deployments still want a dedicated platform engineer.
confidence: high · status: open · basis: public-data · last reviewed: 2026-05-25 · review cadence: every 6 months
This is a classic shipped-vs-claim delta. Vendor marketing says one thing; practitioner reports from Huntress engineering and my own engagement experience say another. The ClickHouse operational_complexity score of 3 captures the honest middle: it would be 4-5 on pure ClickHouse Cloud at small scale, and 2 on self-hosted at petabyte scale. Mid-market deployments at petabyte scale still want a dedicated platform engineer.
Refutation. An independent measurement at 10+ TB/day shows ClickHouse Cloud is operable by a one-engineer team without significant on-call burden, or ClickHouse Cloud ships product changes that materially reduce SRE involvement across all deployment shapes (autoscaling and auto-tuning).
Dependent MDRs: MDR-0007 (engines Archetype A weights), MDR-0021 (engines Archetype C weights).
A-12 · AWS Security Lake hidden-cost delta
AWS Security Lake's 20-40% hidden-cost delta above advertised pricing is structurally persistent for multi-cloud teams.
confidence: medium · status: open · basis: public-data · last reviewed: 2026-05-25 · review cadence: every 6 months
The hidden costs come from egress for federated queries, third-party connectors, and OCSF mapping gaps. The advertised 60-80% savings vs. DIY collapse to 40-60% after those costs land. The headline ratio depends on workload shape (single-cloud vs. multi-cloud, source-class diversity), so I hold confidence at medium. Egress is the central piece; AWS's pricing on cross-cloud egress is the rate-limiting factor.
Refutation. AWS reduces cross-cloud egress costs for federated catalog access by 50%+, or AWS Glue adds native Polaris/REST-catalog federation that eliminates the multi-cloud egress penalty.
Dependent MDRs: MDR-0009 (formats + catalogs Archetype A weights), MDR-0020 (Archetype C weights).
A-15 · Vendor benchmarks are Tier C until corroborated
Vendor self-published benchmarks stay Tier C until independently corroborated by practitioner measurement or named production reference.
confidence: high · status: open · basis: convention · last reviewed: 2026-05-25 · review cadence: every 24 months
This is a methodology meta-assumption embedded in the evidence-tier system itself, not a state-of-the-world claim. StarRocks's 5-30s multi-table benchmark and Cribl's 70-90% reduction are examples; both stay Tier C until I see practitioner corroboration or a named production reference. The long 24-month review cadence reflects that reviewing more often would just be re-affirming the convention. The convention is what underwrites the credibility of the tier system as a whole.
Refutation. Not applicable in the usual sense. Refuting this would require redesigning the evidence-tier system itself per MDR-0003.
Dependent MDRs: MDR-0003 (evidence tiers A-D).
A-16 · Detection-survivability corpus sample is representative
The sampled detection corpus a survivability score is computed over represents the client's live detection coverage closely enough that the measured fraction transfers to the unsampled remainder.
confidence: medium · status: open · basis: inference · last reviewed: 2026-06-17 · review cadence: every 6 months
A precise detection-survivability percentage (MDR-0029) is the coverage of the client's sampleddetection corpus, not of every detection they will ever run, so the headline number generalizes only as far as the sample is representative: the rule shapes that silently degrade (correlation windows, ordered sequences, distinct-counts) have to appear in the sample at roughly the rate they appear in the full corpus. The mechanism underneath the score is corpus-independent, because a backend that drops a window primitive drops it for every rule that uses one (the SIGMA-EXEC evidence), so the per-rule verdict is robust; what the sample can skew is the fraction. That is why the correlation-subset score is always reported alongside the whole-corpus number, so a single-event-heavy sample cannot hide the dangerous tail.
Refutation. A client engagement where scoring a detection sample yields a materially different survivability fraction than scoring the full corpus (the sample over- or under-represents correlation and sequence rules), which would mean survivability must be scored on the full corpus rather than a sample. Separately, if real-world corpora turn out to be almost entirely single-event rules, the whole-corpus headline would sit near 100% everywhere and lose discriminating power, leaving the correlation-subset score as the criterion's real output.
Dependent MDRs: MDR-0029 (detection-survivability criterion), MDR-0034 (C5 band re-score), MDR-0035 (C5 band plugin contingency).
A-17 · Compliance gate is binary and the regime is known at scoping
A mandated compliance control is binary (a candidate either clears the gate or is disqualified), and the client's applicable regime is knowable at engagement scoping, so the gate can be set.
confidence: medium · status: open · basis: inference · last reviewed: 2026-06-17 · review cadence: every 6 months
The gate-capability in MDR-0030 rests on two linked premises. First, that under a named regime (a 17a-4(f) WORM mandate, a residency mandate) a candidate either satisfies the mandated control or it does not, with no partial-credit middle a strong cost score could trade against, which is why the right behavior is disqualification before weighted totals rather than a low score inside them. Second, that the client's applicable regime is knowable at engagement scoping, so the gate can actually be set. The named regimes (SEC 17a-4(f), Reg SCI, DORA) are Tier-A public obligations, but whether a specific requirement binds a specific client as a hard gate, versus being satisfiable by a compensating control or a scoped exception, is a legal judgment made per client, not a property the matrix measures. WORM under 17a-4(f) is close to genuinely binary for a broker-dealer, which is the canonical case the gate is built around; residency and retrieval-SLA are softer, often satisfiable by more than one control, so treating them as hard gates can over-disqualify.
Refutation. A real engagement where a candidate that fails a sub-criterion is nonetheless acceptable to the client's auditor via a compensating control (a non-WORM store fronted by an external immutability ledger, say), which would make that sub-criterion a high-weighted score rather than a hard gate; or a client whose applicable regime cannot be pinned at scoping (multi-jurisdiction, unsettled exposure), which would mean the gate cannot be set until compliance counsel rules.
Dependent MDRs: MDR-0030 (compliance/WORM gate axis).
Audit the premises. The history is the audit trail.
Each assumption above carries the same shape: a claim, a basis, a confidence level, a status, a refutation criterion, and the MDRs that depend on it. If a premise flips, the MDRs flag for re-examination on the next revalidation cycle, not later.