Capability Matrix · Cross-archetype synthesis
Five validation patterns. One non-obvious finding per archetype.
Nine scoring runs across three components (engines, formats-plus-catalogs, pipelines) and three archetypes (Zeek-heavy SOC, federated lakehouse, AWS-native serverless) produce more than nine ordered lists. They produce a structural picture of which candidates carry the most weight where, and why the same scoring taxonomy reorders cleanly when archetype weights change.
The scoring tables alone do not drive a decision. The tables tell you which candidate wins a cell. The synthesis tells you why the ranking holds, where two candidates legitimately tie, and which differences between adjacent scores are procurement-defensible versus measurement noise. That is what this page is.
Part one · Methodology validation
Five patterns the archetype-conditional-weights discipline produced.
The methodology promise (different archetypes get different weights, applied to the same criterion taxonomy and the same evidence) only earns its keep if it produces noticeably different orderings when the workload framing changes. If every archetype produced the same ranking with cosmetically different totals, the weights would be decoration. Across nine runs, five distinguishable patterns show up. Each one validates a different claim about what the weights are doing.
Pattern 1 · Dramatic re-ranking
ClickHouse #1 at A, falls to #3 at C. The weights actually distinguish.
Engines is the strongest illustration. ClickHouse sits at #1 (3.65–3.85) for Archetype A, the Zeek-heavy SOC where raw single-query latency dominates the weighting. The same ClickHouse drops to #2 at B (3.50) and #3 at C (3.65) as concurrency, federation, and managed-serverless ops climb up the weight stack. Trino holds #2 at A (3.41–3.61) on the published four-candidate table, though the v1.4 concurrency re-score (MDR-0038) opened a clear gap to ClickHouse rather than the narrow one v1.3 showed, because single-host Trino measured weakest of the arms under multi-user load; StarRocks and DuckDB follow at #3 (2.87–3.07, after the v1.4 concurrency re-score lifts StarRocks past DuckDB) and #4 (2.63–2.83). Trino takes the federated archetype outright, #1 at B (3.90) on federation breadth (a 5 at weight 20) plus Iceberg-native support (a 5 at weight 25), then holds its footing at C, landing #2 (3.98) a clear 0.31 behind Athena's #1 (4.29). Trino runs the steady inverse of ClickHouse's slide: #2 at A, #1 at B, #2 at C, the engine that gains ground as concurrency and federation weighting climb rather than losing it.
The v1.3 cut is what set those orderings. Adding a native IP type criterion lifts the engines that carry first-class IP types (ClickHouse, Trino, Athena) and trims the ones that lean on string or numeric stand-ins (DuckDB, StarRocks), which is what keeps DuckDB and StarRocks out of the top two at every archetype despite their other strengths.
Same candidates. Same criterion taxonomy. Same evidence. Different weights. The orderings reshape.
What this validates: the archetype-conditional-weights discipline is not decorating a single underlying ranking with different totals. It is producing genuinely different recommendations when the workload framing changes the design center. ClickHouse is the right call for a Zeek-heavy SOC and the wrong call for a hundred-analyst federated environment, and the matrix surfaces that rather than burying it in a total-points calculation that hides the structural reason.
Pattern 2 · Modest re-ranking with absolute lift
Delta+Unity climbs 0.65 to its home archetype. Magnitudes shift honestly.
Formats-plus-catalogs Archetype A to B is subtler than the Engines story. Delta+Unity scores 3.30 at A (multi-engine open) and 3.95 at B (single-vendor managed), a +0.65 absolute lift to its home archetype. Polaris does not collapse on the same move; it scores 4.25 at A and 4.10 at B, edging Delta+Unity at both, narrowly at B.
The ranking does not flip dramatically (Polaris is competitive at both), but the magnitudes do shift, and they shift in the direction the design center predicts. Delta+Unity earns its absolute total at Archetype B because the criteria that anchor B (RBAC concentration, managed-catalog ops, Databricks ecosystem depth) carry more weight under B's framing than under A's.
What this validates: the weights produce real absolute movements, not just relative re-orderings. A candidate's score should rise when the archetype matches its design center even if the relative ordering shifts by only one position or not at all. That is the discipline behaving correctly. The alternative, a single ranking with cosmetic re-shuffling, would not produce this pattern.
Pattern 3 · Position-preserving with total preserving
Tenzir holds #1 at both Pipelines A and B. The gap widens.
The subtlest validation. Tenzir wins Pipelines Archetype A (cost-reduction-led, 3.75–3.95) and also wins Archetype B (OCSF schema-normalization-led, 3.70–3.90). Same #1. Roughly the same absolute total. The candidates immediately below (Cribl, Vector) also hold their positions across the A-to-B move.
So nothing changes? Look at the gap. Tenzir's lead over Cribl widens under B's framing because OCSF normalization fidelity weighs more heavily under B, and Tenzir's architectural OCSF-native posture cashes in. The ranking is identical; the confidence in the recommendation is not.
What this validates: even when archetype weights do not change the winner, they should change the separation. A flat-line ranking with identical gaps across archetypes would be the symptom of weights that do not matter. Tenzir's widening lead is the opposite signal. The discipline is doing real work on the magnitudes even when it does not produce a flip at the top of the cell.
Pattern 4 · Monotonic home-archetype confirmation
Every formats-plus-catalogs candidate has a clear home. AWS Glue climbs 1.50 to its.
Iceberg+AWS Glue is the cleanest illustration. Score 2.70 at Archetype A (multi-engine open), 3.45 at B (single-vendor managed), 4.20 at C (AWS-native lake). Monotonic climb across three archetypes; +1.50 from A to C, the largest single-candidate archetype shift in v1.x. Under A's framing AWS Glue is a compromise: closed-managed-by-cloud-vendor when the design center prizes open and multi-engine. Under C's framing, AWS lock-in is accepted upfront, AWS-managed-serverless is the design center, AWS-ecosystem depth is the relevant maturity measure, and Glue is the strongest fit.
Polaris runs the inverse: 4.25 at A, 4.10 at B, 3.90 at C, a monotonic decline from its home at A. Delta+Unity peaks at B (3.95), falls at C (3.60) where Databricks is not AWS-native. At formats-plus-catalogs Archetype B specifically, Polaris narrowly leads Delta+Unity, 4.10 to 3.95.
What this validates: across three archetypes, every scored candidate has a home archetype where it scores highest, even though only some of them win that home outright. Among the formats-plus-catalogs pairs, Polaris peaks at A and wins it, Glue peaks at C and wins it, while Delta+Unity peaks at B and Nessie peaks at B without either taking the top slot from Polaris, so a peak is not always a win. That the peaks land in different archetypes is not luck; it is the methodology working as designed, because distinct design centers should pull distinct candidates toward their best fit, and a candidate that never peaks anywhere is one the matrix probably should not be carrying.
Pattern 5 · Category-winner change across archetypes
Tenzir wins A and B. Vector wins C. The recommendation itself changes.
The most consequential pattern, because it changes the actual tool recommendation rather than the magnitudes around it. Pipelines Archetype A and B both rank Tenzir #1. Pipelines Archetype C ranks Vector #1 (3.55) and Tenzir #2 (3.40), with Kinesis Firehose+Lambda entering as the fifth candidate at #3 (3.35), promoted from v1.1 implicit reference to explicit scoring under C's framing.
The mechanism: under Archetype C, pricing weight (20) and lock-in weight (15) combine to 35 points of criterion influence. Vector's OSS economics and minimum-lock-in posture cash in heavily on those two criteria, and the gain is enough to overtake Tenzir's OCSF-fidelity lead, which weighed heavily under B and modestly under A but is reweighted down under C, where AWS-native lake archetypes typically defer OCSF normalization to the storage layer or downstream engine.
What this validates: archetype-conditional weights can change which tool the matrix recommends, not just how confidently it recommends the same tool. That is the strongest possible demonstration that the weights are not cosmetic. If the right answer for a workload genuinely changes when the workload framing changes, the matrix surfaces the new answer rather than hiding it under marginal-total adjustments. Pipelines C is the first cell across v1.x where this pattern shows up cleanly. I expect more as the archetype set grows.
Part two · Per-archetype non-obvious findings
One finding per archetype that the cell-by-cell scoring does not surface.
The scoring tables present nine rankings as if each cell were independent. The synthesis layer is where I read across cells — what is the structural finding for Archetype A as a whole? What does the full row of B columns say? The three findings below are the calls I would not make from a single scoring table.
Archetype A · Zeek-heavy SOC
The component winners are not paired-recommendable together.
Archetype A's component winners read cleanly on the table: ClickHouse for Engines (3.65–3.85), Iceberg+Polaris for Formats-plus-Catalogs (4.25), Tenzir for Pipelines (3.75–3.95). All three are Tier-A or top-tier home wins under A's framing.
But I would not hand a customer that triple as a paired recommendation without surfacing the asymmetry. Cribl scores #2 at Pipelines A (3.35–3.55) on a substantially stronger production-evidence basis: Tier A named references at scale across multiple regulated industries, versus Tenzir's Tier B/C reference posture. For an Archetype-A customer where procurement defensibility carries organizational weight (regulated industry, board-level scrutiny, mandatory third-party reference checks), Cribl is the procurement-defensible default even though Tenzir scores higher on technical merit at the same cell.
The non-obvious finding: the cell winner and the engagement default can legitimately differ when the evidence tier under the score is asymmetric. The matrix surfaces both, the technical leader and the production-evidence leader, and the engagement layer reconciles them against the customer's tolerance for non-Tier-A references. A customer who can sustain Tier-B reference deployment risk gets Tenzir. A customer who cannot gets Cribl. The synthesis is what makes that distinction visible.
Archetype B · Federated lakehouse, single-vendor managed
Iceberg+Polaris narrowly leads Delta+Unity at Archetype B, 4.10 to 3.95.
Iceberg+Polaris scores 4.10 at Archetype B and Delta+Unity 3.95, a 0.15 gap that makes this a near-call rather than a clean separation. The two land so close because the methodology is working as intended: two distinct architectural commitments trade off the same criteria in opposite directions, so under B's framing they end up a hair apart rather than far apart.
Polaris carries broad multi-engine portability, OSS governance, cloud-agnosticism, and federation breadth, strengths that hold up under B's framing even where B downweights the multi-engine criterion relative to A. Delta+Unity carries deeper RBAC concentration, Databricks ecosystem depth, and Unity's attribute-based access control, strengths that lift under B's framing where governance depth outweighs portability breadth. The two arrive within 0.15 of each other by opposite routes, so they are not interchangeable, and the narrow lead is the kind the client's own RBAC weight can flip: raise that weight and Delta+Unity closes the gap or overtakes.
The honest read at B is two valid finalists a hair apart, Polaris narrowly ahead, with the choice made procurement-defensible by the scoring and the client's own governance weighting.
Archetype C · Analyst-led shop, AWS-native serverless
No dominant winner at Pipelines C. The right move is per-workload selection.
Pipelines C carries five candidates with a tight spread. Vector 3.55, Tenzir 3.40, Firehose+Lambda 3.35, Cribl 2.95, Kafka Connect 2.75. Even after Vector takes #1 and the matrix records the category-winner change (Pattern 5), the lead is 0.15 over Tenzir and 0.20 over Firehose+Lambda. Inside the noise floor I would expect from a Tier-B scoring run with this much heterogeneity in evidence sources.
The naive reading: pick Vector, ship it. The synthesis reading: at this spread, with this candidate set, the right move is per-workload selection rather than a catalog default. An AWS-native customer shipping high-volume telemetry where pricing is the binding constraint gets Vector. The same customer with a smaller telemetry footprint and tight integration into existing Lambda-based event handlers gets Firehose+Lambda. The same customer with an OCSF-first detection pipeline where schema fidelity is contractual gets Tenzir even at #2.
The non-obvious finding: a 0.15-spread top is a signal to widen the engagement question set, not to pick the cell winner. The matrix surfaces this honestly by carrying five candidates rather than truncating to a top-three; the synthesis is what tells the engagement layer to treat Pipelines C as a per-workload decision rather than a default-vendor decision. The candidate set is large enough and the gap is small enough that a customer's workload mix should drive selection, not the cell ordering.
Part three · v1.3 evidence integration
In the v1.3 evidence cut the scores held. The evidence under each firmed up.
Between the v1.2 scoring and the v1.3 evidence-strengthening pass (landed 2026-05-25), three first-party measurements moved the basis under three foundational assumptions without moving any score. That evidence-not-scoring pattern is what this section reads; the one place a later benchmark did move a score is the engines concurrency re-score in the v1.4 cut (MDR-0037), covered in the meta-finding below.
A-05 confirmed — Unity Catalog OSS immaturity now backed by two surfaces.
The Unity Catalog self-hosted immaturity assumption flipped from open to confirmed on the strength of two independent first-party signals. First, deployment maturity — the OSS Helm chart at v0.4.0 ships pre-semver with default H2 persistence and an image-tag discrepancy that forced a downgrade from the chart-recommended 0.4.1 to the actually-published 0.4.0. Second, functional gap — the Iceberg REST API surface at v0.4.0 is a read-only subset. Standard pyiceberg writes (POST/DELETE on namespaces and tables) return NotImplementedError from the server. Reads work; writes only flow through the native Unity client plus a Delta UniForm round-trip.
The synthesis implication: Delta+Unity's score at Archetype B (3.95) is built on the managed Unity Catalog product, not the OSS server. The disclosure block on the Delta+Unity candidate now carries first-party evidence of the OSS gap — anyone reading the recommendation knows the managed product is the scored entity, and the OSS path is two layers from production-ready (deployment plus functional completeness) rather than one. Sharpens the procurement-defensible posture; does not move the score.
MDR-0025 — Nessie and Polaris metadata-fetch parity at toy scale.
Status Proposed, anchored on a first-party Tier-B measurement: nine queries times three runs times two catalogs, with byte-identical data replicated across both.
| Catalog | Median | p95 |
|---|---|---|
Nessie | 98.9 ms | 241.1 ms |
Polaris | 96.4 ms | 236.0 ms |
All nine queries within 5% on median. At toy 1K-rows-per-table scale, the two catalogs are statistically indistinguishable on metadata-fetch latency.
The synthesis implication is sharper than the scoring implication. Polaris scores 4.25 at Archetype A versus Nessie's 3.10, a 1.15 spread. With MDR-0025 in the basis, that 1.15 is now explicitly characterized as RBAC posture plus ecosystem maturity, not catalog-driver overhead on the read path. The two catalogs are interchangeable on metadata fetch at small scale; Polaris's lead is structural (governance depth, ecosystem trajectory) rather than mechanical (driver latency). Customers reading the Archetype A Polaris-versus-Nessie call now see the structurally important dimension explicitly rather than inferring it from a total. Tier-A promotion is gated on the Q3 2026 catalog benchmark producing a greater-than-1M-row counterpart measurement.
A-06 first-party — Iceberg V3 row-lineage works on Spark+Hadoop, blocks on pyiceberg+Nessie.
Iceberg V3 row-lineage is the strongest counter-evidence on the null-hypothesis side of H-SEC-CATALOG-01 — the catalog-versus-audit-trail question for security data lakes. A focused first-party smoke against the OSS Python plus Nessie stack hit two stacked blockers on the write path (pyiceberg 0.11.1 cannot write V3, issue #1551 still open, and Nessie 0.107.5 silently downgrades V3 metadata to V2), plus two independent gaps beside it: DuckDB's iceberg_scan() does not surface _row_id or _last_updated_sequence_number on the reader side, and pyiceberg's overwrite and upsert paths collapse snapshot history. Closing any one does not close the others.
The synthesis implication for Engines Archetype A: Spark-based pipelines and DuckDB-based pipelines have substantially different V3 reachability today. Spark plus Hadoop-catalog plus Iceberg ≥1.11 exercises V3 row-lineage end-to-end; the equivalent pyiceberg-plus-Nessie path does not. A future v1.4 weight refresh may differentially weight Spark-based pipelines higher under criteria that depend on V3 features (audit-trail completeness, row-level lineage for incident reconstruction). For v1.3 the basis updates and the score holds.
The honest meta-finding
Accumulated evidence shifts confidence in the foundational premises. Scores wait for benchmarks.
The matrix scoring itself does not change in the v1.3 cut. That is not a defect. That is exactly how the methodology is designed to work. Accumulated evidence shifts confidence in the foundational assumptions (the explicit premises A-05, A-04, A-06 that the per-archetype weights and the per- candidate scores depend on) without arbitrarily rescoring the candidates every quarter.
Scoring changes wait for the Q3 2026 catalog benchmark and its Tier-A counterpart measurements. Until then, the v1.2 scores hold and the v1.3 basis firms up around them. A customer reading the synthesis between v1.2 scoring and the Q3 benchmark sees the same scores their procurement reviewed in May, backed by a thicker evidence basis than the basis those scores landed on. That is the asset improving between major revisions without the score-churning that destroys reference value.
The alternative, rescoring every quarter on every new datum, would produce numbers that look more current and would be less useful. Reference data needs revision discipline. The matrix's revision discipline is: scores move on benchmarks, assumptions move on evidence, and the version number labels which kind of motion just landed. The v1.4 cut is that rule in action: the 2026-06-15 multi-user benchmark measured StarRocks as the best-behaved engine under concurrency, so the engines concurrency scores moved on that benchmark (2 to 4 at Archetypes A and B), while the formats and pipelines scores, which have no new benchmark under them yet, stayed put.
Intellectual honesty · What the synthesis does not do
Three things this synthesis is not.
Not a cross-engagement applicability claim. The matrix is conditional on workload classification. A customer whose workload does not fit Archetype A, B, or C cleanly should get a scoping conversation, not a synthesis lift. The three archetypes are the ones I scored; they are not the universe of security data architectures. The methodology generalizes to other archetypes when I author the weights for them; the v1.2 synthesis applies to the three I have authored.
Not a vendor recommendation for an unstated workload. Every per-archetype finding carries an explicit workload framing. "Vector for Pipelines C" is a recommendation for an AWS-native serverless customer where pricing is the binding constraint, not a blanket Vector recommendation. A customer reading the synthesis should map their workload to an archetype before applying any candidate-level call. If the mapping is uncertain, the engagement layer resolves it before the recommendation lands.
Not a conflation of the Q3 catalog benchmark with the v1.2 scoring. The v1.2 scores are Tier-B confidence. The Q3 catalog benchmark is the planned Tier-A counterpart. Customers who need Tier-A confidence on the catalog choice specifically (Archetype B catalog tie-break; Archetype C Glue-versus-Polaris) should wait for the Q3 benchmark or commission a scoped benchmark replication rather than treating the v1.2 synthesis as Tier-A evidence. The synthesis is honest about which tier of evidence supports which call; that honesty is what makes the asset usable.
The synthesis points into the rest of the matrix.
Component criteria carries the per-component scoring rubric. Decision records (MDRs) carry the meta-decisions about how the matrix itself is built. The decision framework carries the four-phase question set the engagement runs against. The vendor evidence page carries the A–D tier discipline behind every cell, which is how you tell a score backed by a named production deployment from one backed by a vendor blog. The scoring decisions, assumptions, and archetypes pages expand on the underlying layers the synthesis references.
- Component criteria — the nine components and per-component scoring criteria.
- Decision records (MDRs) — the meta-decisions behind how the matrix scores, each cited to its MDR.
- Scoring decisions — the metric decision records (MDRs) that drive per-component scoring.
- Assumptions — the foundational premises the scores depend on, with status.
- Archetypes — Zeek-heavy SOC, federated lakehouse, AWS-native serverless — definitions and weights.
- Vendor evidence — how to read the A–D evidence tiers behind every scored cell.
- Decision framework — the four-phase, fourteen-question framework the engagement runs.