Security Data Works

The Capability Matrix

A scoring matrix for security data tools.

Once you require open formats, more than one engine qualifies — and every vendor's deck says theirs wins. That decision is where independent measurement earns its keep, because the measured spread is compressed at single-node SOC scale: every engine I tested answered the SOC join suite in under 1.5 s (Tier B, single host), so the pick usually turns on catalog maturity, concurrency behavior, and operational cost rather than raw speed. The Capability Matrix is the instrument for making that call: seven components, criteria weighted to your workload, scored from lab measurements rather than datasheets.

The full scored Matrix is public and stays public, weighted totals and criteria-by-criterion reasoning and the vendor-claim-versus-shipped-reality deltas included, all of it live on the component pages and refreshed quarterly, so anyone can check the evidence without paying for it. What a Security Data Works engagement (see /engagements) adds is the client-specific application: the scores are the same for everyone, but the weights are not, so your archetype weights turn those public scores into your per-workload bundle, your migration sequencing, and the reversibility cost of each choice, scored against your own stack and workload. That is how an engagement chooses between the compliant options without taking any vendor's word for it.

The scored matrix, at a glance

Three components. Three archetypes. Nine cells.

The matrix is public: every score, weight, and evidence tier below is open to read. Each cell is one component-by-archetype intersection, showing the top candidate for that intersection and colorized by how decisively it wins, and clicking any cell opens the full scoring table with weights, evidence tiers, and the discussion that follows the numbers. What you engage me for is applying this to your environment, not access to the scores.

Archetype A Self-hosted lakehouse. Capable platform team, operational ownership.
Archetype B Managed lakehouse. Splunk modernization, hybrid stack, vendor-supported.
Archetype C Cloud-native lakehouse. AWS-anchored, serverless-leaning.

Evidence and audit state

Most cells currently sit at evidence Tier B–C. Five first-party benchmarks are now published, four of them with public code you can rerun, and the headline is the flagship Zeek run. The lab evidence behind these cells is two-regime: the schema-on-read index wins the simple indexed lookups, and the lakehouse engines win the hunting-shaped aggregations by 5–62× — a 46.8× native / 10.1× Iceberg average across the five-query suite (Tier B, single host, 10M events, CV-gated, identical answers verified). All cells are illustrative single-host Tier B; the TB-scale multi-node regime is unmeasured. The Tier-A upgrades are still the named Q3 2026–Q1 2027 bake-offs (catalog RBAC, multi-node OLAP federated-join, pipeline OCSF-lossiness), not claims already in hand. The lab's first external review is a Q4 2026 forward commitment; the reviewer is named on the lab page when it completes, not before.

The Capability Matrix method: five candidate products scored across components on a 1-poor to 5-best scale, with the component set (Lakehouse, Catalog, Engine, Route, Graph, Storage). Cells are illustrative, not real vendor scores; the full per-vendor scores are published in the matrix.
How the scoring works, shown with illustrative cells and generic product labels; the full per-vendor scores, weights, and evidence tiers are published in the matrix itself.

Read the v1 scoring writeup → — per-archetype orderings, five validation patterns, refresh triggers. 2026-05-25.

Components

These seven components are the MOAR reference architecture. The matrix is the scoring view of the same vocabulary, not a second vocabulary. The 3x3 grid above scores three of them at a glance (engines, formats + catalogs, pipelines); the architecture narrative, component by component, is on the MOAR thesis page. Component 0 (Platform Pattern) is the composed-vs-managed decision that frames the other six.

Methodology

Evidence → Matrix → recommendation

The Matrix is the destination every lab benchmark rolls up into. The lab does not publish numbers for their own sake; each reproducible result becomes a 1–5 score on a component criterion, the scores are weighted by your workload archetype, and the weight-adjusted totals produce a defensible recommended bundle. The benchmarks are the evidence; the Matrix is the decision the evidence supports.

How the Matrix works · evidence → recommendation

How lab evidence becomes a weighted Matrix recommendationSix public reproducible lab benchmarks — zeek-flagship two-regime, engine-join-specialization, concurrency-multiuser, workload-interference, cost-to-serve-retention and pipeline-ocsf-fidelity — fan into criterion scores of 1 to 5 per candidate. Those scores are multiplied by archetype weights that sum to 100 and are set by your workload, producing weight-adjusted totals and then a per-workload bundle with sequencing and reversibility. The Matrix scores are public evidence; the client-specific per-workload bundle is the paid engagement deliverable.LAB EVIDENCE(public, reproducible)MATRIX(public evidence)RECOMMENDATIONzeek-flagship (two-regime)engine-join-specializationconcurrency-multiuserworkload-interferencecost-to-serve-retentionpipeline-ocsf-fidelitycriterion scores 1–5per candidate× archetype weights(sum to 100, set byyour workload)weight-adjusted totalsper-workload bundle+ sequencing + reversibilitythe paid deliverable
Public reproducible lab benchmarks become 1–5 criterion scores, multiplied by archetype weights that sum to 100, producing the weight-adjusted total that is the recommended per-workload bundle. A cell-winner is not the recommendation; the weighted total for your archetype is.

A cell-winner is not a procurement-defensible default; the recommendation is the weighted total for your archetype, not the best score on any single criterion. The worked example below shows the whole path on one component, and shows the winner change when the archetype changes.

A worked scorecard (illustrative)

One worked example of the method, end to end, on a single component, the Query Engine, for a Zeek-heavy SOC archetype (high-volume network time-series, long retention, sub-5 s p99 on hunting aggregations, a handful of concurrent analysts). The capability scores are grounded in the public lab (Tier B, single host) where the criterion is measured, and qualitative where it is client-specific. It is illustrative: the weights are an example archetype profile, not a client's, and the per-cell vendor-claim-vs-shipped-reality delta, which is published on the component pages, is left out of this illustrative table to keep it readable.

CriterionWtCH/IceStarRkTrino
Analytical perf (hunting aggregations)20544
Cost-to-serve (compute $/effective-TB)18543
Concurrency / multi-tenant15453
Iceberg native vs connector12345
Operational simplicity12433
Routability (deterministic front end)8444
Semantic-layer / MV rewrite5342
Federation breadth4325
SPL→SQL dialect distance3344
Existing in-house skill base (client-specific)3433
Weight-adjusted total (max 500)100414392358

For this archetype the recommendation is ClickHouse-over-Iceberg (414), with StarRocks a close alternative (392). The gap is carried by the criteria where the lab actually separates the engines (analytical aggregation performance and compute cost) more than by any single headline. StarRocks wins the concurrency criterion outright (5 vs 4), which is the whole point of the next paragraph. The per-cell delta between each vendor's published claim and its shipped reality is public, and it rides every row above; where an engagement earns its keep is the step after this table, turning these public scores into your weighted bundle, your sequencing, and the reversibility cost of each choice against your own stack.

The winner changes with the archetype, the same cells under different weights, a mixing board rather than a leaderboard. Re-weight for a many-concurrent-analyst environment (concurrency 32, cost 15, perf 10, ops 11, routability 8, Iceberg-native 10, semantic 4, federation 4, dialect 3, skill 3) and the same cells re-total StarRocks 410 > ClickHouse 404: the ranking flips because the concurrency bench found StarRocks degrades most gracefully under load while the single-query latency edge that wins the Zeek-hunting archetype erodes to a throughput tie at the host CPU ceiling. No cell changed; the workload did. That is why the engagement scores against your archetype rather than publishing a single "best engine."

Honesty boundary: the lab evidence behind these cells is two-regime: the schema-on-read index wins the simple indexed lookups, and the lakehouse engines win the hunting-shaped aggregations by 5–62× — a 46.8× native / 10.1× Iceberg average across the five-query suite (Tier B, single host, 10M events, CV-gated, identical answers verified). All cells are illustrative single-host Tier B; the TB-scale multi-node regime is unmeasured.

Augment vs replace — the decision path

The component scores answer which open stack to pick if you move; the decision path answers whether to move, and how far — the question a board votes on. The incumbent and the partial moves are scored candidate paths, not a baseline: Stay (status-quo SIEM), Augment (keep the SIEM for hot/detection, offload cold retention to the lakehouse), Hybrid-tiered, and Full-replace. The breakeven where each path pays back is computed, not asserted — migration cost divided by the monthly cost-to-serve saving against staying — so a board sees the cheapest path at its own retention horizon and the month at which the call flips. Each recommendation carries a four-part board-defensibility read: the call, what being wrong costs (the reversibility kill-switch), the evidence tier behind the crossover, and the one assumption that flips it. And the recommendation is the lowest risk-adjusted cost: a path that silently breaks a meaningful share of your detections, or fails a mandatory compliance control like WORM, is ineligible to win on cost alone. The method is here; the per-client crossover and the scored gates are the engagement.

The full scored tables

The three component matrices, in full.

Public — every score, weight, and evidence tier

Interpretation — the reasoning

Read after at least one scoring table.

Cross-cutting views, the synthesis patterns that emerge across the nine cells, the thirty-one scoring decisions, and the fifteen foundational assumptions with their refutation criteria.

MethodologyThe structural layer. Audit-grade depth on how the matrix is built.

The decision framework, the seven-component criteria taxonomy, the 90-vendor evidence-tiered database, and the six tool-architecture ADRs. Most clients never need this depth on first read; it's here for the assessment-grade audit when one is requested.

Engagement evidence

Case studies.

The named teardowns below are public analyses of on-the-record deployments. The anonymized client case studies and the client-specific architecture decision records stay behind the client-materials gate.

Public referenceCitable and shareable.

Getting the applied Matrix

The scored Matrix is already public; what an engagement delivers is the applied Matrix, your own scores re-weighted to your archetype and scored against your stack and your workload, with the recommended bundle, the migration sequencing, and the reversibility cost of each choice. The four-phase decision framework and the per-component scoring criteria are public too, published on this site. Only the anonymized case study and the client-specific architecture decision records sit behind a client-materials gate, indexed alongside the public reference catalog so nothing gets lost. To get the applied Matrix for your stack, book a scoping call.

The research page carries the hypotheses these scores rest on, and the engagements page describes how a Security Data Works engagement applies the Matrix to your own stack and workload.