Security Data Works

A fair broker for security data engineering. Methodology and code in the open.

Trustworthy

Measured, not asserted.

Well-connected

Context resolved cleanly.

Performant

Detection + hunting at scale.

The problem

The SIEM tax scales faster than your data.

$1 of sensor  →  ~$5 of SIEM

Observed in production: $1 of sensor → ~$5 of SIEM just to ingest it. That tax reads as a security blocker, but it’s data-engineering debt the data world mostly solved.

Cost outpaces data

Schema-on-read SIEM query time degrades ~8× for 10× more data — and the per-GB licensing line scales worse than that.

Budget ceiling

Analytics tiers running $70K/mo and climbing — before the next data source is even onboarded.

AI can’t run on it

The same closed architecture that taxes ingest can’t feed the AI initiatives the board is now asking for.

5:1 ratio observed in one enterprise network-telemetry deployment — directional, single-source (Tier B). The reproducible benchmark later in this deck carries the proof weight.

02 / 19

Why now

The data layer is open by default now.

Two open standards from the data-engineering world — Apache Iceberg for data at rest, Apache Arrow for data in motion — have remade how Snowflake, Databricks, BigQuery, ClickHouse, DuckDB, Polars, and Trino store and process information. Vendor-locked monoliths are the old model.

2023

Format wars

Iceberg vs. Delta vs. Hudi. Pick one, bet your platform on it.

2024

Databricks acquires Tabular

The Iceberg creators join the Delta vendor. The wars effectively end.

2025

Delta UniForm reads Iceberg

Convergence shipping in production. Format choice loses meaning.

2026

Format interoperability

AWS S3 Tables, Snowflake Polaris GA, BigQuery managed Iceberg. The data layer is open by default.

Choose ecosystem, not format. Security data architects who don't track this are being sold yesterday's architecture.
03 / 19

Why this persists

Almost everyone selling you a fix has a stake in the problem.

Three structural conflicts — which is why the evidence has to be reproducible, not asserted.

SIEM vendors

  • The ingest tax is the revenue model
  • Open formats erode the lock-in — little incentive to ship them

Resellers & Big-4

  • Resell the products they’re assessing — margin rides on the recommendation
  • Bench depth often shallower than the logo implies

Nearly everyone

  • No one hands you reproducible evidence on your own workload
  • Claims, not measurements — you can’t re-run a slideware number
04 / 19

Goals for security data

Security data goals are testable.

Whether the framing is AI readiness or fundamentals done well, every security organization should already be doing all three — and your analysts and detection engineers already do, by hand, every time they pull data from the SIEM. Each goal is earned empirically: source by source, claim by claim, query by query.

Data health, measured

  • Four layers: source health, flow health, data quality, cross-tool gap
  • Completeness, freshness, schema conformance measured per source
  • Best tool sees 47.7% of true state; cross-tool merge reaches 75.6%
  • Single-store normalization loses ~2× adversary-tail recall; a fidelity store recovers it for ~1.8× storage

Attack surface, validated

  • Asset and identity inventory reconciled across CMDB, EDR, vuln scan, identity
  • Coverage gaps named, not hidden; authoritative source per attribute
  • Business owners certify what they own, on a quarterly cadence

Data architecture, built deliberately

Standards
Arrow Iceberg D3FENDsecurity ontology OCSF Sigma
  • Every layer, storage to detection, stays swappable
  • The same data serves sub-second detection and petabyte hunting
  • Vendor lock-in becomes a deliberate trade-off, not a default
05 / 19

Why all three rooms matter

Data engineering is the unlock.

Venn of three overlapping roles: Security Analyst (domain-specific analysis), Data Engineer (scalable pipelines, governance, performance optimization, schema evolution), and Data Scientist (behavioral analysis, machine learning, statistical analysis, algorithm development). The analyst–engineer seam is ingest, validate, parse, normalize. Threat Hunter spans all three. Toolset spine beneath: SIEM, Data Engineering, MLOps.
Analysts in domain depth. Scientists already deep in data engineering. Hunters need both. The shared seam is ingest–validate–parse–normalize; the toolset spine runs SIEM → Data Engineering → MLOps. Excellence in data engineering compounds every specialist’s domain expertise.
06 / 19

Services · three principles

Services: trustworthy data, well-connected insights, performant architecture.

operational tracks
Service Offering 1 · Foundation · 2-4 weeks Data Health Validation

Layer 1 · upstream

Source health

Uptime, drop rate, time-sync drift, capture volume.

Layer 2 · in motion

Flow health

SRE golden signals per pipeline stage.

Layer 3 · at rest

Data quality

DAMA dimensions plus retention, per source.

Principle 1

Trustworthy

Failures surface before queries.

trustworthy data enables ↓
Service Offering 2 · Enablement Attack Surface Reporting Enablement

Input

Trustworthy sources

Each source measured through Offering 1.

Combined context

Composite entities mapped

· Asset · User · Configuration · Vulnerability

Certification

Entity owners certify

Owners confirm what they own. Coverage gaps actively measured, not hidden.

Principle 2

Well-connected

Coverage demonstrated, not asserted.

design track
Service Offering 3 · Architecture Modular Open Architecture

Wedge · 2-3 weeks

Splunk-to-MOAR Migration Assessment

Sized migration path anchored on the two-regime engine benchmark (open lakehouse ~10× over a schema-on-read SIEM foil on hunting-shaped queries; ClickHouse native 46.8×).

Foundation · 2-4 weeks

Architecture Assessment

Vendor-neutral review against the matrix. Storage, engine, 3-year TCO.

Principle 3

Performant

Validated on your workload, not the brochure.

07 / 19

Trustworthy data, measured source by source.

Service Offering 1: Data Health Validation

Every index and sourcetype scored across three layers before analysts ever depend on it — source health, flow health (SRE golden signals), and data quality (DAMA dimensions plus retention), rolled to one composite per source. This is the deliverable, not a dashboard screenshot.

Source health
Flow health
Data quality
Score
zeek:conn
suricata:alert
wineventlog
sysmon
okta:auth
aws:cloudtrail
gcp:audit
m365:audit
proxy:web
dns:resolver

Source health uptime · drop% · time-sync · volume · parse%  •  Flow health (SRE golden signals) latency · traffic · errors · saturation  •  Data quality (DAMA) completeness · uniqueness · timeliness · validity · accuracy · consistency · retention  •  Score composite, per index / sourcetype.

08 / 19

From tool-siloed inventories to a reconciled attack surface.

Service Offering 2: Attack Surface Reporting Enablement

Trustworthy sources feed composite entities — asset, user, configuration, vulnerability — mapped across CMDB, EDR, vulnerability, and identity. Entity owners certify what they own; coverage gaps are measured, not papered over.

Force-directed graph: thousands of asset, identity, configuration, and vulnerability nodes and their relationships resolved into one connected attack-surface graph.

Assets, identities, configurations, vulnerabilities resolved in navigable relationship graphs (depicted here in Splunk).

Composite entities, resolved

Asset, user, configuration, and vulnerability records reconciled across CMDB, EDR, vulnerability scanning, and identity — one entity per real-world thing, not four tool-local copies.

Owners certify, on a cadence

Every composite entity has an accountable owner who attests to what they own each quarter — the rhythm that turns connected data into demonstrated coverage.

Gaps measured, not hidden

Unowned assets and unmapped relationships are quantified and tracked over time — coverage becomes a number you can move, not an assertion.

09 / 19

Modular, open, swappable at every layer.

Service Offering 3: Modular Open Architecture

Security data infrastructure flow: Source data through Ingest, Store, and Analysis stages to security Tasks. Built on Arrow and Iceberg for data, OCSF for schema, and Sigma for portable detection logic. Matrix reports score tools at each layer.
Source through analysis, every layer swappable. Require these open standards in what you buy and renew — Arrow and Iceberg for data, OCSF for schema, Sigma for portable detection logic at the analysis layer. Matrix-scored throughout. The same data serves sub-second detection and petabyte-scale hunting; layers change without re-platforming, including the detections.
10 / 19

Public production teardowns · at scale

Proven at scale, in regulated industries.

The discussed architectures are not theoretical.

Internet & cloud scale

Cloudflare

Quadrillion-row scale. 1.61 Q events queried in < 2 s. DNS analytics, bot management, security logging.

ClickHouse

Comcast

10+ PB security data fabric. Hot retention > 1 year. 50 K IOCs swept across 10 PB in < 30 min.

Snowflake

Pinterest

Zero Trust FGAC at the Trino query layer. Credential Vending Service issues per-user temporary STS tokens for security analysts querying massive S3 telemetry.

Trino / Presto

Regulated industries

Bank Hapoalim

Banking SOC modernization. Federated legacy security telemetry across disparate sources without centralizing.

Trino (via Starburst)

Dremio + VAST Data

Joint Cyber Lakehouse for multi-PB security analytics; sub-second queries. HIPAA, SOC 2, and ISO 27001:2022 readiness for healthcare and financial deployments.

Dremio

DNB

Norway's largest financial group. Cyber Defense Center moved to a composable Ibis-fronted stack — analysts switch backends (DuckDB, Spark, Snowflake, Splunk) per investigation.

Ibis · marimo

Security-specific deployments

Palo Alto Networks

Cortex XSIAM real-time security monitoring via stream processing. Mitigates threats with minimal delay at extreme event volumes.

RisingWave

RunReveal

Security data platform built natively on ClickHouse for HTTP analytics and massive log aggregations.

ClickHouse

Ziggiz.ai

Cyber Lakehouse-as-a-Service. Vendor claims ~90% security cost reduction and ~90% onboarding-time reduction (Alpha Level partnership).

Databricks

Sources: Cloudflare, Snowflake, and Databricks customer case studies; ClickHouse and DuckDB engineering blogs; published case material per company.

clickhouse.com/blog/cloudflare · snowflake.com/customers/comcast · medium.com/pinterest-engineering/securely-scaling-big-data-access-controls · starburst.io/resources/bank-hapoalim-case-study · dremio.com/blog/why-a-cyber-lakehouse-dremio-vast-data · marimo.io/blog/case-study-dnb · paloaltonetworks.com/cortex/cortex-xsiam · clickhouse.com/blog/runreveal · databricks.com/blog/transforming-cybersecurity-data-intelligence

11 / 19

Reference catalog summary

Today: pick two.

Need: focused, evidence-based advocacy on every one.

Performancelatencyquery response
Scaleconcurrencyretentioningest
Inter­operabilityopen standardscross toolon-prem
Cost
Manage­ability
12 / 19

Product

The Capability Matrix.

The constraint that governs platform adoption in security is operational health: engineers who know both security and data infrastructure are scarce and expensive, and teams are rightly unwilling to bet the SOC on unproven, complex-to-operate tools, so a change has to win large on both axes, technical and operational, or the risk doesn't make sense. One product turns the fair-broker discipline into a recurring artifact built for that calculus — candidate tools scored per platform component, with the lab benchmark as structured performance evidence underneath. Public methodology; reproducible on your own workload. The scored join bench (single host, Tier B) backs the weighting: with the SOC join suite under 1.5 s on every engine, catalog maturity, concurrency behavior, and operational cost carry more of the separation than raw join latency, and the specialization that remains is real but small (StarRocks ahead on the deep multi-table joins, ClickHouse on the aggregation- and lookup-shaped SOC queries).

13 / 19

Candidate tools scored per platform component, weighted by client workload.

The Capability Matrix

Product 1
Product 2
Product 3
Product 4
Product 5
1 poor
3 fit
5 best

Components scored

Lakehouse Catalog Engine Route Graph Storage

The methodology and candidate catalog are public; the engagement-internal version — weighted scoring, criterion-by-criterion reasoning, vendor-claim-vs-shipped-reality deltas, and recommended bundles per workload archetype — is the paid asset.

Refreshed quarterly.

Annual sanity-check pass with one external practitioner reviewing for bias creep.

14 / 19

The benchmark anyone can re-run.

The benchmark behind the matrix.

Engine Format Avg query vs. baseline
ClickHouse Native MergeTree 0.061 s 46.8× faster
ClickHouse Iceberg (open format, in place) 0.282 s 10.1× faster
StarRocks Iceberg (async-MV rewrite, open format) 0.343 s 8.3× faster
Trino Iceberg (federated REST, open format) 0.795 s 3.6× faster
Schema-on-read SIEM Native Index (foil) 2.854 s 1× (baseline)
Workload
10M Zeek conn.log events · five standardized queries · re-run 2026-06-10 against an OpenSearch 2.18.0 schema-on-read foil · Tier B
Averages
Five-query mean over the foil: ClickHouse native 46.8×, ClickHouse-over-Iceberg 10.1×, StarRocks-over-Iceberg 8.3×, Trino-over-Iceberg 3.6×
Hardware
Single-node Docker on WSL2 · 32 GB RAM · 7 trials, CV-gated
Methodology
Published spec; reproducible on equivalent hardware. Answers verified identical across all arms.
Two regimes
The result is not a blanket win. On the simple indexed lookups the foil holds its ground (protocol distribution 3.4×, long-duration filter 1.8× in the index's favor); on the hunting-shaped aggregations the lakehouse pulls ahead 5–62× depending on query and realization (ClickHouse native 21×/62×, over-Iceberg 5.4×/14×).
Framing
The averages are single-host on synthetic Zeek data, not a cluster-scale or production-SIEM promise. The open-format tax is real too — ClickHouse's native store beats it over byte-identical Iceberg files by about 4.6× on average — so the matrix weighs catalog maturity, concurrency behavior, and operational cost ahead of the latency race.
15 / 19

The benchmark · design & scope

Trade-offs documented up front.

What this benchmark proves

  • Iceberg-based engines beat the schema-on-read SIEM baseline 10×+ on identical hardware
  • ZSTD-22 + LowCardinality compresses 8.2× vs. raw JSON
  • Schema-on-read SIEM query time degrades ~8× for 10× more data on the same hardware
  • Five engines (DuckDB, Trino, ClickHouse, StarRocks, Dremio) return identical answers on the gated workloads — and the store, catalog, format, and router each swap answer-clean
  • At 100M events the per-workload winner differs — ClickHouse on the full scan, DuckDB on the needle and group-by — so ‘no single engine wins’ is measured on a single host, not asserted
  • The scored engine-join bench (2026-06; single host, Tier B) — five TPC-H-derived joins plus a SOC correlation suite across five arms, every arm answering every SOC-suite join in 0.069–1.411 s, with StarRocks measurably ahead on the deep multi-table joins and ClickHouse on the aggregation- and lookup-shaped SOC queries, so engine choice rides on catalog maturity, concurrency behavior, and operational cost
  • The methodology is reproducible — anyone can re-run on their own workload

What it doesn't prove (yet)

  • Multi-node cluster concurrency — single-host concurrency (1–16 clients) is measured, where the server engines scale throughput and the embedded engine's tail degrades; a distributed cluster is still untested
  • TB-scale vendor claims — a different regime this bench neither confirms nor refutes
  • Other log types — network and identity covered; cloud audit and EDR still to run
  • Write amplification on production media — modeled, not yet metered end to end
10M — index hunt-edge
hunt-shaped aggregation crosses 10 s; lookups stay ms
100M — crossover + hazards
per-workload winner flips to the servers; one engine silently undercounts
1B — single-node ceiling
shape-dependent: top-N 79–91 s, selective filters still 0.4 s
beyond — unmeasured
zero lab data by design

Architectures fail at an edge, not at a row count — and every measured edge except the cliffs gives warning. These limitations are the next iteration's scope, and the engagement starting point stands: benchmark your workload against the edges, then migrate one piece at a time. One host, Tier B — what travels is ordering and shape.

16 / 19

Principal

Jeremy Wiley

Cybersecurity architect and data scientist focused on security data architecture and emerging-tech research. Contributor to international / industry standards orgs.

Previously at Corelight (Professional Services Engineer, building international customer reference architectures), Southern Company (Senior Security Architect / Security Data Scientist), and Accenture (SIEM migrations in regulated environments).

U.S. Marine Corps veteran · Atlanta Metro

17 / 19

How to engage

Fixed price. Scoped to fit.

Each engagement quotes a fixed fee scoped to deliverables — no body-shop hours, no surprise invoices. Engagements above $25K include a 6-month matrix subscription plus two quarterly reports.

Service Offering 1Data Health Validation

2–4 weeks · $25K–$60K

Service Offering 2Attack Surface Reporting Enablement

Sized to scope · Custom

Service Offering 3Modular Open Architecture

Wedge 2–3 wks $30K–$50K · Foundation 2–4 wks $40K–$80K

Entry wedgeModernization Discovery

2 weeks · $20K · 100% credit toward the MOAR Migration Assessment within 90 days

ContinuityAdvisory Retainer

Ongoing · $5K–$40K/mo

Next step — a 30-minute discovery call scopes which offering fits.

18 / 19

The takeaway

‘The next Splunk’ is already here; it’s just not evenly distributed.

Open standards let you pick data engineering tools that match your work.

You need evidence that your data works for AI’s speed, scale, and sophistication.

Performance & scale

Two regimes, measured: on the hunting-shaped aggregations the open lakehouse runs ~10× over a schema-on-read foil (ClickHouse-over-Iceberg, five-query mean; native store 46.8×), with the open-format tax about 4.6×; on simple indexed lookups the index still wins. Single host, Tier B. Detection-as-code lifts the per-rule ceiling — 1,000+ concurrent detections, maintenance debt flat.

Cost

$70K → $5K/mo — 93% at the analytics tier. 40–99% volume reduction in production.

Standards to test

Does the vendor ship Arrow / Iceberg / OCSF / Sigma or just market them? Portability is how you keep owning your data.

Force multiplier

Ten analysts on a data platform they trust outperform fifty on one they privately distrust.

The same structural move ATT&CK Evals made for endpoint security — replace vendor claims with reproducible, workload-real evidence — applied to security data infrastructure. Not MITRE’s institutional scale: the integrity guard is reproducibility (public benchmark methodology, re-runnable on your own data), not a claim of being unconflicted. Partnerships disclosed.

Measured. Not asserted. securitydataworks.com  ·  jeremy@securitydataworks.com
19 / 19

Reference material

Appendix: reference architectures

Appendix · Reference architectures

What’s actually working in production.

The SDW reference catalog, ordered by how much you can trust each pattern: public production teardowns (real, named, validated — reconstructed from the public record), one close case study of an organization running the pattern in production (Atlassian’s Project Banyan, drawn from its published record), component references (single engines), vendor blueprints (no named production validator yet — including the two pipeline-tier shapes, warehouse ELT and the SDPP category), and methodologies (detection-as-code, SQL-metadata lakehouse, MLOps for model-assisted threat hunting). Refreshes quarterly; shared through SDW products and the CISA Joint Cyber Defense Collaborative.

What’s working now

  • Reference architectures from real customer deployments
  • Named where public, measured outcomes
  • Only what shipped; no vendor claims

Through SDW products

  • Each architecture paired with the matrix scores that justified its tool choices
  • Methodology open; results reproducible on your own workload
  • Same fair-broker discipline as the benchmark

Through CISA JCDC

  • Working patterns shared with the federal cyber-defense community
  • Fair-broker thesis applied at industry scale
  • Reference architectures as public knowledge, not vendor IP

Public production architecture teardown

Huntress on ClickHouse.

MDR/EDR business operating at fleet scale. Replaced Elasticsearch with ClickHouse Cloud on the same workload. The move was driven by economics, not vendor advocacy. Ruby-on-Rails application stack on top, Vector.dev as the routing tier, columnar OLAP underneath.

$5K/mo
Monthly bill on ClickHouse Cloud. Was $70K on Elasticsearch. Same retention envelope, more events ingested. Roughly 93% cost reduction at the analytics tier, while throughput grew to 200 K records/sec.

Sources

Endpoints & identities

3 M endpoints
1 M identities

Route

Vector.dev (HTTP)

Batched, templated; 200 K rec/sec

Store

ClickHouse Cloud

MergeTree; columnar compression

Aggregate

MV + AggregatingMergeTree

Hourly + daily roll-ups

Serve

Huntress SOC tier

SIEM + analyst dashboards

16 B events/day ingested across the fleet.
Compression: terabytes → dozens of GB on sorted tag data.
Why Vector.dev: lightweight, templatable, mature HTTP insert path.
Why ClickHouse: columnar OLAP beats inverted-index store on this workload.
What survives: SQL-based detection content portable to other engines.
What's brittle: Vector batching tuning; egress at very high scale.

Sources: ClickHouse case study · Huntress engineering blog · ClickHouse video "Lessons Learned Building the Huntress SIEM with ClickHouse"

clickhouse.com/blog/how-huntress-improved-performance-and-slashed-costs-with-clickhouse · huntress.com/blog/scalable-edr-advanced-agent-analytics-with-clickhouse · clickhouse.com/videos/lessons-learned-building-siem-with-clickhouse

Appendix · 1 / 17

Public production architecture teardown

Okta on DuckDB-in-Lambda.

Security data platform built around serverless OLAP. DuckDB runs inside AWS Lambda for normalization and operational metadata harvesting, eliminating the per-query warehouse cost that traditional ETL stacks accumulate. Mini databases per invocation, not one shared engine.

250 GB/min
Peak normalization throughput, sustained by AWS Lambda concurrency — DuckDB embedded per invocation. Daily volume swings 1.5–50 TB/day (CloudTrail + VPC Flow); 7.5 trillion records normalized over six months across 130M files at production scale.

Sources

AWS logs

CloudTrail
VPC Flow

Ingest

Kinesis / S3 raw

Buffered event streams

Transform

Lambda + DuckDB

SQL normalization in-function

Store

S3 normalized

Durable result set

Serve

Downstream engines

Detection + investigation

7.5T records / 6 months: cumulative production scale across 130M files.
Why DuckDB: embedded OLAP with full SQL; no cluster to operate.
Why serverless: auto-scales with event volume; pay-per-invocation.
What composes: normalized result feeds other engines for query-time work.
What's distinctive: mini databases per invocation, not one shared engine.
What's brittle: Lambda cold-start tail; DuckDB version pinning across deploys.

Sources: Data Council talk "Processing Trillions of Records at Okta with Mini Serverless Databases" · Motherduck case study · Julien Hurault, "Okta's Multi-Engine Data Stack"

datacouncil.ai/talks/processing-trillions-of-records-at-okta-with-mini-serverless-databases · motherduck.com/blog/15-companies-duckdb-in-prod · juhache.substack.com/p/oktas-multi-engine-data-stack

Appendix · 2 / 17

Case study · public reference · analyzed by SDW

Atlassian’s Project Banyan on Databricks + OCSF.

Lakehouse-native security data platform on the Databricks medallion architecture, with telemetry normalized to OCSF on the Silver layer. Atlassian built and operates it; this is SDW’s reading of the architecture, drawn from the public record — Atlassian’s own Databricks customer story and its Data + AI Summit talk.

21B+
Security events queryable in under a minute, with OCSF normalized once on the Silver layer. Ingest costs down roughly 80%, hot retention extended from 30 days to 12 months, high-frequency detection latency cut from 17 s to 5 s on Photon.

Sources

Full telemetry breadth

Endpoint, network, identity, cloud, app — OCSF on Silver

Ingest

Lakeflow + streaming

Declarative pipelines + Expectations

Bronze

Raw (Delta Lake)

Source-fidelity retention

Silver

OCSF-normalized

Unified schema; entity resolution

Gold

Detections + cases

SOC surfaces; detection logic on PySpark

Unity Catalog for governance, lineage, and time-travel audit across all three layers.
Published outcomes: ingest cost down ~80%; file sizes ~20% smaller once OCSF-standardized into leaner Parquet.
Why medallion: raw fidelity preserved while the analyst surface stays normalized.
Why OCSF on Silver: detection portability across sources; vendor-agnostic, normalized once and early.
What this validates: lakehouse-native security at this scale is production-proven, not theoretical.
What's hard: mapping coverage at long-tail sources; semantic drift over time — ongoing cost, not one-time setup.

Atlassian built and operates Project Banyan; SDW authored this case study / analysis. Figures per Databricks’ Atlassian customer story (Tier C) · Niels Heijmans, Chief Security Architect; David Cross, CISO.

Appendix · 3 / 17

Reference catalog

Component reference.

A single engine or store, in isolation — what it does, where it fits the pipeline, what it composes with, and what’s brittle at scale.

Appendix · 4 / 17

Component reference

Dremio — semantic layer + Reflections.

Engine-anchored architecture for organizations with mixed BI and security analytics on shared Iceberg data. Reflections pre-materialize the hot paths; the semantic layer keeps the SQL surface clean for analyst and engineer alike.

Iceberg-native semantic layer over open tables. Dremio queries open Iceberg in place with Reflections off by design here; the semantic layer keeps the SQL surface clean for analyst and engineer alike, and detections stay portable to Trino, StarRocks, and ClickHouse-Iceberg.

Sources

Security telemetry

Network, endpoint, cloud, identity

Route

Vector / Cribl / Kafka

OCSF normalization on ingress

Store

Iceberg on S3

Polaris / Nessie catalog

Engine

Dremio + Reflections

Semantic layer; materialized accelerations

Serve

BI + SOC UIs

Grafana, Superset, notebooks

Iceberg-native: queries hit the lake without copy-out.
Reflections: accelerate hot queries transparently; engineers tune, analysts benefit.
Semantic layer: reduces SQL complexity for SOC analysts.
Best fit: mixed BI + security analytics on shared data; reusable accelerations.
Trade-off: slower raw scan than ClickHouse native MergeTree (Reflections off) — the semantic layer and Reflections are where it earns its keep, not raw scan.
What survives: standard SQL detections portable to Trino, StarRocks, ClickHouse-Iceberg.

Source: Dremio engineering documentation. Dremio benchmark results are withheld on public surfaces (EULA); this slide is architecture only.

docs.dremio.com · dremio.com/blog/why-a-cyber-lakehouse-dremio-vast-data

Appendix · 5 / 17

Component reference

Trino — federation breadth.

Engine-anchored architecture for multi-region, multi-system SOCs that cannot, or should not, centralize all data into one store. Federate first, then query. Standard SQL throughout, so detection content stays portable.

3.6×
Faster than the schema-on-read SIEM foil (OpenSearch, 2.854 s) on the same Zeek workload at 0.795 s average query, with federation reach Splunk cannot match. Iceberg, Postgres, cloud DWs, and Kafka queried in one SQL plane, without copying data into a central store.

Iceberg

Lakehouse

Long-term retention · hunting

Postgres / DWs

Operational stores

Snowflake, BigQuery, Postgres

Engine

Trino federation plane

Single SQL surface, all sources

Kafka

Streams

Near-real-time joins

Serve

SOC + analytics

Investigation, ad-hoc hunt

No central store required. Data stays where it sits; query reaches across.
Why Trino: mature connector ecosystem; standard SQL; horizontal scale-out.
Best fit: multi-region, data-residency, cross-jurisdiction SOCs.
Trade-off: slower than single-engine OLAP (0.795 s vs. ClickHouse native 0.061 s).
What survives: standard SQL detections portable across engines.
What's distinctive: federation reach Splunk cannot match at any price.

Source: Trino documentation · benchmark numbers per the methodology shown on later slides.

trino.io/docs

Appendix · 6 / 17

Component reference

RisingWave — streaming threat detection.

Engine-anchored architecture for SOCs that cannot afford batch latency. PostgreSQL-compatible streaming database; continuous SQL computation over Kafka and CDC streams; writes natively to Iceberg without a separate buffer layer.

<100ms
Event-to-rule firing. Replaces the batch-index methodology that creates a five-minute vulnerability window before scheduled queries run. Detection fires the moment an event arrives.

Sources

Kafka / CDC

Auth logs, API calls, network telemetry, Postgres/MySQL CDC.

Engine

RisingWave streaming compute

Continuous SQL; incremental view maintenance.

Store

Iceberg on S3

Exactly-once writes to open table format.

Serve

SOAR + dashboards

Sub-100ms alerts to automated containment.

Production at scale: Palo Alto Networks (Cortex XSIAM), China Telecom Bestpay.
PostgreSQL-compatible: standard SQL, no proprietary dialect to learn.
Native CDC ingestion: Postgres, MySQL, MongoDB streams as first-class inputs.
Why this matters: closes the open-format streaming-ingest gap at security scale.
Materialized views ground LLMs and agentic AI on live telemetry, not stale snapshots.
Best fit: sub-minute detection that must run on lakehouse-native data.

Honorable mention · Feldera. Incremental view maintenance for streaming SQL; conceptually adjacent to RisingWave, with a different theoretical foundation (DBSP). Worth evaluating alongside when streaming is the binding architectural constraint.

Source: RisingWave engineering blog · arXiv streaming database research

risingwave.com/blog/building-a-real-time-cybersecurity-solution · risingwave.com/blog/privilege-escalation-detection-sql

Appendix · 7 / 17

Component reference

Cribl Search — query in place, never rehydrate.

Engine-anchored architecture that queries telemetry directly in cheap object storage. Cribl Stream writes Parquet partitioned by date and source; Cribl Search dispatches ephemeral operators to where the data sits. No ingestion into the hot SIEM tier required for historical analysis.

40%
Log volume reduction at Yale New Haven Health across 30,000 endpoints during their migration to Microsoft Sentinel. A separate Fortune 1000 IT Services deployment achieved 99.99% reduction in Virtual SOC traffic on the same pattern.

Edge

Cribl Edge / Stream

Route, reshape, reduce; mask PII; normalize to OCSF.

Store

Object storage (Parquet)

S3 / Azure Blob / Cribl Lake; partitioned by date and source.

Engine

Cribl Search operators

Dispatched to data; partition pruning; no egress.

Serve

SIEM + SOC

Only high-fidelity alerts forwarded to Sentinel / Splunk.

Three-plane model: Management, Control, Data planes; operators are ephemeral.
Regional / Geo Split: Search co-located with storage bucket bypasses cloud egress fees.
Route, reshape, reduce before expensive destinations; G-Cloud TCO: 35% savings at 100 GB/day, 64% at 1 TB/day.
Why this matters: historical telemetry stays queryable without SIEM hot-tier rates.
Best fit: 7-year compliance retention, forensics, ad-hoc hunts on cold data.
What's brittle: partition discipline; cardinality choices materially affect query performance.

Sources: Cribl engineering · Yale New Haven Health published case study · Cribl + Azure Data Explorer joint architecture documentation

cribl.io/customers/yale-new-haven-health · cribl.io/resources/wp/eight-steps-to-mastering-your-microsoft-sentinel-migration

Appendix · 8 / 17

Reference catalog

Vendor blueprint.

Vendor-proposed patterns with no named production validator yet — what ships today, what doesn’t, and the honest critique.

Appendix · 9 / 17

Vendor blueprint · prerelease

Splunk Machine Data Lake — the Cisco Data Fabric data layer.

Announced September 8, 2025 at .conf25; alpha confirmed February 2026; no GA date public. Splunk's response to lakehouse-native security: a schema-less, AI-ready landing zone inside Splunk Cloud / Enterprise, plus Borderless Real-Time Search federating across S3, Snowflake (GA July 2026), and (announced, unshipped) Iceberg, Delta Lake, and Azure.

1

What ships today

Cisco Time Series Foundation Model: 250M params, Apache 2.0 open weights, decoder-only multiresolution, 16.12% MASE improvement on observability data. Trained on 300B+ data points. Runs anywhere via PyTorch. PyPI package cisco-tsm.

2

What doesn’t ship yet

Machine Data Lake itself (alpha, no GA date). Federated Search for S3 re-architected (alpha). Snowflake federation (GA target July 2026). Iceberg, Delta Lake, Azure (announced, unshipped). Borderless Real-Time Search engine architecture undisclosed in public docs.

3

What it changes for architects

The lakehouse pivot is real but pre-shipped. Query plane still routes through Splunk Cloud / Enterprise. Splunk co-founded OCSF and the Cisco TSM is open-source — ecosystem alignment is genuine. The catalog and query-plane control remain Splunk's.

4

The honest critique

Forrester flagged the timing gap publicly: "competing platforms already deliver these offerings." No public pricing meter. No named beta customers (Singapore Airlines is a general Splunk reference, not an MDL reference). Decision today: wait vs. run the benchmark yourself.

Sources: Cisco/Splunk press release (2025-09-08) · Forrester .conf25 recap · Splunk "Complete Guide to Data Management" · arXiv 2511.19841 (Cisco Time Series Model) · GitHub splunk/cisco-time-series-model

splunk.com/blog/artificial-intelligence/introducing-the-cisco-time-series-model · arxiv.org/abs/2511.19841 · github.com/splunk/cisco-time-series-model

Appendix · 10 / 17

Vendor blueprint · prerelease

Databricks Lakewatch — open, agentic SIEM.

Announced March 24, 2026, Private Preview. Databricks publicly positions Lakewatch on Unity Catalog, Delta Lake, and Apache Iceberg — open table formats throughout, with OCSF on the Silver layer. Agentic triage (Mosaic AI + Anthropic Claude), Detection-as-Code, Genie NL-to-SQL.

1

What ships today

Open Agentic SIEM on Unity Catalog. OCSF on the Silver layer. Lakeflow Declarative Pipelines + Expectations for data quality. Lakebase (serverless Postgres) for case management. DASF 2.0 (62 risks / 64 controls) as the governance overlay.

2

What doesn’t ship yet

First-party asset / identity graph. Productized OCSF conformance (the DataBahn whitespace). CISO-language maturity model. Lakehouse-fluent buyer enablement curriculum. These are the partner and PS gaps where independent practitioners add value.

3

What it changes for architects

“Lakehouse-native security” stops being a self-build conversation. Reference architectures move from one-off engagements to a vendor-supported product motion. The TAM widens, and the buyer-education gap widens with it.

4

The honest critique

HFS Research: 80% TCO reads as “more efficient, not automatically cheaper.” Hugo Lu: “build your own SIEM” is structurally different from the GTM Splunk built its base on. InfoTech: reframe as existing-infrastructure utilization, not net-new vendor. All three are landings to plan for, not pitches to repeat.

Two acquisitions closed into launch: Antimatter (agent AuthN/AuthZ) and SiftD.ai (SPL → Lakewatch translation by the original SPL author).

Sources: Databricks Lakewatch announcement (2026-03-24) · Databricks press release

databricks.com/blog/databricks-announces-lakewatch-new-agentic-siem · databricks.com/company/newsroom/press-releases/databricks-enters-security-market-launch-lakewatch

Appendix · 11 / 17

Vendor blueprint · ELT pattern

Fivetran + dbt — ELT for the security data lake.

Managed extraction (Fivetran) plus in-warehouse transformation (dbt) as the ELT spine of a security data lake: Fivetran lands cloud, identity, and SaaS logs into Snowflake / BigQuery with privacy controls; dbt normalizes raw logs to OCSF inside the warehouse, with no separate transform compute. GA products — the security-specific public evidence is real, but thinner than the general data-engineering reputation suggests.

1

What ships today

Fivetran: managed connectors (AWS, GitHub, Jira, identity providers) into security data lakes; column blocking, hashing, and RBAC before data lands; Hybrid Deployment keeps pipelines inside the network perimeter; webhook push of Fivetran's own logs to Google Security Operations. dbt: version-controlled ELT mapping disparate logs to OCSF inside Snowflake.

2

Where the public evidence is

Concrete and security-specific: Heritage Environmental Services (Fivetran Hybrid Deployment for regulated data); the documented Fivetran → Google Security Operations webhook integration; the published dbt → OCSF normalization pattern in Snowflake. The pattern is verifiable.

3

Where it isn’t

“Brex, Coinbase, Rippling use Fivetran/dbt for security” is reputational, not a named public security case. They are cited for advanced data-engineering practice generally; treat the security-specific attribution as unproven until a public case names it.

4

What it changes for architects

ELT-to-OCSF moves normalization out of the SIEM and into the warehouse the org already runs. The honest critique: managed ingestion is a recurring per-connector cost and a data-egress decision — weigh it against Cribl / Vector routing, and validate on your own sources before assuming Brex-tier polish.

Sources: Fivetran documentation (Hybrid Deployment; Google Security Operations webhook integration; column blocking / hashing / RBAC) · dbt + OCSF normalization public references (dbt Labs / Snowflake security data lake guides).

fivetran.com/data-movement/hybrid-deployment · fivetran.com/press/fivetran-launches-hybrid-deployment

Appendix · 12 / 17

Vendor blueprint · SDPP category

Security data pipeline platforms — the in-flight tier.

The pipeline tier has two shapes. Warehouse ELT (Fivetran + dbt, prior slide) extracts, loads, then transforms in the warehouse. SDPP route, reduce, reshape, and normalize telemetry in flight — before it lands — sitting between sources and the lake/SIEM. This is where volume economics and OCSF-on-ingress actually happen. The category is crowded and mostly vendor-claimed for security specifics; the public evidence concentrates in one or two players.

1

What the category is

An in-flight tier between sources and the store: route, reduce (drop / sample / aggregate), reshape, mask PII, normalize to OCSF before landing. Players: Cribl, Vector (CNCF, Datadog-stewarded), Datadog Observability Pipelines, Tenzir, DataBahn, Observo AI, Monad, Abstract Security. Distinct from warehouse ELT — pre-landing, not in-warehouse transform.

2

Where the public evidence is

Concentrated, not broad. Cribl has named public security cases (Yale New Haven Health, 30K endpoints) — it carries its own component reference in this catalog. Vector is production-validated inside the Huntress teardown on this site (the routing tier at 200K rec/sec). Those two are real and independently citable.

3

Where it isn’t

Datadog Observability Pipelines, Tenzir, DataBahn, Observo AI, Monad, and Abstract Security are credible and largely GA, but their security-specific production validation is vendor-claimed or analyst-relayed, not a named public security case. Treat the newer entrants as unproven for security until a public case names them — the same discipline applied to Fivetran/dbt.

4

What it changes for architects

The SDPP tier is where 30–70% volume reduction and OCSF-on-ingress land, ahead of SIEM/lake economics. The honest critique: a fast-consolidating category with heavy vendor claims and thin independent security validation. Choose on measured reduction against your telemetry and on pipeline-config portability — not logo count. Cribl and Vector are the proven anchors; the rest are bets.

Sources: Cribl engineering blog · Yale New Haven Health published case (see Cribl Search component reference) · Vector (CNCF / Datadog) production-validated in the Huntress teardown on this site · Datadog Observability Pipelines product documentation · Tenzir / DataBahn / Observo AI / Monad / Abstract Security vendor materials and analyst commentary — security-specific production validation vendor-claimed pending a named public case.

vector.dev · datadoghq.com/blog/observability-pipelines-stream-logs-in-ocsf-format · tenzir.com/company/press/tenzir-launches-security-data-pipeline-platform · databahn.ai/blog/security-data-pipeline-platforms · softwareanalyst.substack.com/p/the-rise-of-security-data-pipeline · forrester.com/blogs/if-youre-not-using-data-pipeline-management-dpm-for-security

Appendix · 13 / 17

Reference catalog

Methodology.

Cross-cutting practice that spans implementations — detection-as-code, the SQL-metadata lakehouse, MLOps for model-assisted threat hunting.

Appendix · 14 / 17

Methodology

DetectFlow — thousands of detections without operational debt.

Detection-as-code architecture pattern for SOCs whose detection backlog grows faster than the team can maintain. Versioned, tested, deployable rules; CI/CD pipelines with regression tests; continuous performance and false-positive measurement. The discipline that lets a SOC scale past the hundreds-of-detections operational ceiling.

1000+
Detections in production with operational debt held flat. Most detection programs cap at low hundreds because each new rule adds maintenance load; DetectFlow keeps the per-rule maintenance cost near zero through automation.

Author

Detection content

Versioned in Git; YAML, SPL, KQL, SQL per target engine.

Test

CI/CD pipeline

Unit tests, regression suites, replay against synthetic and historical telemetry.

Deploy

Detection engine

Splunk ES, Sentinel, Chronicle, custom — engine-agnostic.

Measure

Feedback loop

Performance metrics, false-positive rate, MTTD — fed back to backlog.

Why this matters: the hundreds-of-detections operational ceiling is real; DetectFlow removes it.
Reference patterns: Anvilogic, Panther, hand-rolled CI/CD on Detection-as-Code repos.
Detection portability: standard SQL detections move between engines without rewrite.
Feedback loop: every rule's performance and FP rate measured per deployment cycle.
Best fit: SOCs where analyst time is consumed by tuning rather than hunting.
What's hard: the cultural shift from "detection content as personal craft" to "detection as production software."

Reference patterns: Anvilogic, Panther, Splunk Enterprise Security content packs; published Detection-as-Code repos including SigmaHQ.

splunk.com/blog/security/peak-threat-hunting-framework · github.com/splunk/PEAK · github.com/SigmaHQ/sigma

Appendix · 15 / 17

Methodology

DuckLake — SQL metadata for the lakehouse.

Methodology for managing lakehouse metadata in a transactional SQL database rather than as JSON/Avro manifest files in object storage. Three-layer separation: data lives in Parquet on blob, metadata lives in any SQL catalog, compute reads both independently. Despite the name, DuckLake is not tied to DuckDB — it's a catalog format, not an engine choice.

926×
Faster queries vs. Iceberg on DuckDB Labs benchmarks; 105× faster ingestion; in streaming workloads, 900× faster reads and 100× faster writes. Open question: independent validation outside DuckDB Labs benchmarks remains pending. The architectural claim — SQL metadata avoids manifest-file proliferation — is sound regardless of the specific multiplier.

Storage

Parquet on blob

S3, Azure Blob, GCS, MinIO. Same files Iceberg uses.

Catalog

SQL metadata database

Postgres, SQLite, DuckDB, MotherDuck — any SQL DB with PKs and transactions.

Compute

Engine-agnostic

Any engine that reads the spec; reference implementation is the DuckDB extension.

Serve

Lakehouse queries

Same Parquet files queryable from multiple compute nodes concurrently.

Production-ready: v1.0 released April 2026 with stable spec and backward-compatibility guarantee.
Replaces the full stack: not just Iceberg or Delta — replaces "Iceberg + Polaris" or "Delta + Unity" together.
Why SQL metadata wins: high-frequency commits don't thrash a Postgres index the way they thrash Iceberg manifest lists.
Native encryption: data files encryptable; keys in the catalog DB. Auth and AuthZ via the catalog (Postgres roles, etc.).
Compatibility: data files exportable to Iceberg if needed; not a one-way bet.
What's not validated yet: independent benchmarks outside DuckDB Labs; security workload validation at TB/day scale.

Sources: ducklake.select v1.0 spec (April 13, 2026); MotherDuck announcement; DuckDB Labs technical blogs; InfoQ coverage; The Register (April 16, 2026).

ducklake.select/2026/04/13/ducklake-10 · motherduck.com/blog/announcing-ducklake-1-0-on-motherduck · duckdb.org/2026/04/13/ducklake-10

Appendix · 16 / 17

Methodology

MLOps data science workflows — for modern, disciplined threat hunting.

Operationalizing Model-Assisted Threat Hunting (M-ATH) from the PEAK framework. A notebook model degrades the moment telemetry shifts or an adversary adapts; MLOps is the engineering discipline — feature store, model registry, orchestrated retraining, drift monitoring — that keeps an algorithmic hunt viable in production.

Level 2
CI/CD-automated retraining is the threshold where M-ATH stops being a one-off notebook hunt. Google Cloud's MLOps maturity model: Level 0 (manual notebook, no drift defense) → Level 2 (pipelines that detect drift and retrain autonomously). Below Level 1, the model is stale before the hunt ends.

Prepare

PEAK + feature store

Frame the hypothesis; engineer features once in Feast — consistent offline training and online inference.

Train

M-ATH model

Supervised classification, clustering, time-series, NLP on petabyte telemetry; experiments tracked in MLflow.

Deploy

Orchestrated retraining

Kubeflow / ClearML pipelines retrain and redeploy on schedule or on drift — continuous delivery, not point-in-time.

Monitor

Drift + MLSecOps

Data- and concept-drift detection (W&B); poisoning / evasion guardrails feed back to the retrain loop.

PEAK framework: Bianco, Fetterman, Marrone (Splunk SURGe). M-ATH is the algorithmic hunt type alongside hypothesis-driven and baseline.
Data vs concept drift: distribution shift vs adversary adaptation — both degrade silently into a false-positive avalanche or a false-negative blind spot.
Why Level 0 fails: a notebook model trained offline and discarded cannot counter non-stationary, adversarial telemetry.
Tooling: Feast (features), MLflow (registry / Detection-as-Code), Kubeflow & ClearML (orchestration), W&B (drift + reasoning observability).
MLSecOps: the hunt infra is itself an attack surface — retrain-loop poisoning, evasion inputs, MLflow CVE-2026-2635, the Kubeflow Doki incident. MITRE ATLAS.
What's hard: continuous-retraining cost; poisoned retraining loops; the org gap between data science and detection engineering.

Sources: Splunk PEAK Threat Hunting Framework (Bianco, Fetterman, Marrone — SURGe); "The Threat Hunter's Cookbook" (Fetterman & Marrone); Google Cloud MLOps maturity model (CI/CD for ML); MITRE ATLAS; CVE-2026-2635 (MLflow authentication bypass); PROID compromise-assessment framework (peer-reviewed, PMC).

splunk.com/blog/security/peak-framework-math-model-assisted-threat-hunting · cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning · atlas.mitre.org · nvd.nist.gov/vuln/detail/CVE-2026-2635 · pmc.ncbi.nlm.nih.gov/articles/PMC12595088

Appendix · 17 / 17