A fair broker for security data engineering. Methodology and code in the open.
Measured, not asserted.
Context resolved cleanly.
Detection + hunting at scale.
The problem
$1 of sensor → ~$5 of SIEM
Observed in production: $1 of sensor → ~$5 of SIEM just to ingest it. That tax reads as a security blocker, but it’s data-engineering debt the data world mostly solved.
Schema-on-read SIEM query time degrades ~8× for 10× more data — and the per-GB licensing line scales worse than that.
Analytics tiers running $70K/mo and climbing — before the next data source is even onboarded.
The same closed architecture that taxes ingest can’t feed the AI initiatives the board is now asking for.
5:1 ratio observed in one enterprise network-telemetry deployment — directional, single-source (Tier B). The reproducible benchmark later in this deck carries the proof weight.
Why now
Two open standards from the data-engineering world — Apache Iceberg for data at rest, Apache Arrow for data in motion — have remade how Snowflake, Databricks, BigQuery, ClickHouse, DuckDB, Polars, and Trino store and process information. Vendor-locked monoliths are the old model.
2023
Iceberg vs. Delta vs. Hudi. Pick one, bet your platform on it.
2024
The Iceberg creators join the Delta vendor. The wars effectively end.
2025
Convergence shipping in production. Format choice loses meaning.
2026
AWS S3 Tables, Snowflake Polaris GA, BigQuery managed Iceberg. The data layer is open by default.
Why this persists
Three structural conflicts — which is why the evidence has to be reproducible, not asserted.
Goals for security data
Whether the framing is AI readiness or fundamentals done well, every security organization should already be doing all three — and your analysts and detection engineers already do, by hand, every time they pull data from the SIEM. Each goal is earned empirically: source by source, claim by claim, query by query.
Why all three rooms matter
Services · three principles
Layer 1 · upstream
Uptime, drop rate, time-sync drift, capture volume.
Layer 2 · in motion
SRE golden signals per pipeline stage.
Layer 3 · at rest
DAMA dimensions plus retention, per source.
Principle 1
Failures surface before queries.
Input
Each source measured through Offering 1.
Combined context
Certification
Owners confirm what they own. Coverage gaps actively measured, not hidden.
Principle 2
Coverage demonstrated, not asserted.
Wedge · 2-3 weeks
Sized migration path anchored on the two-regime engine benchmark (open lakehouse ~10× over a schema-on-read SIEM foil on hunting-shaped queries; ClickHouse native 46.8×).
Foundation · 2-4 weeks
Vendor-neutral review against the matrix. Storage, engine, 3-year TCO.
Principle 3
Validated on your workload, not the brochure.
Trustworthy data, measured source by source.
Every index and sourcetype scored across three layers before analysts ever depend on it — source health, flow health (SRE golden signals), and data quality (DAMA dimensions plus retention), rolled to one composite per source. This is the deliverable, not a dashboard screenshot.
Source health uptime · drop% · time-sync · volume · parse% • Flow health (SRE golden signals) latency · traffic · errors · saturation • Data quality (DAMA) completeness · uniqueness · timeliness · validity · accuracy · consistency · retention • Score composite, per index / sourcetype.
From tool-siloed inventories to a reconciled attack surface.
Trustworthy sources feed composite entities — asset, user, configuration, vulnerability — mapped across CMDB, EDR, vulnerability, and identity. Entity owners certify what they own; coverage gaps are measured, not papered over.
Assets, identities, configurations, vulnerabilities resolved in navigable relationship graphs (depicted here in Splunk).
Asset, user, configuration, and vulnerability records reconciled across CMDB, EDR, vulnerability scanning, and identity — one entity per real-world thing, not four tool-local copies.
Every composite entity has an accountable owner who attests to what they own each quarter — the rhythm that turns connected data into demonstrated coverage.
Unowned assets and unmapped relationships are quantified and tracked over time — coverage becomes a number you can move, not an assertion.
Modular, open, swappable at every layer.
Public production teardowns · at scale
The discussed architectures are not theoretical.
Internet & cloud scale
Quadrillion-row scale. 1.61 Q events queried in < 2 s. DNS analytics, bot management, security logging.
ClickHouse
10+ PB security data fabric. Hot retention > 1 year. 50 K IOCs swept across 10 PB in < 30 min.
Snowflake
Zero Trust FGAC at the Trino query layer. Credential Vending Service issues per-user temporary STS tokens for security analysts querying massive S3 telemetry.
Trino / Presto
Regulated industries
Banking SOC modernization. Federated legacy security telemetry across disparate sources without centralizing.
Trino (via Starburst)
Joint Cyber Lakehouse for multi-PB security analytics; sub-second queries. HIPAA, SOC 2, and ISO 27001:2022 readiness for healthcare and financial deployments.
Dremio
Norway's largest financial group. Cyber Defense Center moved to a composable Ibis-fronted stack — analysts switch backends (DuckDB, Spark, Snowflake, Splunk) per investigation.
Ibis · marimo
Security-specific deployments
Cortex XSIAM real-time security monitoring via stream processing. Mitigates threats with minimal delay at extreme event volumes.
RisingWave
Security data platform built natively on ClickHouse for HTTP analytics and massive log aggregations.
ClickHouse
Cyber Lakehouse-as-a-Service. Vendor claims ~90% security cost reduction and ~90% onboarding-time reduction (Alpha Level partnership).
Databricks
Sources: Cloudflare, Snowflake, and Databricks customer case studies; ClickHouse and DuckDB engineering blogs; published case material per company.
clickhouse.com/blog/cloudflare · snowflake.com/customers/comcast · medium.com/pinterest-engineering/securely-scaling-big-data-access-controls · starburst.io/resources/bank-hapoalim-case-study · dremio.com/blog/why-a-cyber-lakehouse-dremio-vast-data · marimo.io/blog/case-study-dnb · paloaltonetworks.com/cortex/cortex-xsiam · clickhouse.com/blog/runreveal · databricks.com/blog/transforming-cybersecurity-data-intelligence
Reference catalog summary
Today: pick two.
Need: focused, evidence-based advocacy on every one.
Product
The constraint that governs platform adoption in security is operational health: engineers who know both security and data infrastructure are scarce and expensive, and teams are rightly unwilling to bet the SOC on unproven, complex-to-operate tools, so a change has to win large on both axes, technical and operational, or the risk doesn't make sense. One product turns the fair-broker discipline into a recurring artifact built for that calculus — candidate tools scored per platform component, with the lab benchmark as structured performance evidence underneath. Public methodology; reproducible on your own workload. The scored join bench (single host, Tier B) backs the weighting: with the SOC join suite under 1.5 s on every engine, catalog maturity, concurrency behavior, and operational cost carry more of the separation than raw join latency, and the specialization that remains is real but small (StarRocks ahead on the deep multi-table joins, ClickHouse on the aggregation- and lookup-shaped SOC queries).
Candidate tools scored per platform component, weighted by client workload.
The methodology and candidate catalog are public; the engagement-internal version — weighted scoring, criterion-by-criterion reasoning, vendor-claim-vs-shipped-reality deltas, and recommended bundles per workload archetype — is the paid asset.
Annual sanity-check pass with one external practitioner reviewing for bias creep.
The benchmark anyone can re-run.
| Engine | Format | Avg query | vs. baseline |
|---|---|---|---|
| ClickHouse | Native MergeTree | 0.061 s | 46.8× faster |
| ClickHouse | Iceberg (open format, in place) | 0.282 s | 10.1× faster |
| StarRocks | Iceberg (async-MV rewrite, open format) | 0.343 s | 8.3× faster |
| Trino | Iceberg (federated REST, open format) | 0.795 s | 3.6× faster |
| Schema-on-read SIEM | Native Index (foil) | 2.854 s | 1× (baseline) |
The benchmark · design & scope
What this benchmark proves
What it doesn't prove (yet)
Architectures fail at an edge, not at a row count — and every measured edge except the cliffs gives warning. These limitations are the next iteration's scope, and the engagement starting point stands: benchmark your workload against the edges, then migrate one piece at a time. One host, Tier B — what travels is ordering and shape.
Principal
Cybersecurity architect and data scientist focused on security data architecture and emerging-tech research. Contributor to international / industry standards orgs.
Previously at Corelight (Professional Services Engineer, building international customer reference architectures), Southern Company (Senior Security Architect / Security Data Scientist), and Accenture (SIEM migrations in regulated environments).
U.S. Marine Corps veteran · Atlanta Metro
How to engage
Each engagement quotes a fixed fee scoped to deliverables — no body-shop hours, no surprise invoices. Engagements above $25K include a 6-month matrix subscription plus two quarterly reports.
2–4 weeks · $25K–$60K
Sized to scope · Custom
Wedge 2–3 wks $30K–$50K · Foundation 2–4 wks $40K–$80K
2 weeks · $20K · 100% credit toward the MOAR Migration Assessment within 90 days
Ongoing · $5K–$40K/mo
Next step — a 30-minute discovery call scopes which offering fits.
The takeaway
Open standards let you pick data engineering tools that match your work.
You need evidence that your data works for AI’s speed, scale, and sophistication.
Two regimes, measured: on the hunting-shaped aggregations the open lakehouse runs ~10× over a schema-on-read foil (ClickHouse-over-Iceberg, five-query mean; native store 46.8×), with the open-format tax about 4.6×; on simple indexed lookups the index still wins. Single host, Tier B. Detection-as-code lifts the per-rule ceiling — 1,000+ concurrent detections, maintenance debt flat.
$70K → $5K/mo — 93% at the analytics tier. 40–99% volume reduction in production.
Does the vendor ship Arrow / Iceberg / OCSF / Sigma or just market them? Portability is how you keep owning your data.
Ten analysts on a data platform they trust outperform fifty on one they privately distrust.
The same structural move ATT&CK Evals made for endpoint security — replace vendor claims with reproducible, workload-real evidence — applied to security data infrastructure. Not MITRE’s institutional scale: the integrity guard is reproducibility (public benchmark methodology, re-runnable on your own data), not a claim of being unconflicted. Partnerships disclosed.
Reference material
Appendix · Reference architectures
The SDW reference catalog, ordered by how much you can trust each pattern: public production teardowns (real, named, validated — reconstructed from the public record), one close case study of an organization running the pattern in production (Atlassian’s Project Banyan, drawn from its published record), component references (single engines), vendor blueprints (no named production validator yet — including the two pipeline-tier shapes, warehouse ELT and the SDPP category), and methodologies (detection-as-code, SQL-metadata lakehouse, MLOps for model-assisted threat hunting). Refreshes quarterly; shared through SDW products and the CISA Joint Cyber Defense Collaborative.
Public production architecture teardown
MDR/EDR business operating at fleet scale. Replaced Elasticsearch with ClickHouse Cloud on the same workload. The move was driven by economics, not vendor advocacy. Ruby-on-Rails application stack on top, Vector.dev as the routing tier, columnar OLAP underneath.
Sources
3 M endpoints
1 M identities
Route
Batched, templated; 200 K rec/sec
Store
MergeTree; columnar compression
Aggregate
Hourly + daily roll-ups
Serve
SIEM + analyst dashboards
Sources: ClickHouse case study · Huntress engineering blog · ClickHouse video "Lessons Learned Building the Huntress SIEM with ClickHouse"
clickhouse.com/blog/how-huntress-improved-performance-and-slashed-costs-with-clickhouse · huntress.com/blog/scalable-edr-advanced-agent-analytics-with-clickhouse · clickhouse.com/videos/lessons-learned-building-siem-with-clickhouse
Public production architecture teardown
Security data platform built around serverless OLAP. DuckDB runs inside AWS Lambda for normalization and operational metadata harvesting, eliminating the per-query warehouse cost that traditional ETL stacks accumulate. Mini databases per invocation, not one shared engine.
Sources
CloudTrail
VPC Flow
Ingest
Buffered event streams
Transform
SQL normalization in-function
Store
Durable result set
Serve
Detection + investigation
Sources: Data Council talk "Processing Trillions of Records at Okta with Mini Serverless Databases" · Motherduck case study · Julien Hurault, "Okta's Multi-Engine Data Stack"
datacouncil.ai/talks/processing-trillions-of-records-at-okta-with-mini-serverless-databases · motherduck.com/blog/15-companies-duckdb-in-prod · juhache.substack.com/p/oktas-multi-engine-data-stack
Case study · public reference · analyzed by SDW
Lakehouse-native security data platform on the Databricks medallion architecture, with telemetry normalized to OCSF on the Silver layer. Atlassian built and operates it; this is SDW’s reading of the architecture, drawn from the public record — Atlassian’s own Databricks customer story and its Data + AI Summit talk.
Sources
Endpoint, network, identity, cloud, app — OCSF on Silver
Ingest
Declarative pipelines + Expectations
Bronze
Source-fidelity retention
Silver
Unified schema; entity resolution
Gold
SOC surfaces; detection logic on PySpark
Atlassian built and operates Project Banyan; SDW authored this case study / analysis. Figures per Databricks’ Atlassian customer story (Tier C) · Niels Heijmans, Chief Security Architect; David Cross, CISO.
Reference catalog
A single engine or store, in isolation — what it does, where it fits the pipeline, what it composes with, and what’s brittle at scale.
Component reference
Engine-anchored architecture for organizations with mixed BI and security analytics on shared Iceberg data. Reflections pre-materialize the hot paths; the semantic layer keeps the SQL surface clean for analyst and engineer alike.
Sources
Network, endpoint, cloud, identity
Route
OCSF normalization on ingress
Store
Polaris / Nessie catalog
Engine
Semantic layer; materialized accelerations
Serve
Grafana, Superset, notebooks
Source: Dremio engineering documentation. Dremio benchmark results are withheld on public surfaces (EULA); this slide is architecture only.
docs.dremio.com · dremio.com/blog/why-a-cyber-lakehouse-dremio-vast-data
Component reference
Engine-anchored architecture for multi-region, multi-system SOCs that cannot, or should not, centralize all data into one store. Federate first, then query. Standard SQL throughout, so detection content stays portable.
Iceberg
Long-term retention · hunting
Postgres / DWs
Snowflake, BigQuery, Postgres
Engine
Single SQL surface, all sources
Kafka
Near-real-time joins
Serve
Investigation, ad-hoc hunt
Source: Trino documentation · benchmark numbers per the methodology shown on later slides.
Component reference
Engine-anchored architecture for SOCs that cannot afford batch latency. PostgreSQL-compatible streaming database; continuous SQL computation over Kafka and CDC streams; writes natively to Iceberg without a separate buffer layer.
Sources
Auth logs, API calls, network telemetry, Postgres/MySQL CDC.
Engine
Continuous SQL; incremental view maintenance.
Store
Exactly-once writes to open table format.
Serve
Sub-100ms alerts to automated containment.
Honorable mention · Feldera. Incremental view maintenance for streaming SQL; conceptually adjacent to RisingWave, with a different theoretical foundation (DBSP). Worth evaluating alongside when streaming is the binding architectural constraint.
Source: RisingWave engineering blog · arXiv streaming database research
risingwave.com/blog/building-a-real-time-cybersecurity-solution · risingwave.com/blog/privilege-escalation-detection-sql
Component reference
Engine-anchored architecture that queries telemetry directly in cheap object storage. Cribl Stream writes Parquet partitioned by date and source; Cribl Search dispatches ephemeral operators to where the data sits. No ingestion into the hot SIEM tier required for historical analysis.
Edge
Route, reshape, reduce; mask PII; normalize to OCSF.
Store
S3 / Azure Blob / Cribl Lake; partitioned by date and source.
Engine
Dispatched to data; partition pruning; no egress.
Serve
Only high-fidelity alerts forwarded to Sentinel / Splunk.
Sources: Cribl engineering · Yale New Haven Health published case study · Cribl + Azure Data Explorer joint architecture documentation
cribl.io/customers/yale-new-haven-health · cribl.io/resources/wp/eight-steps-to-mastering-your-microsoft-sentinel-migration
Reference catalog
Vendor-proposed patterns with no named production validator yet — what ships today, what doesn’t, and the honest critique.
Vendor blueprint · prerelease
Announced September 8, 2025 at .conf25; alpha confirmed February 2026; no GA date public. Splunk's response to lakehouse-native security: a schema-less, AI-ready landing zone inside Splunk Cloud / Enterprise, plus Borderless Real-Time Search federating across S3, Snowflake (GA July 2026), and (announced, unshipped) Iceberg, Delta Lake, and Azure.
Cisco Time Series Foundation Model: 250M params, Apache 2.0 open weights, decoder-only multiresolution, 16.12% MASE improvement on observability data. Trained on 300B+ data points. Runs anywhere via PyTorch. PyPI package cisco-tsm.
Machine Data Lake itself (alpha, no GA date). Federated Search for S3 re-architected (alpha). Snowflake federation (GA target July 2026). Iceberg, Delta Lake, Azure (announced, unshipped). Borderless Real-Time Search engine architecture undisclosed in public docs.
The lakehouse pivot is real but pre-shipped. Query plane still routes through Splunk Cloud / Enterprise. Splunk co-founded OCSF and the Cisco TSM is open-source — ecosystem alignment is genuine. The catalog and query-plane control remain Splunk's.
Forrester flagged the timing gap publicly: "competing platforms already deliver these offerings." No public pricing meter. No named beta customers (Singapore Airlines is a general Splunk reference, not an MDL reference). Decision today: wait vs. run the benchmark yourself.
Sources: Cisco/Splunk press release (2025-09-08) · Forrester .conf25 recap · Splunk "Complete Guide to Data Management" · arXiv 2511.19841 (Cisco Time Series Model) · GitHub splunk/cisco-time-series-model
splunk.com/blog/artificial-intelligence/introducing-the-cisco-time-series-model · arxiv.org/abs/2511.19841 · github.com/splunk/cisco-time-series-model
Vendor blueprint · prerelease
Announced March 24, 2026, Private Preview. Databricks publicly positions Lakewatch on Unity Catalog, Delta Lake, and Apache Iceberg — open table formats throughout, with OCSF on the Silver layer. Agentic triage (Mosaic AI + Anthropic Claude), Detection-as-Code, Genie NL-to-SQL.
Open Agentic SIEM on Unity Catalog. OCSF on the Silver layer. Lakeflow Declarative Pipelines + Expectations for data quality. Lakebase (serverless Postgres) for case management. DASF 2.0 (62 risks / 64 controls) as the governance overlay.
First-party asset / identity graph. Productized OCSF conformance (the DataBahn whitespace). CISO-language maturity model. Lakehouse-fluent buyer enablement curriculum. These are the partner and PS gaps where independent practitioners add value.
“Lakehouse-native security” stops being a self-build conversation. Reference architectures move from one-off engagements to a vendor-supported product motion. The TAM widens, and the buyer-education gap widens with it.
HFS Research: 80% TCO reads as “more efficient, not automatically cheaper.” Hugo Lu: “build your own SIEM” is structurally different from the GTM Splunk built its base on. InfoTech: reframe as existing-infrastructure utilization, not net-new vendor. All three are landings to plan for, not pitches to repeat.
Two acquisitions closed into launch: Antimatter (agent AuthN/AuthZ) and SiftD.ai (SPL → Lakewatch translation by the original SPL author).
Sources: Databricks Lakewatch announcement (2026-03-24) · Databricks press release
databricks.com/blog/databricks-announces-lakewatch-new-agentic-siem · databricks.com/company/newsroom/press-releases/databricks-enters-security-market-launch-lakewatch
Vendor blueprint · ELT pattern
Managed extraction (Fivetran) plus in-warehouse transformation (dbt) as the ELT spine of a security data lake: Fivetran lands cloud, identity, and SaaS logs into Snowflake / BigQuery with privacy controls; dbt normalizes raw logs to OCSF inside the warehouse, with no separate transform compute. GA products — the security-specific public evidence is real, but thinner than the general data-engineering reputation suggests.
Fivetran: managed connectors (AWS, GitHub, Jira, identity providers) into security data lakes; column blocking, hashing, and RBAC before data lands; Hybrid Deployment keeps pipelines inside the network perimeter; webhook push of Fivetran's own logs to Google Security Operations. dbt: version-controlled ELT mapping disparate logs to OCSF inside Snowflake.
Concrete and security-specific: Heritage Environmental Services (Fivetran Hybrid Deployment for regulated data); the documented Fivetran → Google Security Operations webhook integration; the published dbt → OCSF normalization pattern in Snowflake. The pattern is verifiable.
“Brex, Coinbase, Rippling use Fivetran/dbt for security” is reputational, not a named public security case. They are cited for advanced data-engineering practice generally; treat the security-specific attribution as unproven until a public case names it.
ELT-to-OCSF moves normalization out of the SIEM and into the warehouse the org already runs. The honest critique: managed ingestion is a recurring per-connector cost and a data-egress decision — weigh it against Cribl / Vector routing, and validate on your own sources before assuming Brex-tier polish.
Sources: Fivetran documentation (Hybrid Deployment; Google Security Operations webhook integration; column blocking / hashing / RBAC) · dbt + OCSF normalization public references (dbt Labs / Snowflake security data lake guides).
fivetran.com/data-movement/hybrid-deployment · fivetran.com/press/fivetran-launches-hybrid-deployment
Vendor blueprint · SDPP category
The pipeline tier has two shapes. Warehouse ELT (Fivetran + dbt, prior slide) extracts, loads, then transforms in the warehouse. SDPP route, reduce, reshape, and normalize telemetry in flight — before it lands — sitting between sources and the lake/SIEM. This is where volume economics and OCSF-on-ingress actually happen. The category is crowded and mostly vendor-claimed for security specifics; the public evidence concentrates in one or two players.
An in-flight tier between sources and the store: route, reduce (drop / sample / aggregate), reshape, mask PII, normalize to OCSF before landing. Players: Cribl, Vector (CNCF, Datadog-stewarded), Datadog Observability Pipelines, Tenzir, DataBahn, Observo AI, Monad, Abstract Security. Distinct from warehouse ELT — pre-landing, not in-warehouse transform.
Concentrated, not broad. Cribl has named public security cases (Yale New Haven Health, 30K endpoints) — it carries its own component reference in this catalog. Vector is production-validated inside the Huntress teardown on this site (the routing tier at 200K rec/sec). Those two are real and independently citable.
Datadog Observability Pipelines, Tenzir, DataBahn, Observo AI, Monad, and Abstract Security are credible and largely GA, but their security-specific production validation is vendor-claimed or analyst-relayed, not a named public security case. Treat the newer entrants as unproven for security until a public case names them — the same discipline applied to Fivetran/dbt.
The SDPP tier is where 30–70% volume reduction and OCSF-on-ingress land, ahead of SIEM/lake economics. The honest critique: a fast-consolidating category with heavy vendor claims and thin independent security validation. Choose on measured reduction against your telemetry and on pipeline-config portability — not logo count. Cribl and Vector are the proven anchors; the rest are bets.
Sources: Cribl engineering blog · Yale New Haven Health published case (see Cribl Search component reference) · Vector (CNCF / Datadog) production-validated in the Huntress teardown on this site · Datadog Observability Pipelines product documentation · Tenzir / DataBahn / Observo AI / Monad / Abstract Security vendor materials and analyst commentary — security-specific production validation vendor-claimed pending a named public case.
vector.dev · datadoghq.com/blog/observability-pipelines-stream-logs-in-ocsf-format · tenzir.com/company/press/tenzir-launches-security-data-pipeline-platform · databahn.ai/blog/security-data-pipeline-platforms · softwareanalyst.substack.com/p/the-rise-of-security-data-pipeline · forrester.com/blogs/if-youre-not-using-data-pipeline-management-dpm-for-security
Reference catalog
Cross-cutting practice that spans implementations — detection-as-code, the SQL-metadata lakehouse, MLOps for model-assisted threat hunting.
Methodology
Detection-as-code architecture pattern for SOCs whose detection backlog grows faster than the team can maintain. Versioned, tested, deployable rules; CI/CD pipelines with regression tests; continuous performance and false-positive measurement. The discipline that lets a SOC scale past the hundreds-of-detections operational ceiling.
Author
Versioned in Git; YAML, SPL, KQL, SQL per target engine.
Test
Unit tests, regression suites, replay against synthetic and historical telemetry.
Deploy
Splunk ES, Sentinel, Chronicle, custom — engine-agnostic.
Measure
Performance metrics, false-positive rate, MTTD — fed back to backlog.
Reference patterns: Anvilogic, Panther, Splunk Enterprise Security content packs; published Detection-as-Code repos including SigmaHQ.
splunk.com/blog/security/peak-threat-hunting-framework · github.com/splunk/PEAK · github.com/SigmaHQ/sigma
Methodology
Methodology for managing lakehouse metadata in a transactional SQL database rather than as JSON/Avro manifest files in object storage. Three-layer separation: data lives in Parquet on blob, metadata lives in any SQL catalog, compute reads both independently. Despite the name, DuckLake is not tied to DuckDB — it's a catalog format, not an engine choice.
Storage
S3, Azure Blob, GCS, MinIO. Same files Iceberg uses.
Catalog
Postgres, SQLite, DuckDB, MotherDuck — any SQL DB with PKs and transactions.
Compute
Any engine that reads the spec; reference implementation is the DuckDB extension.
Serve
Same Parquet files queryable from multiple compute nodes concurrently.
Sources: ducklake.select v1.0 spec (April 13, 2026); MotherDuck announcement; DuckDB Labs technical blogs; InfoQ coverage; The Register (April 16, 2026).
ducklake.select/2026/04/13/ducklake-10 · motherduck.com/blog/announcing-ducklake-1-0-on-motherduck · duckdb.org/2026/04/13/ducklake-10
Methodology
Operationalizing Model-Assisted Threat Hunting (M-ATH) from the PEAK framework. A notebook model degrades the moment telemetry shifts or an adversary adapts; MLOps is the engineering discipline — feature store, model registry, orchestrated retraining, drift monitoring — that keeps an algorithmic hunt viable in production.
Prepare
Frame the hypothesis; engineer features once in Feast — consistent offline training and online inference.
Train
Supervised classification, clustering, time-series, NLP on petabyte telemetry; experiments tracked in MLflow.
Deploy
Kubeflow / ClearML pipelines retrain and redeploy on schedule or on drift — continuous delivery, not point-in-time.
Monitor
Data- and concept-drift detection (W&B); poisoning / evasion guardrails feed back to the retrain loop.
Sources: Splunk PEAK Threat Hunting Framework (Bianco, Fetterman, Marrone — SURGe); "The Threat Hunter's Cookbook" (Fetterman & Marrone); Google Cloud MLOps maturity model (CI/CD for ML); MITRE ATLAS; CVE-2026-2635 (MLflow authentication bypass); PROID compromise-assessment framework (peer-reviewed, PMC).
splunk.com/blog/security/peak-framework-math-model-assisted-threat-hunting · cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning · atlas.mitre.org · nvd.nist.gov/vuln/detail/CVE-2026-2635 · pmc.ncbi.nlm.nih.gov/articles/PMC12595088