Security Data Works

Writing · Provenance

How each essay was made.

A practice that publishes benchmarks nobody else will run cannot be vague about which of its own essays a model drafted. Every essay on this site carries one of three badges saying who held the pen and who checked the work, per essay rather than as one notice at the bottom of the site, because a sitewide disclaimer tells you nothing about the piece actually in front of you.

The badge describes process, not quality. A state 1 essay is not a worse essay; it is one whose claims have not been checked, so the useful way to read the badge is as a discount rate on the evidence rather than a rating of the writing.

  • 01

    AI-drafted, human observed
    Asserts
    A model produced the draft. A human read it end to end and let it stand.
    Does not
    That any claim, number, or citation was checked.
    Essays
    79 of 98
  • 02

    Human revised
    Asserts
    A human rewrote the argument, the structure and the language. The reasoning is the author's.
    Does not
    That sources were verified or numbers reconciled against the lab.
    Essays
    13 of 98
  • 03

    Fully human vetted
    Asserts
    Every claim traces to a named source; every number was checked against the lab result that produced it.
    Does not
    That the essay is correct — only that it has been checked and the author stands behind it.
    Essays
    6 of 98

The ladder only goes one way.

State 3 includes everything state 2 asserts and state 2 everything state 1 asserts, so an essay only moves up once the work behind the higher state has actually been done. An essay whose provenance nobody can reconstruct is a state 1, which is why the whole catalog started at the bottom and each promotion has to be earned one at a time. A badge applied in advance would cost more trust than it bought.

Who assigns these, and how far to trust that.

I assign them myself, from memory of how each piece was written, and that is the weakest link in the whole scheme so it is worth saying plainly rather than leaving you to infer it. The repository history cannot settle it either: every essay here was revised after it was first committed, 96 of the 98 across two or more separate days, and 533 of this repository’s 600 commits ran through an AI session, so revision volume cannot separate an essay I rewrote from one I had a model rewrite. That distinction is exactly the line between state 1 and state 2, and it rests on my word.

State 3 is the exception, and it is the reason the ladder has a third rung at all: it means the numbers trace to a named result file in the public lab, so you can open that file and disagree with me. Six essays are there now, each ending in a Source data block naming the exact file its figures were read back out of, down to the seed the corpus was generated under where the run records one. Twelve of the nineteen carrying first-party measured results now name their benchmark; the remaining seven do not, and that is the gap the third rung is waiting on rather than any doubt about the numbers themselves. Some will never close it, because their figures are not lab measurements to begin with: the D3FEND and ATT&CK coverage percentages come from the mapping work rather than a benchmark run, and pointing them at a lab directory would be citing the wrong source instead of filling in a missing one.

Checking has now been wrong twice on the same essay, in opposite directions, and both are worth recording. An earlier pass reported that the write-pattern essay quoted an inlining win of 3.93× with no run under it, and held the essay at state 2 on that basis. The number was sourced the whole time. I had searched the benchmark whose name matched the claim, which counts data files and times nothing, instead of the one that ran the measurement: the ratio is a row in the streaming-cadence sweep, 200 commits at five rows each. The correction that followed then got the metric wrong. It called 3.93× a commit-latency ratio, when it is a sustained-throughput ratio, 273 rows per second against 70; the commit-latency gap in that same run is wider still, 133 ms against 11.9. So the essay was promoted to state 3 on the strength of a citation that misnamed what it cited. It now names all three benchmarks it rests on and says which metric each number belongs to, and it has gone back to state 2 until I have read it through against the lab myself rather than on the strength of that repair. The first failure was searching by name instead of by measurement. The second was citing the right file and describing the wrong column, which is the harder one to catch, because a citation that names a real file and a real number looks like a check that happened.

Three panels. In the first, a machine is at the keyboard writing while a person watches the screen. In the second, the person is at the keyboard rewriting and the machine is set aside. In the third, the person is at the keyboard and the finished document carries a seal.

Why disclose it at all.

A sitewide footer saying AI was used somewhere is cheaper to maintain and carries almost no information. Doing it per essay lets you weight a measured lab result differently from an argument sketched out at speed, which is a distinction the writing depends on anyway.

The evidence tiers inside the essays do a related job for the claims themselves, Tier A through D on how a number was produced. The provenance badge is about how the prose was produced.

← All pillars