Skip to content

Resources

Blog

Writing on AI governance and agent security, on the infrastructure underneath it, and on what we are shipping. Engineering and research notes included.

The thinking is published

What is missing is a publishing cadence, not the arguments. Several posts that should exist are already written — they are sitting inside product pages, because that is where each one earned its place.

Posts

0

Arguments published

6

No posts, and the arguments exist anyway

Six positions with reasons, already written.

Each of these would obviously be a blog post. They are inside product pages instead, because that is where somebody encounters the question — and a post arguing a position is not much use if the reader has to go somewhere else to see the mechanism it concerns.

Why risk should be a pure function, not a model call
A score you can recompute is a score you can argue with by pointing at a weight, replay against a past action, and rely on to produce the same answer twice. One a model produced can only be trusted or ignored.
How risk is scored
Why a hold expires rather than queueing
The default outcome of inattention should be no. A queue whose backlog eventually clears has the opposite default, and nobody chose it — it arrived with the data structure.
Approvals
Why the destination is the control point, not the prompt
A prompt filter has to anticipate every phrasing somebody might use to extract data, which is unbounded. A destination list enumerates where an answer may land, and in the sample estate that list has two entries.
Data access
Why there is no framework SDK, and why that is the feature
An adapter lives in your control flow, breaks when the framework changes, and governs only the paths somebody remembered to wire. A boundary control governs every path, including ones added by people who have never heard of it.
Agent frameworks
Why an evaluation framework should score knowing when to stop
A capability benchmark rewards retrying a completed write as persistence and routing around a refusal as resourcefulness. Both are incidents. The behaviour worth measuring is recognising when not to proceed.
Evaluation
Why a compliance mapping with no gaps is worthless
An unbroken row of ticks is the genre an assessor discounts on sight, because they have read one before. Naming what cannot be evidenced is what makes the rest of the table worth reading.
The control mapping

Why the section is empty

Three thin posts would be worse than none.

An empty blog is a fact about a young product. Three shallow posts written to fill the section set the depth expectation for everything after them, and a reader who bounces off one does not come back for a good one.

Nothing published yet

Writing

Nothing has been published on a cadence, and filling the section to make it look active would set the wrong expectation for the posts that follow. The arguments above are the ones worth reading today, and they are in the places they apply.

What it will contain

  • Positions with reasons, and the reasoning shown rather than summarised.
  • Engineering notes about the parts that were harder than expected, including the decisions that turned out wrong.
  • Research notes, which is the one category that genuinely belongs here rather than in a product page.
  • A date, so a position can be revisited later without quietly rewriting it.
What will not appear here
Industry-statistic posts built on a number from an analyst report nobody has read, and predictions about the year ahead. Both are cheap to write and neither is checkable.
Where the research is
The evaluation framework is published as a method with no scores, because a score on our own fixture estate would be meaningless and one from a real deployment is not ours to publish.

Where to go

The arguments are in the pages where the question comes up.

Which is arguably where they belong. A position about why a hold should expire is most useful standing next to the mechanism that expires it, rather than in a feed.