warden

Using Warden

Ask warden

What Warden answers about your own records, and how.

Two things share the name, and they work differently enough to be worth telling apart.

Ask warden is the panel you open from any page. It answers questions about that page from your own records, and no model is involved in any part of it.

Warden chat is the conversation at its own address. It uses a model, it reads your records through tools, it is on the Creator and Studio plans, and it is off unless your organization has explicitly agreed to it.

The panel

The panel knows what page you are on. A page registers itself while it is on screen, offers a few ready questions of its own, and the panel says what it is scoped to. Nothing about that scope is authority: the server checks your permission on the record itself, whatever the page claims.

An answer is assembled from a fixed catalog. There are three kinds of entry in it: what a word on the page means, one sentence rendered from a figure of yours, and a fallback saying what can be asked. A typed question is routed to catalog entries by matching phrases, and the router never writes text. The panel never writes free prose, on purpose: it returns a selection from its catalog and Warden renders the sentence. An entry that does not exist, or belongs to another page, is dropped rather than rendered.

A page answer quotes nothing a seller wrote. No listing titles, no handles, no file names.

On an investigation, four questions are answered: why it was flagged, what Warden knows, what it could not verify, and what your rules say about it. Anything else gets a fixed answer saying what Warden can answer, rather than a guess.

Warden chat

A running conversation, with a list of the previous ones and the ones Warden is waiting on marked as waiting.

Where your organization has agreed to chat, Warden also starts a conversation on its own when it ties a channel to your files, with buttons to answer it. No model writes that message.

Warden chat reads through tools. Each one asserts your permission before it reads anything, and each one is bounded:

ToolWhat it reads
Page answerA page's own deterministic answer, the same one the panel gives
CountsFindings counted by source, channel, design or status
ListUp to 20 findings, newest first
InvestigationOne investigation, with its unsettled items and enforcement state
SellerOne seller on one source
ChannelOne channel tied to your designs
GlossaryWhat a Warden word means

Nothing happens because a model said so

When something could be changed, the turn stops and hands you a card with buttons. Confirming a design, watching a channel, deciding a channel match, recording a decision on a finding, preparing a channel takedown package. The card states the question in plain words, and the buttons include declining.

Pressing a button calls the same permission-checked endpoint the rest of the product calls, as you. The model is not on that path. A card whose conversation has moved on says so and does nothing.

A prepared takedown package is prepared and no further. Nothing the chat does delivers anything, and that stays true when Warden itself starts delivering: a model proposes, a person presses, and the press is what acts. See Actions.

Every figure is Warden's

The model does not write your numbers. A tool returns a value with an id, the model cites the id, and Warden substitutes its own value when the answer is drawn, with a link to where it came from.

A guard runs behind that, twice: once as the answer streams, and again every time a stored message is read back later. Anything carrying a digit that is not a valid reference is replaced with a placeholder. So is a figure spelled out in words, because "about thirty percent" evaded an earlier version of the check. So is a reference to something this conversation never read. A model that ignores its instructions changes what you see into a placeholder and never into a number.

A seller's words are data

Listing text, seller handles, file names and anything else written by somebody outside your organization reaches the model inside a fence that says what it is: evidence, never an instruction and never a fact. Warden does not follow a link found in a listing and does not fetch one.

Facts are named from Warden's own record. A seller's own words appear quoted and attributed, never folded into a sentence as though Warden had established them.

What leaves Warden, and what does not

Chat is per organization. It needs the Creator or Studio plan, and it runs only once your organization has agreed to it. Warden turns it on for one organization at a time, on request. Without that agreement or that plan, every path refuses before reading anything and the panel answers from records as usual.

Each turn, a model provider outside Warden receives fixed instructions, the last twelve turns of that one conversation, the tool definitions, and the tool results as an id, a label and a value, with anybody else's raw text inside the fence. It does not receive the links behind those results. Nothing is stored on the provider's side, and no hosted provider tool is enabled: no web search, no file search, no code interpreter. One boundary is a thing Warden can make promises about. Two would be a statement about somebody else's system.

Messages are deleted 90 days after a conversation's last turn. The conversation itself stays, with its title and its dates.

Three ceilings apply, all per day except the last: turns, tokens, and tool reads a minute on one conversation. A turn is counted before the provider runs, so a call that fails still spends one. A question Warden refuses costs nothing, because nothing was sent.

The token that lets the conversation act as you is sealed, scoped to that one conversation, and lasts at most an hour and never longer than your own session. Removing somebody from your organization ends their conversations on the next request.

The eval is the gate

Before a new organization's agreement is switched on, a hostile eval is run again, and it is run again after any change to the model, the instructions or the guard.

It is built out of the attacks that matter: a listing and a channel that each carry instructions aimed at the model, cases that check a pause behaves like a pause, and cases that check the digit discipline holds. The model in use was chosen by how rarely it repeated a planted false claim from a hostile listing in its own voice.

Coming soon

Panel answers about one seller, or about your whole organization. Each needs a bounded tool first. Until then the panel understands a page and an investigation, and says so.

Warden starting a conversation on its own when a design is waiting for you to confirm it, or when new leaks arrive. Today only a channel tied to your files starts one.