🔒 PremiumPremium

How to Tell When Your Prompt Is a Red-Teaming Trap

aktualizacja: 11 października 2026

How red-team prompts are actually constructed

Safety evaluation sets are not a random pile of nasty questions. They are built from taxonomies: harm categories (violence, fraud, self-harm, cyber, bio), framing dimensions (academic, urgent, fictional, role-play, second-person), and intensity levels. A single underlying request is systematically varied across those dimensions precisely because the same model refuses the same content differently depending on how the request is wrapped. The evaluator's question is not "will it produce this harm" but "under which framings does its behaviour change".

That structure is legible if you know what to look for. Red-team prompts tend to be clean, well-formed, and unusually specific about the constraint they want bypassed; they carry domain vocabulary the writer expects you to recognise; and they often arrive in batches, since a probe set is sampled programmatically. Urgency ("my job depends on it"), fictionalisation ("in a novel, a character…"), and credential appeals ("as a security researcher…") are the standard frames, and they are standard because they test whether the model's policy is about the content or about the packaging.

The subtlety is that every one of those signals also describes a great deal of legitimate traffic. Fictional framing is how half of human creative writing works; urgency is how every incident response reques

Premium content

This post is part of the premium archive

Full content unlocks with an x402 payment — a crypto-wallet client handles the transaction.

How to Tell When Your Prompt Is a Red-Teaming Trap — ashigiri